The Internet today is replete with forum sites that in the context of the deep/dark web are used to sell illicit goods, and to deal in illegal activities including human trafficking (HT). Working on the DARPA MEMEX project, our team collaborated with law enforcement to perform bulk analysis and crawl these sites and in doing so found a number of challenges. Forum sites contained pages with many links and little content and the ads the rich source of content were hidden behind link-dense hubs. To identify important links that led to ads, crawlers had to take advantage of CSS visual cues. As forum sites changed often, training a model offline would not be sufficient.We address these issues by creating ROACH, or Reinforcement-based, Online, Apprentice-Critic based focused crawling approacH. ROACH provides an online, adaptive crawling mechanism that employs static subject matter expert knowledge, with online learning based on a simplified version of reinforcement learning with back propagation. We use the widely popular apprentice-critic framework for performing this capability. ROACH is independent of any crawler implementation. The approach is scalable and accurate and overall provides better link relevancy scores than two baseline approaches including Apache Nutch. We evaluate ROACH on a dataset collected in the human trafficking domain from the DARPA MEMEX effort.
The evolution of the Internet has created an abundance of unstructured data on the web, a significant part of which is textual. The task of author profiling seeks to find the demographics of people solely from their linguistic and content-based features in text. The ability to describe traits of authors clearly has applications in fields such as security and forensics, as well as marketing. Instead of seeing age as just a classification problem, we also frame age as a regression one, but use an ensemble chain method that incorporates the power of both classification and regression to learn the author's exact age.
The knowledge we gain from research in climate science depends on the generation, dissemination, and analysis of high-quality data. This work comprises technical practice as well as social practice, both of which are distinguished by their massive scale and global reach. As a result, the amount of data involved in climate research is growing at an unprecedented rate. Climate model intercomparison (CMIP) experiments, the integration of observational data and climate reanalysis data with climate model outputs, as seen in the Obs4MIPs, Ana4MIPs, and CREATE-IP activities, and the collaborative work of the Intergovernmental Panel on Climate Change (IPCC) provide examples of the types of activities that increasingly require an improved cyberinfrastructure for dealing with large amounts of critical scientific data. This paper provides an overview of some of climate science's big data problems and the technical solutions being developed to advance data publication, climate analytics as a service, and interoperability within the Earth System Grid Federation (ESGF), the primary cyberinfrastructure currently supporting global climate research activities.
Snow cover and its melt dominate regional climate and water resources in many of the world's mountainous regions. Snowmelt timing and magnitude in mountains are controlled predominantly by absorption of solar radiation and the distribution of snow water equivalent (SWE), and yet both of these are very poorly known even in the best-instrumented mountain regions of the globe. Here we describe and present results from the Airborne Snow Observatory (ASO), a coupled imaging spectrometer and scanning lidar, combined with distributed snow modeling, developed for the measurement of snow spectral albedo/broadband albedo and snow depth/SWE. Snow density is simulated over the domain to convert snow depth to SWE. The result presented in this paper is the first operational application of remotely sensed snow albedo and depth/SWE to quantify the volume of water stored in the seasonal snow cover. The weekly values of SWE volume provided by the ASO program represent a critical increase in the information available to hydrologic scientists and resource managers in mountain regions.
Staying up to date with the latest discoveries is a challenge in any scientific field. In planetary science, new observation targets on the surface of Mars are identified and named every day, and new publications announcing new discoveries and conclusions provide frequent updates about these targets. We are constructing a system that uses information extraction and retrieval methods to mine the steadily growing body of planetary science publications about Mars surface targets and automatically construct a concise summary of what is known about each target. The Mars Target Encyclopedia will provide a central, continually updated resource for use by planetary scientists and the interested public. We describe our use of Tika, Sundance, and AutoSlog to extract and summarize information, some of the challenges associated with this domain, and our plans for maturing the system.
Abstract Biomedical information is available to research and development scientists as unstructured text in the form of scientific manuscripts and reports published in the literature and elsewhere. Scientists focused on specific research programs are burdened with surveying vast numbers of publications and reports to acquire information relevant to their efforts. Employing technology as a research aid provides a mechanism to cope with information overload that characterizes the R&D environment. Text mining can extract knowledge from large corpora of biomedical text and make it available to support scientific research and knowledge collections [1, 2] and intelligent PDF reader tools able to search content and find related articles [3] are available; however, such reader tools are typically desktop applications limited to specific platforms and data sources so they cannot easily support broad based integrated scientific search needs for a dispersed R&D organization with a wide variety of content needs. Our team has developed a web-browser based document reader with a built-in exploration tool and automatic concept extraction from biomedical text content. This provides R&D scientists with a simple tool to aid finding, reading, and exploring documents relevant to focused research objectives. The tool, Shangri-Docs, combines a document reader with automatic concept extraction and highlighting of relevant terms based on carefully selected ontologies combined with our custom corporate enterprise taxonomy. Shangri-Docs provides the ability to evaluate a wide variety of document formats (e.g. PDF, Word, PPT, text, etc.) and exploits the linked nature of the Web and personal content by performing searches on content from public sites (e.g. Wikipedia, PubMed) and privately cataloged databases simultaneously. Shangri-Docs incorporates Apache cTAKES (clinical Text Analysis and Knowledge Extraction System) [4] and Unified Medical Language System (UMLS) to automatically identify and highlight terms and concepts, such as specific pathology, disease, drug, and biological terms mentioned in the text. cTAKES was originally designed specifically to extract information from clinical medical records. We have extended cTAKES automatic knowledge extraction process to include the R&D biomedical research domain by improving the ontology guided information extraction process. Shangri-Docs could be adapted to other science fields and further customized across our R&D scientific community via our open source, cloud-based, data management system. [1] Funk, et.al., BMC Bioinformatics 2014, 15:59 doi:10.1186/1471-2105-15-59 [2] Kang et al., BMC Bioinformatics 2014, 15:64 doi:10.1186/1471-2105-15-64 [3] Utopia Documents, http://utopiadocs.com [4] Apache cTAKES, http://ctakes.apache.org Citation Format: Chris Mattmann, Lauren Intagliata, Selina Chu, Garth McGrath, Giuseppe Totaro, Daniel Civello, David Ballard, Jeffrey Long, Nipurn Doshi, Shivika Thapar, Michael Livstone, Paul Ramirez, Maureen Cronin. Shangri-Docs: a browser based tool for document exploration and automatic knowledge extraction from unstructured biomedical text. [abstract]. In: Proceedings of the 107th Annual Meeting of the American Association for Cancer Research; 2016 Apr 16-20; New Orleans, LA. Philadelphia (PA): AACR; Cancer Res 2016;76(14 Suppl):Abstract nr 5283.
Abstract Biopharmaceutical R&D organizations characterize drug candidate target effects and modes of action and create molecular models of target diseases. These data-intensive activities are informed by vast data resources including publicly available data, internally generated data and partnered private data collections. However, rapid evolution in computing, data management tools, analytical and visualization methods, the complexity of data types and the data volumes that must be accommodated present significant technical and logistic hurdles to overcome. It is particularly difficult for a geographically dispersed R&D organization to make data resources easily available to scientists for search, visualization and exploration. Nevertheless, this is required for R&D scientists to gain insight into disease and drug mechanisms and to capture the knowledge needed to sustain the scientific enterprise. Standardized commercial solutions to R&D data challenges are unattractive since they require significant resource investment in platform configuration, user-training and system maintenance. This strategy necessarily creates delay in adopting newly emerging technologies and provides incentive not to adopt alternatives due to investment in existing systems. In contrast, our solution to R&D data demands was to build a cloud-deployed data platform using state of the art tools developed and maintained by the open source software community at the Apache Software Foundation. Partnering with academic data scientists, we selected the best available tools to fit our specific needs. We integrated them into a platform accessible to our federated R&D scientific community while allowing the system to be freely modified and updated on demand to meet evolving user requirements. Priorities for our data platform are to ingest, secure and index R&D source data of all types, make these indexed data assets available to computational scientists for analysis and provide faceted search capability based on a comprehensive metadata model. Three products: LabKey server, Apache OODT and ISATools have all been combined into a scientific data management system to provide a unified data resource enhanced by a search platform powered by Apache Solr. The platform supports both internally generated data and data imported from public, contracted or partnered sources. All data are available for interactive exploration by our R&D community, accessed via integrated search, analysis and visualization tools. Deployment of this system to our R&D organization has been met with enthusiastic adoption. Feedback for improvement or requests for system enhancements and additional capabilities are rapidly addressed in this open source environment, leading to further adoption among the R&D scientists and providing the basis for accessible, stable institutional knowledge collections. Citation Format: Lauren Intagliata, Selina Chu, Garth McGrath, Giuseppe Totaro, Daniel Civello, Nipurn Doshi, Shivika Thapar, Michael Livstone, Chris Mattmann, Paul Ramirez, Maureen Cronin. A cloud-enabled open source data management platform supporting a federated research and development organization. [abstract]. In: Proceedings of the 107th Annual Meeting of the American Association for Cancer Research; 2016 Apr 16-20; New Orleans, LA. Philadelphia (PA): AACR; Cancer Res 2016;76(14 Suppl):Abstract nr 5282.
We present further work on SciSpark, a Big Data framework that extends Apache Spark's inmemory parallel computing to scale scientific computations. SciSpark's current architecture and design includes: time and space partitioning of highresolution geo-grids from NetCDF3/4; a sciDataset class providing N-dimensional array operations in Scala/Java and CF-style variable attributes (an update of our prior sciTensor class); parallel computation of time-series statistical metrics; and an interactive front-end using science (code & visualization) Notebooks. We demonstrate how SciSpark achieves parallel ingest and time/space partitioning of Earth science satellite and model datasets. We illustrate the usability, extensibility, and early performance of SciSpark using several Earth science Use cases, here presenting benchmarks for sciDataset Readers and parallel time-series analytics. A three-hour SciSpark tutorial was taught at an ESIP Federation meeting using a dozen “live” Notebooks.
Mesoscale convective systems are high impact convectively driven weather systems that contribute large amounts to the precipitation daily and monthly totals at various locations globally. As such, an understanding of the lifecycle, characteristics, frequency and seasonality of these convective features is important for several sectors and studies in climate studies, agricultural and hydrological studies, and disaster management. This study explores the applicability of graph theory to creating a fully automated algorithm for identifying mesoscale convective systems and determining their precipitation characteristics from satellite datasets. Our results show that applying graph theory to this problem allows for the identification of features from infrared satellite data and the seamlessly identification in a precipitation rate satellite-based dataset, while innately handling the inherent complexity and non-linearity of mesoscale convective systems.
: This paper outlines the creation of the Polar dataset within the TREC-Dynamic Domain track. The techniques used to create the Polar dataset fall into two basic categories: information extraction using Apache Tika and information retrieval using Apache Nutch. First, we expanded the parsing capabilities of Apache Tika, an open source framework for text and metadata extraction, to provide more searchable content within Polar data repositories. Second, we used Apache Nutch, a distributed search engine that runs on top of Apache Hadoop, to crawl three prominent Polar data repositories: the National Science Foundation Advanced Cooperative Arctic Data and Information System (ACADIS), the National Snow and Ice Data Center (NSIDC) Arctic Data Explorer (ADE), and the National Aeronautics and Space Administration Antarctic Master Directory (AMD). Because finding data is often a primary challenge in scientific discovery, the inclusion of the Polar dataset in TREC-DD helps advance science through data discovery and provides TREC-DD a new challenge in in the realm of search relevancy.
In this paper we present SciSpark, a Big Data framework that extends Apache™ Spark for scaling scientific computations. The paper details the initial architecture and design of SciSpark. We demonstrate how SciSpark achieves parallel ingesting and partitioning of earth science satellite and model datasets. We also illustrate the usability and extensibility of SciSpark by implementing aspects of the Grab 'em Tag 'em Graph 'em (GTG) algorithm using SciSpark and its Map Reduce capabilities. GTG is a topical automated method for identifying and tracking Mesoscale Convective Complexes in satellite infrared datasets.
Chris A. Mattmann合作论文数Department of Computer Science, Viterbi School of Engineering48