Although systematic reviews are intended to provide trusted scientific knowledge to meet the needs of decision-makers, their reliability can be threatened by bias and irreproducibility. To help decision-makers assess risks in systematic reviews they intend to use as the foundation of their action, we designed and tested a new approach to analyze the evidence selection of a review: its coverage of the primary literature and its comparison to other reviews. Our approach could also help anyone using or producing reviews understand diversity or convergence in evidence selection. The basis of our approach is a new network construct called the inclusion network, which has two types of nodes: primary study reports (PSRs, the evidence) and systematic review reports (SRRs). The approach assesses risks in a given systematic review (the target SRR) by first constructing an inclusion network of the target SRR and other systematic reviews studying similar research questions (the companion SRRs) and then applying a three-step assessment process that utilizes visualizations, quantitative network metrics, and time series analysis. This paper introduces our approach and demonstrates it in two case studies. Risks we identified: missing potentially relevant evidence, epistemic division in the scientific community, and recent instability in evidence selection standards. We also compare our inclusion network approach to knowledge assessment approaches based on another influential network construct, the claim-specific citation network, discuss current limitations of the inclusion network approach, and present directions for future work.
The Maritime Transportation System (MTS) accounts for more than 80% of global merchandise trade in volume and roughly one-sixth of the Total Gross Output of the United States. Given that national and global economies depend upon efficient supply chains, port stakeholders must develop security plans to respond to all hazards, natural and manmade. Given recent cyber-attacks affecting shipping ports, along with the multi-billion dollar cyber insurance gap, ports need to understand the tradeoffs between increased competitiveness and higher risk through investment in automation and advanced logistics technologies. This article addresses the need to understand the economic impact of cyber-attacks that affect shipping port operations and thereby enable risk assessments that holistically evaluate interactions among port Information Technology (IT) and Operational Technology (OT) systems. Using a Nearly-Orthogonal Latin Hypercube (NOLH) experimental design, we construct transportation disruption profiles based on actual cyber-attacks that specify the range of operational effects of IT/OT dependencies on stakeholder transportation assets. To capture the costs of the physical disruption, we extend Boland et al’s Dynamic Discretization Discovery (DDD) algorithm to capture capacity constraints and enable delay modeling to accommodate commodities arriving late due to disruption. Economic loss functions for seven commodity categories based on the willingness to pay literature are used to compute delay costs so that stakeholders can estimate the range of economic and operational impacts within a disruption profile. Results based on data for cyber-attacks on landlord port and terminal operator assets provided by Port Everglades, FL illustrate impacts at $80,000 and $1.2M on average during one week in October 2017 and at $141,000 and $2.8M for May 2017 respectively. The runtime performance of our enhanced DDD algorithm improves on the state of the art by an order of magnitude and on larger problem sizes based on real-world port networks.
This study used statistical topic modeling to examine 800,000 documents within HathiTrust and JSTOR databases to identify the kinds of discourses in books, poetry, newspapers, and journals related to African American women. We examined a range of conversations that emerged, between 1746 and 2014, revealing insights about, and from African American women. We identified a metadata revision methodology that served to rescue 150 documents for or about Black women that were either not previously cataloged or cataloged in such a way that Black women's experiences are either lost or erased. This project’s use of computation is unique in that it allows for the quantitative surveying of such a large dataset while charting a qualitative assessment to determine if and how texts capture the experiences of African American women. Using a technique called ‘intermediate reading’, texts are verified for their applicability. This strategy of search, recognition, rescue and recovery (SeRRR) may aid curators of information in making Black women’s voices more accessible within the digitized record. The SeRRR strategy will allow scholars to use a form of ‘call and response’ with metadata to understand the lived (and death) experiences of Black women as Alice Walker did during her search for Zora Neal Hurston.
The Maritime Transportation System is crucial to the global economy, accounting for more than 80% of global merchandise trade in volume and 67% of its value in 2017. Within the US economy alone, this system accounted for roughly a quarter of GDP in 2018. This paper defines an approach to measure the degree to which individual stakeholders, when disrupted, affect the commodity flows of other stakeholders and the entire port. Using a simulation model based on heterogeneous datasets gathered from fieldwork with Port Everglades in FL, we look at the effect of varying the timing and location of disruptions, as well as response actions, on the flow of imported commodities. Insights based upon our model inform how and when stakeholders can impact one another’s operations and should thereby provide a data-driven, strategic approach to inform the security plans of individual companies and shipping ports as a whole.
Computational analysis and digital humanities are far from neutral processes and sites unimpeded by the political, social and economic context in which they emerged and are utilized. As an interdisciplinary field, the digital humanities have transformed the relationship of humans to computers broadly conceived. At the same time, the methods, theories, perspectives and the concomitant digital tools developed are being criticized for reproducing the social divisions that exist in society. The effort to recover Black women's subjectivities from the digital minefield is not without its challenges, reflected in our study which searched approximately 800,000 books, newspapers, and articles in the HathiTrust and JSTOR Digital Libraries. The goal was to identify perceptions and lived experiences of Black women that emerged and the resulting knowledge that developed. The project team discovered multiple challenges related to the rescue and recovery of Black women's standpoints or group knowledge. This essay explores how even as computational analysis has embedded biases, it can be utilized to recover the experiences of Black women from within the digitized record. Thus, computational analysis and all that it encompasses not only makes visible Black women's experiences, but also expands the scope of the digital humanities.
Topic modeling is a widely used approach for analyzing large text collections. In particular, Latent Dirichlet Allocation (LDA) is one of the most popular topic modeling approaches to aggregate vocabulary from a document corpus to form latent "topics". However, learning meaningful topic models with massive document collections which contain millions of documents, billions of tokens is challenging, given the complexity of the data involved, the difficulty in distributing the computation across multiple computing nodes. In recent years some data processing frameworks, such as Spark, Mallet, others have been developed to address the issues associated with analyzing large volumes of unlabeled text pertaining to various domains in a scalable, efficient manner. In this paper, we will present a preliminary case study demonstrating the scholarship achieved in the study of political consumerism via XSEDE resources. The experimental study will showcase the use of digitized social sciences data, text analytics toolkits to generate topic models, visualize topics for empowering intersectional research engaging the relationship between consumption, race, class, gender in the area of sociology. Consequently, this comparative big data textual analysis involving use of JSTOR data, LDA modeling toolkit's, visualization techniques, computational components is of paramount importance, especially for researchers from academic domain dealing with social science applications involving big data.
This study employs Latent Dirichlet allocation (LDA) algorithms and comparative text mining to search 800,000 periodicals in JSTOR (Journal Storage) and HathiTrust from 1746 to 2014 to identify the types of conversations that emerge about Black women's shared experience over time and the resulting knowledge that developed called standpoint We used MALLET to interrogate various genres of text (poetry, science, psychology, sociology, African American Studies, policy, etc.). We also used comparative text mining (CTM) to explore latent themes across collections written in different time periods by analyzing the common and expert models. We used data visualization techniques, such as tree maps, to identify spikes in certain topics during various historical contexts such as slavery, reconstruction, Jim Crow, etc. We identified a subset of our corpus (20,000) comprised of articles about or by or Black women and compared patterns of words in the subset against the larger 800,000 corpus. Preliminary findings indicate that when we pulled 300,000 volumes, about 800,000 (~27%) do not have subject metadata. This appears to suggest that if a researcher searched for volumes about Black women, they may not have access to a significant amount of data on the topic. When volumes are not tagged properly, researchers would have to know that these texts exists when they do their searches. The recovery nature of this project involves identifying these untagged volumes and making the corpus publicly available to librarians and others with copyright considerations.
The scientific visualization community increasingly questions the use of rainbow colormaps. This is not unfounded as significant problems are readily seen in a luminance plot of the rainbow colormap. Many good, generally applicable colormaps are proposed as direct replacements for the rainbow. However, there are still many who choose rainbows and like them. Would a colormap with perfect luminance and the chromaticity of a rainbow find a wider audience? This was our motivation in studying the range of chromatic effects arising from luminance corrections. Consequently we developed a framework for adjusting colormaps to various degrees which produces favorable results on a wide range of colormaps. In this work we will detail this framework and demonstrate its effectiveness on several colormaps.
Electron and x-ray diffraction are well-established experimental methods used to explore the atomic scale structure of materials. In this work, a computational algorithm is presented to produce electron and x-ray diffraction patterns directly from atomistic simulation data. This algorithm advances beyond previous virtual diffraction methods by utilizing an ultra high-resolution mesh of reciprocal space which eliminates the need for a priori knowledge of the material structure. This paper focuses on (1) algorithmic advances necessary to improve performance, memory efficiency and scalability of the virtual diffraction calculation, and (2) the integration of the diffraction algorithm into a workflow across heterogeneous computing hardware for the purposes of integrating simulations, virtual diffraction calculations and visualization of electron and x-ray diffraction patterns.