Large Language models have demonstrated promising performance in research ideation across scientific domains. Hypothesis development, the process of generating a highly specific declarative statement connecting a research idea with empirical validation, has received relatively less attention. Existing approaches trivially deploy retrieval augmentation and focus only on the quality of the final output ignoring the underlying reasoning process behind ideation. We present HypER ( Hyp othesis Generation with E xplanation and R easoning), a small language model (SLM) trained for literature-guided reasoning and evidence-based hypothesis generation. HypER is trained in a multi-task setting to discriminate between valid and invalid scientific reasoning chains in presence of controlled distractions. We find that HypER outperformes the base model, distinguishing valid from invalid reasoning chains (+22% average absolute F1), generates better evidence-grounded hypotheses (0.327 vs. 0.305 base model) with high feasibility and impact as judged by human experts ( > 3.5 on 5-point Likert scale).
Automatic medical text simplification can assist providers with patient-friendly communication and make medical texts more accessible, thereby improving health literacy. But curating a quality corpus for this task requires the supervision of medical experts. In this work, we present Med-EASi (Medical dataset for Elaborative and Abstractive Simplification), a uniquely crowdsourced and finely annotated dataset for supervised simplification of short medical texts. Its expert-layman-AI collaborative annotations facilitate controllability over text simplification by marking four kinds of textual transformations: elaboration, replacement, deletion, and insertion. To learn medical text simplification, we fine-tune T5-large with four different styles of input-output combinations, leading to two control-free and two controllable versions of the model. We add two types of controllability into text simplification, by using a multi-angle training approach: position-aware, which uses in-place annotated inputs and outputs, and position-agnostic, where the model only knows the contents to be edited, but not their positions. Our results show that our fine-grained annotations improve learning compared to the unannotated baseline. Furthermore, our position-aware control enhances the model's ability to generate better simplification than the position-agnostic version. The data and code are available at https://github.com/Chandrayee/CTRL-SIMP.
In this case study, we present a mixed methodology approach to needfinding, integrating in-depth qualitative interview data with machine learning-powered analysis of a larger dataset. The research is motivated by the high failure rates and low involvement of consumers in the food startup industry’s product design process. To help food startups design products in a more consumer-friendly and timely manner, we are developing a novel framework called the Food Personality Framework (FPF). This framework categorizes eaters according to their eating habits, preferences, motivations and constraints. To better understand the complex relationships between motivations, we chose Grounded Theory as the most pertinent approach and interviewed 14 singles with full autonomy over their food choices. We further leveraged the availability of large online food-related datasets to inform and reinforce our findings from the qualitative work. We analyzed 6687 user behaviors of Food.com, a popular recipe recommendation site, according to the 18 influencers of food choice identified from the qualitative interviews. We found three meaningful clusters of user behavior: Minimalists, Social Butterflies and Conscious eaters. The interview data, thus, enabled grounded classification of the large scale user behavior and provided a grounded way to interpret the relationship among the top motivators identified in each clusters. The cluster analysis will inform future sampling of interviewees and will provide new insightful questions for the qualitative research. The case study delineates a dynamic interplay of qualitative and quantitative data used to investigate human food choice, a novel domain in the Human-Food Interaction literature.
Health Literacy is the degree to which individuals can comprehend basic health information needed to make appropriate health decisions. The topmost reason for low health literacy is the vocabulary gap between providers and patients. Automatic medical text simplification can contribute to improving health literacy by assisting providers with patientfriendly communication, improving health data search, and making online medical texts more accessible. It is, however, extremely challenging to curate quality corpus for this natural language processing (NLP) task. In this position paper, we observe that, despite recent research efforts, existing open corpora for medical text simplification are poor in quality and size. In order to match the progress in general text simplification and style transfer, we must leverage careful crowdsourcing. We discuss the challenges of naive crowd-sourcing. We propose that careful crowd-sourcing for medical text simplification is possible, when combined with automatic data labeling, a well-designed expert-layman collaboration framework, and context-dependent crowd-sourcing instructions. Low health literacy has been associated with non-adherence to treatment plans and regimens, poor patient self-care, lack of timely communication of health issues, and increased risk of hospitalization and mortality (King 2010). Simplification of medical documents, of online communications like email messages and patient instructions can go a long way to mitigate health literacy challenges. While the consumer versions of medical journals, news articles, and a few trusted websites (NIA 2018; Savery et al. 2020) are written by trained experts, they are by no means exhaustive. Automated approaches are necessary to keep pace with the rapidly growing body of biomedical literature. In this work, we evaluate some of the open corpora that power automated text simplification in the medical domain. We define text simplification, following Siddharthan (2014), as the process of reducing the linguistic complexity of a text, while still retaining the original information content and meaning. A domain-specific expert text undergoes various kinds of transformations to reach the final simple Copyright © 2021for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0). form. Research in automatic non-medical text simplification has been burgeoning, with the introduction of large parallel corpora (Zhu, Bernhard, and Gurevych 2010; Woodsend and Lapata 2011; Coster and Kauchak 2011; Xu, CallisonBurch, and Napoles 2015; Paetzold and Specia 2017). Creation of multi-references enabled models that can learn different kinds of textual transformations separately, viz. lexical changes (e.g. paraphrasing), syntactic modifications (e.g. reordering of concepts, splitting texts, reducing sentence length etc.) and compression (e.g. deleting peripheral information irrelevant to the target domain) (Alva-Manchego et al. 2020). References are gold standard human generated simplifications, used to validate model outputs. The success of the automatic text simplification and style transfer hinges on large amounts of crowd-sourced multiple references. However, crowd-sourcing even a single set of references for medical texts is challenging. It requires the recruitment of a specific sub-population with a certain degree of domain expertise. For example, Nye et al. (2018) described an elaborate process of recruiting MDs and medical experts from Upwork, for PICO data annotation. Naturally, we observe a dearth of high-quality parallel training corpus in medical AI. Furthermore, text simplification task has additional challenges. Only the expert knows what content of the domain-specific text is relevant to the laymen, whereas the laymen or medical writers trained to translate medical texts can judge the quality and accessibility of the simplified versions. In this work, we make the following contributions: • identify the open-source datasets for medical text simplification • characterize the datasets by their quantity, quality, diversity, and representativeness • identify challenges of scaling high-quality corpus generation for medical text simplification Assumptions: We treat summarization as a subset of text simplification. We only consider corpora that represent composite textual transformations (simple text is derived after a combination of syntactic, semantic, thematic, and lexical transformations of the expert text) (Lyu et al. 2021) for further analysis. Datasets for Medical Text Simplification Datasets for medical text simplification support two kinds of document simplification: sentence-level and paragraphlevel. We focus on sentence-level and short paragraph-level simplification. After an elaborate search, we found three datasets in English for medical text simplification: two parallel corpora SIMPWIKI (Van den Bercken, Sips, and Lofi 2019) and PARASIMP (Devaraj et al. 2021), and one nonparallel corpus MSD (Cao et al. 2020). Next, we delve deeper into how these datasets are created and the potential artifacts of the data collection and annotation processes. Artifacts of Corpus Curation In the absence of reliable crowd-sourcing of medical texts, researchers resort to crawling medical websites. The expert texts are sampled from the online articles and checked posthoc for adequate corpus representativeness. The layman texts are retrieved from the layman or consumer versions of the professional articles, based on the alignment of section titles and text content. The alignment is either checked manually for a small fraction of the corpus or automatically derived using different algorithms. Only a few of the automatically aligned pairs are validated by the experts. Automatic alignment is not always reasonable (Alva-Manchego, Scarton, and Specia 2020). Random sampling of expert texts from larger articles and unreliable automatic retrieval can lead to text pieces that are not stand-alone (Choi et al. 2021). We found that the process of expert verification is insufficient for quality data curation and could still lead to pairs lacking correspondence. On the other end, models trained using highly aligned text pairs may exhibit limited generalizability. A more recent trend is to generate large volumes of non-parallel corpus, obviating validation of automatically aligned pairs. This follows similar approaches in nonmedical text style transfer (Shen et al. 2017; He and McAuley 2016; Madaan et al. 2020). Some researchers distinguish between text simplification and text style transfer tasks. We consider text simplification as a sub-domain of text style transfer where the goal is to transform text from the expert style to the layman style.
Enabling robots to act according to human preferences across diverse environments is a crucial task, extensively studied by both roboticists and machine learning researchers. To achieve it, human preferences are often encoded by a reward function which the robot optimizes for. This reward function is generally static in the sense that it does not vary with time or the interactions. Unfortunately, such static reward functions do not always adequately capture human preferences, especially, in non-stationary environments: Human preferences change in response to the emergent behaviors of the other agents in the environment. In this work, we propose learning reward dynamics that can adapt in non-stationary environments with several interacting agents. We define reward dynamics as a tuple of reward functions, one for each mode of interaction, and mode-utility functions governing transitions between the modes. Reward dynamics thereby encodes not only different human preferences but also how the preferences change. Our contribution is in the way we adapt preference-based learning into a hierarchical approach that aims at learning not only reward functions but also how they evolve based on interactions. We derive a probabilistic observation model of how people will respond to the hierarchical queries. Our algorithm leverages this model to actively select hierarchical queries that will maximize the volume removed from a continuous hypothesis space of reward dynamics. We empirically demonstrate reward dynamics can match human preferences accurately.
We focus on learning the desired objective function for a robot. Although trajectory demonstrations can be very informative of the desired objective, they can also be difficult for users to provide. Answers to comparison queries, asking which of two trajectories is preferable, are much easier for users, and have emerged as an effective alternative. Unfortunately, comparisons are far less informative. We propose that there is much richer information that users can easily provide and that robots ought to leverage. We focus on augmenting comparisons with feature queries, and introduce a unified formalism for treating all answers as observations about the true desired reward. We derive an active query selection algorithm, and test these queries in simulation and on real users. We find that richer, feature-augmented queries can extract more information faster, leading to robots that better match user preferences in their behavior.
We focus on learning the desired objective function for a robot. Although trajectory demonstrations can be very informative of the desired objective, they can also be difficult for users to provide. Answers to comparison queries, asking which of two trajectories is preferable, are much easier for users, and have emerged as an effective alternative. Unfortunately, comparisons are far less informative. We propose that there is much richer information that users can easily provide and that robots ought to leverage. We focus on augmenting comparisons with feature queries, and introduce a unified formalism for treating all answers as observations about the true desired reward. We derive an active query selection algorithm, and test these queries in simulation and on real users. We find that richer, feature-augmented queries can extract more information faster, leading to robots that better match user preferences in their behavior.
In the age of autonomous driving, researchers and companies are getting ever-so-close to enabling cars to generate driving behavior that reaches the destination and satisfies safety constraints, like not colliding with other cars or pedestrians. Imagine we got there. Initially, cars might be able to generate only one solution that satisfies these constraints. But really, many solutions exist – there are many ways to drive. We have an existence proof for that. Some of us are more aggressive drivers, valuing efficiency and being comfortable getting close to other cars on the road. Others are more defensive, a bit more conservative when it comes to safety, leaving a large distance to the next car for example, or quickly braking when someone attempts to merge in front. Soon after we are able to generate one feasible solution, we will be asking ourselves which solution we should try to generate: what driving style should an autonomous car have? In this project, we think the answer is simple: the car should drive like the end-user wants it to:
Several ongoing research projects in Human autonomous car interactions are addressing the problem of safe co-existence for human and robot drivers on road. Automation in cars can vary across a continuum of levels at which it can replace manual tasks. Social relationships like anthropomorphic behavior of owners towards their cars is also expected to vary according to this spectrum of autonomous decision making capacity. Some researchers have proposed a joint cognitive model of a human-car collaboration that can make the best of the respective strengths of humans and machines. For a successful collaboration, it is important that the members of this human - car team develop, maintain and update each others behavioral models. We consider mutual trust as an integral part of these models. In this paper, we present a review of the quantitative models of trust in automation. We found that only a few models of humans’ trust on automation exist in literature that account for the dynamic nature of trust and may be leveraged in human car interaction. However, these models do not support mutual trust. Our review suggests that there is significant scope for future research in the domain of mutual trust modeling for human car interaction, especially, when considered over the lifetime of the vehicle. Hardware and computational framework (for sensing, data aggregation, processing and modeling) must be developed to support these adaptive models over the operational phase of autonomous vehicles. In order to further research in mutual human - automation trust, we propose a framework for integrating Mutual Trust compu- tation into standard Human - Robot Interaction research platforms. This framework includes User trust and Agent trust, the two fundamental components of Mutual trust. It allows us to harness multi-modal sensor data from the car as well as from the user’s wearable or handheld device. The proposed framework provides access to prior trust aggregate and other cars’ experience data from the Cloud and to feature primitives like gaze, facial expression, etc. from a standard low-cost Human - Robot Interaction platform.
Occupancy count in rooms is valuable for applications such as room utilization, opportunistic meeting support, and efficient heating-cooling operations. Few buildings, however, have the means of knowing occupancy beyond simple binary presence-absence. In this paper we present the PerCCS algorithm that explores the possibility of estimating person count from CO2 sensors already integrated in everyday room air-conditioning infrastructure. PerCSS uses task-driven Sparse Non-negative Matrix Factorization (SNMF) to learn a nonnegative low-dimensional representation of the CO2 data in the preprocessing stage. This denoised CO2 acts as the predictor variable for estimating occupancy count using Ensemble Least Square Regression. We tested the algorithm to estimate 15 minutes average occupancy count from a classroom of capacity 42 and compared its performance against existing methods from the literature. PerCSS estimates occupancy with a normalized mean squared error (NMSE) of 0.075 and outperformed our comparative methods in predicting occupancy count with 91 % and 15 % for exact occupancy estimation, when the room was unoccupied and occupied respectively, whereas the competing methods failed mostly.
Indoor tracking has all-pervasive applications beyond mere surveillance, for example in education, health monitoring, marketing, energy management and so on. Image and video based tracking systems are intrusive. Thermal array sensors on the other hand can provide coarse-grained tracking while preserving privacy of the subjects. The goal of the project is to facilitate motion detection and group proxemics modeling using an 8 x 8 infrared sensor array. Each of the 8 x 8 pixels is a temperature reading in Fahrenheit. We refer to each 8 x 8 matrix as a scene. We collected approximately 902 scenes with different configurations of human groups and different walking directions. We infer direction of motion of a subject across a set of scenes as left-to-right, right-to-left, up-to-down and down-to-up using cross-correlation analysis. We used features from connected component analysis of each background subtracted scene and performed Support Vector Machine classification to estimate number of instances of human subjects in the scene.
Wireless sensor networks (WSN) have great potential to enable personalized intelligent lighting systems while reducing building energy use by 50%-70%. As a result WSN systems are being increasingly integrated in state-of-art intelligent lighting systems. In the future these systems will enable participation of lighting loads as ancillary services. However, such systems can be expensive to install and lack the plug-and-play quality necessary for user-friendly commissioning. In this paper we present an integrated system of wireless sensor platforms and modeling software to enable affordable and user-friendly intelligent lighting. It requires similar to 60% fewer sensor deployments compared to current commercial systems. Reduction in sensor deployments has been achieved by optimally replacing the actual photo-sensors with real-time discrete predictive inverse models. Spatially sparse and clustered sub-hourly photo-sensor data captured by the WSN platforms are used to develop and validate a piece-wise linear regression of indoor light distribution. This deterministic data-driven model accounts for sky conditions and solar position. The optimal placement of photo-sensors is performed iteratively to achieve the best predictability of the light field desired for indoor lighting control. Using two weeks of daylight and artificial light training data acquired at the Sustainability Base at NASA Ames, the model was able to predict the light level at seven monitored workstations with 80%-95% accuracy. We estimate that 10% adoption of this intelligent wireless sensor system in commercial buildings could save 0.2-0.25 quads BTU of energy nationwide.
Solar power prediction remains an important challenge for renewable energy integration primarily due to its inherent variability and intermittency. In this work, a neural network based solar power forecasting framework is developed for the NASA Ames Sustainability Base (SB) solar array using the publicly available National Oceanic and Atmospheric Administration (NOAA) weather data forecasts. The prediction inputs include temperature, irradiance and wind speed obtained through the NOAA NOMADS server in real-time. The neural network (ANN) is trained and tested on input-output data from on-site sensors. The NOAA archived forecast data is then input to the trained ANN model to predict power output spanning over nine months (June 2013-March 2014). The efficacy of the model is determined by comparing predicted power output against on-site sensor data.
Studies show that if we retrofit all the lighting systems in the buildings of California with dimming ballasts, then it would be possible to obtain a 450 MW of regulation, 2.5 GW of nonspinning reserve, and 380 MW of contingency reserve from participation of lighting loads in the energy market. However, in order to guarantee participation, it will be important to monitor and model lighting demand and supply in buildings. To this end, wireless sensor and actuator networks have proven to bear a great potential for personalized intelligent lighting with reduced energy use at 50%-70%. Closed-loop control of these lighting systems relies upon instantaneous and dense sensing. Such systems can be expensive to install and commission. In this paper, we present a sensor-based intelligent lighting system for future grid-integrated buildings. The system is intended to guarantee participation of lighting loads in the energy market, based on predictive models of indoor light distribution, developed using sparse sensing. We deployed ~92 % fewer sensors compared with state-of-art systems using one photosensor per luminaire. The sensor modules contained small solar panels that were powered by ambient light. Reduction in sensor deployments is achieved using piecewise linear predictive models of indoor light, discretized by clustering for sky conditions and sun positions. Day-ahead daylight is predicted from forecasts of temperature, humidity, and cloud cover. With two weeks of daylight and artificial light training data acquired at the sustainability base at NASA Ames, our model was able to predict the illuminance at seven monitored workstations with 80%-95% accuracy. Moreover, our support vector regression model was able to predict day-ahead daylight at ~92% accuracy.
U.S.- India Joint Center for Building Energy Research and Development (CBERD) Enabling Efficient, Responsive, and Resilient Buildings: Collaboration Between the United States and India Chandrayee Basu and Girish Ghatikar Lawrence Berkeley National Laboratory Prateek Bansal Johns Hopkins University Presented at the IEEE Great Lakes Symposium on Smart Grid and the New Energy Economy 2012, Chicago, IL September 23-25, 2013 and Published in the Proceedings March 2014
by energy, the additional hourly variation is less than 0.5percent and 4.5percent of the peak demand respectively for 99percent of the time. For wind penetration of 15percent by energy, Tamil Nadu system is found to be capable of meeting the additional ramping requirement for 98.8percent of the time. Potential higher uncertainty in net load compared to load is found to have limited impact on ramping capability requirements of the system if coal plants can me ramped down to 50percent of their capacity. Load and wind aggregation in Tamil Nadu and Karnataka is found to lower the variation by at least 20percent indicating the benefits geographic diversification. These findings suggest modest additional flexible capacity requirements and costs for absorbing variation in wind power and indicate that the potential capacity support (if wind does not generate enough during peak periods) may be the issue that has more bearing on the economics of integrating wind
Smart lighting systems in low energy commercial buildings can be expensive to implement and commission. Studies have also shown that only 50% of these systems are used after installation, and those used are not operated at full capacity due to inadequate commissioning and lack of personalization. Wireless sensor networks (WSN) have great potential to enable personalized smart lighting systems for realtime model predictive control of integrated smart building systems. In this paper we present a framework for using a WSN to develop a real-time indoor lighting inverse model as a piecewise linear function of window and artificial light levels, discretized by sub-hourly sun angles. Applied on two days of daylight and ten days of artificial light data, this model was able to predict the light level at seven monitored workstations with accuracy sufficient for daylight harvesting and lighting control around fixed work surfaces. The reduced order model was also designed to be used for long term evaluation of energy and comfort performance of the predictive control algorithms. This paper describes a WSN experiment from an implementation at the Sustainability Base at NASA Ames, a living laboratory that offers opportunities to test and validate information-centric smart building control systems.