Large Language Models (LLMs) excel academically but struggle with social intelligence tasks, such as creating good compromises. In this paper, we present methods for generating empathically neutral compromises between two opposing viewpoints. We first compared four different prompt engineering methods using Claude 3 Opus and a dataset of 2,400 contrasting views on shared places. A subset of the gen erated compromises was evaluated for acceptability in a 50-participant study. We found that the best method for generating compromises between two views used external empathic similarity between a compromise and each viewpoint as iterative feedback, outperforming stan dard Chain of Thought (CoT) reasoning. The results indicate that the use of empathic neutrality improves the acceptability of compromises. The dataset of generated compromises was then used to train two smaller foundation models via margin-based alignment of human preferences, improving efficiency and removing the need for empathy estimation during inference.
Although in-car touchscreens expand interaction possibilities, they risk compromising driver safety and vigilance. We propose a data- and expert-informed framework for designing adaptive touchscreens that respond to a driver’s usage profile and cognitive state, maximizing usability while mitigating safety risks. First, in a driving simulator study, we find that cognitive load slows touchscreen button selections by 20% and produced shorter, more frequent off-road glances. We also find that enlarging buttons improves selection speeds by 0.3 seconds but at the cost of requiring more display pages. Next, these findings informed a co-design session with expert in-cabin designers, generating guidelines for adaptive interfaces that balance usability and safety. These guidelines form the basis of our Profile-State Adaptive (PSA) framework, which integrates driver profiles with cognitive states to guide interface adaptations. We then extend the framework to include a quantitative Time-Cost model as well as design patterns for adaptive layouts across usage profiles and cognitive demands.
This study investigates the interplay between a driver’s cognitive load, touchscreen interactions, and driving performance. Using an N-back task to induce four levels of cognitive load, we measured physiological responses (pupil diameter, electrodermal activity), subjective workload (NASA-TLX), touchscreen performance (Fitts’ law), and driving metrics (lateral deviation, throttle control). Our results reveal significant mutual performance degradation, with touchscreen pointing throughput decreasing by over 58.1% during driving conditions and lateral driving deviation increasing by 41.9% when touchscreen interactions were introduced. Under high cognitive load, participants demonstrated a 20.2% increase in pointing movement time, 16.6% decreased pointing throughput, and 26.3% reduced off-road glance durations. We identified a prevalent "hand-before-eye" phenomenon where ballistic hand movements frequently preceded visual attention shifts. These findings quantify the impact of cognitive load on multitasking performance and demonstrate how drivers adapt their visual attention and motor-visual coordination when cognitive resources are constrained.
The heterogeneity of outcomes in behavioral research has long been perceived as a challenge for the validity of various theoretical models. More recently, however, researchers have started perceiving heterogeneity as something that needs to be not only acknowledged but also actively addressed, particularly in applied research. A serious challenge, however, is that classical psychological methods are not well suited for making practical recommendations when heterogeneous outcomes are expected. In this article, we argue that heterogeneity requires a separation between basic and applied behavioral methods, and between different types of behavioral expertise. We propose a novel framework for evaluating behavioral expertise and suggest that selective expertise can easily be automated via various machine learning methods. We illustrate the value of our framework via an empirical study of the preferences towards battery electric vehicles. Our results suggest that a basic multiarm bandit algorithm vastly outperforms human expertise in selecting the best interventions.
Objective: A study involving 7 experiment stations evaluated the effects of a second iron injection administered before weaning on growth and hematological measures of pigs. Materials and Methods: Pigs (n = 514) were given an iron injection (100-200 mg) on the first day of life. Piglets were then allotted to pairs of similar-weight, samesex siblings 3 to 5 d before weaning (on d 18-24) with one piglet from each pair receiving a second iron injection. All pigs received common station-specific postweaning diets. Data were subjected to ANOVA with the model containing the terms treatment, station, pair within station, and treatment x station interaction. Results and Discussion: Postweaning ADG was greater for the added-injection group during during 0 to 14 d after weaning, but the response (212.5 vs. 202.6 g) was largely influenced by a single station as evidenced by a treatment x station interaction. The tendency for a treatment x station interaction for overall ADG (d -4 to 28) indicated that iron status was not the most limiting factor for growth at all stations. Hemoglobin concentration was greater for the added-injection group at weaning and d 14 after weaning. Implications and Applications: An additional iron injection before weaning may lead to improved early nursery growth; however, the beneficial effects of an additional iron injection are not universal and are likely dependent on unique herd characteristics including timing and total dosage of iron injections as well as nursery diet supplementation.
Conversations around public spaces are fractured and often circle around designs and their implications for the space. Contributing to this problem is the fact that people often talk past one another. Recent advances in generative AI may help to democratize this process of designing for public spaces and enable people to meaningfully converse with one another. Here, we develop a platform that asks users to submit a photo and story about a place in their community that they cherish. The system then pairs each user with a target person whose preferences conflict with their own preferences. Users read the target person’s account and use our generative AI system to redesign a space to accommodate the target person. Preliminary studies demonstrate that the act of redesigning the space may lead to greater empathy for the target person, which may help advance conversations around public spaces.
Abstract Pelvic organ prolapse is one of the top leading causes of sow mortality today and widely considered a multifactorial issue. Previous reports observed prolapsed sows having decreased serum concentrations of various trace minerals and increased tumor necrosis factor-alpha compared with non-prolapsed sows. Additionally, collagen, hormones, and previous litter performance are other factors proposed to be associated with POP. A study was completed to determine differences in serum collagen type 1 (COL1), matrix metalloproteinase 1 (MMP-1), tissue inhibitor of metalloproteinase 1 (TIMP-1), as well as relaxin and estradiol differences in the serum concentration between prolapsed and non-prolapsed sows. Furthermore, previous litter performance, lactation days, pigs born, stillborns, mummified, and pigs weaned on prolapse rate were evaluated. The study utilized 44 prolapsed sows and 44 non-prolapsed sows of similar parity, location, management, and stage of production. Blood was collected from the prolapsed and non-prolapsed sows upon discovery. Serum samples were analyzed using radioimmunoassay for estradiol content. Additionally, COL1, MMP-1, TIMP-1, and relaxin concentrations were determined from the serum using ELISA testing. Data were analyzed using the PROC GLIMMIX procedure in SAS. Prolapsed sows had decreased (P < 0.05) COL1 and TIMP-1 concentrations in the serum compared with the non-prolapsed sows. Prolapsed sows had marginally greater (P < 0.10) MMP-1 concentration compared with non-prolapsed sows. There were no differences (P > 0.10) between prolapsed and non-prolapsed serum concentraions of sows for estradiol and relaxin. Prolapsed sows had longer lactation days (P < 0.05) in the litter before prolapsing compared with non-prolapsed sows. Additionally, stillborns in the previous litter were marginally greater (P = 0.06) in prolapsed sows when compared with the non-prolapsed sows. There were no differences (P > 0.10) in pigs born, mummified, or pigs weaned between prolapsed and non-prolapsed sows. This study suggests that there is a reduced concentration of serum COL1 and TIMP-1 found in prolapsed sows compared with non-prolapsed sows and a marginally increased MMP-1 concentration in prolapsed sows. This could influence the structural integrity of the connective tissue that supports the pelvic organs. We found no differences in serum estradiol or relaxin levels. Additionally, lactation days and stillborn pigs in the litter prior to prolapsing could be contributing factors to prolapse. This work further supports the idea that prolapse is a multifactorial issue.
Since 2014, pelvic organ prolapse (POP) has become a growing concern in the U.S. swine industry. Pelvic organ prolapse is one of the three leading causes of sow mortality today. While research has increased over the years investigating the cause and occurrence of POP, the reason for the increase in POP is not fully understood. Trace mineral status is one area of interest proposed to be associated with POP. We conducted a study to determine if there was a difference in serum trace minerals and tumor necrosis factor alpha (TNF-α) concentration between prolapsed and non-prolapsed sows. The study utilized 44 prolapsed sows and 44 non-prolapsed sows of the same parity, location, management, and stage of production. Blood was collected from the prolapsed and non-prolapsed sows upon discovery and processed at a local lab where serum was collected and stored until further analyses were performed. Serum samples were analyzed using inductively coupled plasma mass spectrometry (ICP-MS) for trace mineral content (Ca, Cu, Fe, K, Mg, Mn, Mo, P, Se, and Zn). Additionally, TNF-α concentration was determined from the serum using ELISA testing. Data were analyzed using the PROC GLIMMIX procedure in SAS. Prolapsed sows had decreased (P < 0.05) Fe (2.77 vs 3.76 ppm), Mo (9.22 vs 13.02 ppb), and Zn (1.07 vs 1.19 ppm) trace mineral concentration in the serum compared with the non-prolapsed sows. However, Mg was greater (P < 0.05) in prolapsed sows compared with non-prolapsed sows (23.24 vs 21.36 ppm). Selenium was marginally (P = 0.08) less in prolapsed sows compared with non-prolapsed sows (191 vs 206 ppb). There were no differences (P > 0.10) between prolapsed and non-prolapsed sows trace mineral concentrations for Ca (97.1 vs 99.1 ppm), Cu (2.07 vs 2.03 ppm), Mn (3.77 vs 3.72 ppb), P (52.5 vs 50.2 ppm), and K (264 vs 264 ppm). There was a difference (P < 0.05) in TNF-α concentration between the prolapsed and non-prolapsed sows (233.1 vs 178.9 pg/mL). This study suggests that there is a decreased concentration of various serum trace minerals found in prolapsed sows compared with non-prolapsed sows and an increased TNF-α response in prolapsed sows.
A growing body of research has explored how to support humans in making better use of AI-based decision support, including via training and onboarding. Existing research has focused on decision-making tasks where it is possible to evaluate “appropriate reliance” by comparing each decision against a ground truth label that cleanly maps to both the AI’s predictive target and the human decision-maker’s goals. However, this assumption does not hold in many real-world settings where AI tools are deployed today (e.g., social work, criminal justice, and healthcare). In this paper, we introduce a process-oriented notion of appropriate reliance called critical use that centers the human’s ability to situate AI predictions against knowledge that is uniquely available to them but unavailable to the AI model. To explore how training can support critical use, we conduct a randomized online experiment in a complex social decision-making setting: child maltreatment screening. We find that, by providing participants with accelerated, low-stakes opportunities to practice AI-assisted decision-making in this setting, novices came to exhibit patterns of disagreement with AI that resemble those of experienced workers. A qualitative examination of participants’ explanations for their AI-assisted decisions revealed that they drew upon qualitative case narratives, to which the AI model did not have access, to learn when (not) to rely on AI predictions. Our findings open new questions for the study and design of training for real-world AI-assisted decision-making.
Visualizations are common methods to convey information but also increasingly used to spread misinformation. It is therefore important to understand the factors people use to interpret visualizations. In this paper, we focus on factors that influence interpretations of scatter plots, investigating the extent to which common visual aspects of scatter plots (outliers and trend lines) and cognitive biases (people's beliefs) influence perception of correlation trends. We highlight three main findings: outliers skew trend perception but exert less influence than other points; trend lines make trends seem stronger but also mitigate the influence of some outliers; and people's beliefs have a small influence on perceptions of weak, but not strong correlations. From these results we derive guidelines for adjusting visual elements to mitigate the influence of factors that distort interpretations of scatter plots. We explore how these guidelines may generalize to other visualization types and make recommendations for future studies.
Ridesharing is a popular choice for personal transportation needs. Although more ecologically-friendly than single-occupancy vehicles, there is an opportunity to further reduce CO2 emissions by offering green choices. Here we examine whether providing people with information about CO2 emissions nudges them to make more eco-friendly rideshare decisions. Our study tested what kind of information works best to inform people about carbon emissions, comparing direct CO2 values with more relatable carbon equivalents (e.g., trees). We conducted an online study with 1000 participants who picked between regular and eco-friendly ride options that detailed various carbon-output equivalency interventions (e.g., pounds of coal, number of smartphones charged, etc.). We found that participants are more likely to choose a green ride when presented with information about direct CO2 emissions than when presented with carbon-equivalencies. This study aims to inform future information-based interventions more broadly, beyond the context of ridesharing.
People consider recommendations from AI systems in diverse domains ranging from recognizing tumors in medical images to deciding which shoes look cute with an outfit. Implicit in the decision process is the perceived expertise of the AI system. In this paper, we investigate how people trust and rely on an AI assistant that performs with different levels of expertise relative to the person, ranging from completely overlapping expertise to perfectly complementary expertise. Through a series of controlled online lab studies where participants identified objects with the help of an AI assistant, we demonstrate that participants were able to perceive when the assistant was an expert or non-expert within the same task and calibrate their reliance on the AI to improve team performance. We also demonstrate that communicating expertise through the linguistic properties of the explanation text was effective, where embracing language increased reliance and distancing language reduced reliance on AI.
A total of 5 experiments were used to determine the relationship between nursery start weight and nursery/finishing performance traits. In each experiment, 48 pens containing 11 pigs/pen (0.65 m2/pig) were utilized. Age entering the nursery was 18-21 d and average BW was 5.9 kg. Pigs were blocked by BW and allotted to pens upon arrival and fed diets meeting or exceeding NRC (2012) recommendations. The nursery phase lasted between 42-45 d and pigs were marketed at a target weight of 136 kg. Pigs remained in the same pens from weaning to market. Data were analyzed using PROC REG/PROC CORR procedure in SAS. Pen served as experimental unit with a total of 240 observations. Nursery BW upon entry had positive (P < 0.01) correlations (r) with nursery end BW (0.57), ADG (0.23), and ADFI (0.30). There was a linear relationship between nursery starting BW and ending nursery BW [Ending BW = 1.09(nursery start BW, kg) + 9.0, R2 = 0.32, P < 0.001]. Starting nursery BW also was positively (P < 0.01) correlated with finishing end BW (0.58), ADG (0.21), ADFI (0.30), and negatively correlated with G:F (-0.25). A linear relationship was noted between starting nursery BW and market BW [Market BW = 3.93(nursery start BW, kg) + 85.6, R2 = 0.34, P < 0.001]. Linear relationships (P < 0.01) also were noted for nursery starting BW on finishing ADG, ADFI, and G:F. Finishing starting BW was positively correlated (P < 0.01) with market BW (0.52), ADG (0.32), and ADFI (0.37) and negatively correlated with G:F (-0.23). Starting finishing BW had a linear relationship with market BW [Market BW = 0.83(finishing start BW, kg) + 94.5, R2 = 0.27, P < 0.001]. Nursery starting BW has positive effects on growth performance in nursery and finishing phases as well as market weight.
Identifying personalized interventions for an individual is an important task. Recent work has shown that interventions that do not consider the demographic background of individual consumers can, in fact, produce the reverse effect, strengthening opposition to electric vehicles. In this work, we focus on methods for personalizing interventions based on an individual's demographics to shift the preferences of consumers to be more positive towards Battery Electric Vehicles (BEVs). One of the constraints in building models to suggest interventions for shifting preferences is that each intervention can influence the effectiveness of later interventions. This, in turn, requires many subjects to evaluate effectiveness of each possible intervention. To address this, we propose to identify personalized factors influencing BEV adoption, such as barriers and motivators. We present a method for predicting these factors and show that the performance is better than always predicting the most frequent factors. We then present a Reinforcement Learning (RL) model that learns the most effective interventions, and compare the number of subjects required for each approach.
John Adcock合作论文数FXPAL16