Public sector agencies are rapidly adopting AI systems to make critical decisions that span across services such as benefits eligibility, housing and homelessness services, child welfare services, mobility, inspections, and more. This foregrounds long‑standing HCI concerns about participatory design in settings marked by institutional constraints, public accountability, and uneven power dynamics. While there is broad agreement that community‑centered practices are necessary, practitioners still wrestle with three practical questions: (1) which participatory methods fit distinct stages of the AI lifecycle; (2) what to measure to show that engagement influenced problem framing, data, model, design, or policy; and (3) how procurement and vendor management can make participation feasible and verifiable when agencies acquire AI tools rather than build in-house. This workshop will focus on two themes: methods and measures, to identify practical ways of showing how participation shapes design and policy, and procurement and contracts, to explore how engagement requirements can be embedded in vendor agreements.
Governments are the primary providers of essential public services and are responsible for delivering them effectively. In high-stakes decision-making domains such as child welfare (CW), agencies must protect children without unnecessarily prolonging a family's engagement with the system. With growing optimism around AI, governments are pushing for its integration but concerns regarding feasibility and harms remain. Through collaborations with a large Canadian CW agency, we examined how LocalLLM and BERTopic models can track CW case progress. We demonstrate how the tools can potentially assist workers in opportunistically addressing gaps in their work by signaling case progress/deviations. And yet, we also show how they fail to detect case trajectories that require discretionary judgments grounded in social work training, areas where practitioners would actually want support to pre-emptively address substantive case concerns. We also provide a roadmap of future participatory directions to co-design language tools for/with the public sector.
Public sector agencies increasingly rely on data and information systems to demonstrate that they “support” families. Yet, the metrics that stand in for support are often misaligned with how they are lived. We draw on interviews with 75 parents involved in the child welfare system (CWS)—an entry point into the broader public safety net—and use an interpretive computational workflow that combines thematic coding, a small language model, and multidimensional scaling to examine parents’ accounts of support and unmet needs. We find anad hoc safety net in which formal services, public assistance, and caseworkers’ efforts are braided together with kin, peers, and employers, but in fragmented and conditional ways. Parents with robust informal networks are better positioned to appear engaged and to receive additional help, while those with few informal supports are more likely to be documented as non-compliant and experience further neglect. These patterns reveal information gaps between sociocentric metrics (e.g., referrals, completions) and parents’ egocentric outcomes (timeliness, trust, feasibility). We discuss implications for designing collaborative information systems that use narrative-rich qualitative accounts to uncover latent patterns in how support and neglect are experienced by parents in the public safety net.
Homelessness systems in North America adopt coordinated data-driven approaches to efficiently match support services to clients based on their assessed needs and available resources. AI tools are increasingly being implemented to allocate resources, reduce costs and predict risks in this space. In this study, we conducted an ethnographic case study on the City of Toronto's homelessness system's data practices across different critical points. We show how the City's data practices offer standardized processes for client care but frontline workers also engage in heuristic decision-making in their work to navigate uncertainties, client resistance to sharing information, and resource constraints. From these findings, we show the temporality of client data which constrain the validity of predictive AI models. Additionally, we highlight how the City adopts an iterative and holistic client assessment approach which contrasts to commonly used risk assessment tools in homelessness, providing future directions to design holistic decision-making tools for homelessness.
Data scientists often formulate predictive modeling tasks involving fuzzy, hard-to-define concepts, such as the "authenticity" of student writing or the "healthcare need" of a patient. Yet the process by which data scientists translate fuzzy concepts into a concrete, proxy target variable remains poorly understood. We interview fifteen data scientists in education (N=8) and healthcare (N=7) to understand how they construct target variables for predictive modeling tasks. Our findings suggest that data scientists construct target variables through a bricolage process, in which they use creative and pragmatic approaches to make do with the limited data at hand. Data scientists attempt to satisfy five major criteria for a target variable through bricolage: validity, simplicity, predictability, portability, and resource requirements. To achieve this, data scientists adaptively apply problem (re)formulation strategies, such as swapping out one candidate target variable for another when the first fails to meet certain criteria (e.g., predictability), or composing multiple outcomes into a single target variable to capture a more holistic set of modeling objectives. Based on our findings, we present opportunities for future HCI, CSCW, and ML research to better support the art and science of target variable construction.
AI systems are often introduced with high expectations, yet many fail to deliver, resulting in unintended harm and missed opportunities for benefit. We frequently observe significant "AI Mismatches", where the system's actual performance falls short of what is needed to ensure safety and co-create value. These mismatches are particularly difficult to address once development is underway, highlighting the need for early-stage intervention. Navigating complex, multi-dimensional risk factors that contribute to AI Mismatches is a persistent challenge. To address it, we propose an AI Mismatch approach to anticipate and mitigate risks early on, focusing on the gap between realistic model performance and required task performance. Through an analysis of 774 AI cases, we extracted a set of critical factors, which informed the development of seven matrices that map the relationships between these factors and highlight high-risk areas. Through case studies, we demonstrate how our approach can help reduce risks in AI development.
AI projects often fail due to financial, technical, ethical, or user acceptance challenges -- failures frequently rooted in early-stage decisions. While HCI and Responsible AI (RAI) research emphasize this, practical approaches for identifying promising concepts early remain limited. Drawing on Research through Design, this paper investigates how early-stage AI concept sorting in commercial settings can reflect RAI principles. Through three design experiments -- including a probe study with industry practitioners -- we explored methods for evaluating risks and benefits using multidisciplinary collaboration. Participants demonstrated strong receptivity to addressing RAI concerns early in the process and effectively identified low-risk, high-benefit AI concepts. Our findings highlight the potential of a design-led approach to embed ethical and service design thinking at the front end of AI innovation. By examining how practitioners reason about AI concepts, our study invites HCI and RAI communities to see early-stage innovation as a critical space for engaging ethical and commercial considerations together.
Local and federal agencies are rapidly adopting AI systems to augment or automate critical decisions, efficiently use resources, and improve public service delivery. AI systems are being used to support tasks associated with urban planning, security, surveillance, energy and critical infrastructure, and support decisions that directly affect citizens and their ability to access essential services. Local governments act as the governance tier closest to citizens and must play a critical role in upholding democratic values and building community trust especially as it relates to smart city initiatives that seek to transform public services through the adoption of AI. Community-centered and participatory approaches have been central for ensuring the appropriate adoption of technology; however, AI innovation introduces new challenges in this context because participatory AI design methods require more robust formulation and face higher standards for implementation in the public sector compared to the private sector. This requires us to reassess traditional methods used in this space as well as develop new resources and methods. This workshop will explore emerging practices in participatory algorithm design - or the use of public participation and community engagement - in the scoping, design, adoption, and implementation of public sector algorithms.
Caseworkers in the child welfare (CW) sector use predictive decision-making algorithms built on risk assessment (RA) data to guide and support CW decisions. Researchers have highlighted that RAs can contain biased signals which flatten CW case complexities and that the algorithms may benefit from incorporating contextually rich case narratives, i.e. - casenotes written by caseworkers. To investigate this hypothesized improvement, we quantitatively deconstructed two commonly used RAs from a United States CW agency. We trained classifier models to compare the predictive validity of RAs with and without casenote narratives and applied computational text analysis on casenotes to highlight topics uncovered in the casenotes. Our study finds that common risk metrics used to assess families and build CWS predictive risk models (PRMs) are unable to predict discharge outcomes for children who are not reunified with their birth parent(s). We also find that although casenotes cannot predict discharge outcomes, they contain contextual case signals. Given the lack of predictive validity of RA scores and casenotes, we propose moving beyond quantitative risk assessments for public sector algorithms and towards using contextual sources of information such as narratives to study public sociotechnical systems.
Research into recidivism risk prediction in the criminal justice system has garnered significant attention from HCI, critical algorithm studies, and the emerging field of human-AI decision-making. This study focuses on algorithmic crime mapping, a prevalent yet underexplored form of algorithmic decision support (ADS) in this context. We conducted experiments and follow-up interviews with 60 participants, including community members, technical experts, and law enforcement agents (LEAs), to explore how lived experiences, technical knowledge, and domain expertise shape interactions with the ADS, impacting human-AI decision-making. Surprisingly, we found that domain experts (LEAs) often exhibited anchoring bias, readily accepting and engaging with the first crime map presented to them. Conversely, community members and technical experts were more inclined to engage with the tool, adjust controls, and generate different maps. Our findings highlight that all three stakeholders were able to provide critical feedback regarding AI design and use - community members questioned the core motivation of the tool, technical experts drew attention to the elastic nature of data science practice, and LEAs suggested redesign pathways such that the tool could complement their domain expertise.
Smartphone enhances healthcare support for everyone, from local to remote patients. Recent advancements in smartphone sensors redefine their usage and the prospect of remote point-of-care tools (e.g., blood diagnostic devices), especially for low-resource settings. This paper studies the sufferings of rural people due to the limited healthcare facilities and figures out the implications. The proliferation of smartphone users suggests converting many smartphones into point-of-care diagnosis devices would be a life-saving decision. Previous studies showed smartphone’s built-in camera captures physiological features (e.g., hemoglobin) from fingertip videos captured under different lights. So, we created a mobile application and attachments (light sources) to record fingertip videos for hemoglobin level calculation. Then we collected feedback on how the rural users interacted with the application. Finally, we applied qualitative and quantitative analysis to investigate their answers. Their invaluable feedback reflected the implications of various aspects of a smartphone-based point-of-care tool. The findings unveil how rural-area people can receive a smartphone's blood diagnostic services. Our results will facilitate mobile health application designers and developers to build a smartphone-based point-of-care tool for any rural area people.
Risk assessment algorithms are being adopted by public sector agencies to make high-stakes decisions about human lives. Algorithms model "risk" based on individual client characteristics to identify clients most in need. However, this understanding of risk is primarily based on easily quantifiable risk factors that present an incomplete and biased perspective of clients. We conducted a computational narrative analysis of child-welfare casenotes and draw attention to deeper systemic risk factors that are hard to quantify but directly impact families and street-level decision-making. We found that beyond individual risk factors, the system itself poses a significant amount of risk where parents are over-surveilled by caseworkers and lack agency in decision-making processes. We also problematize the notion of risk as a static construct by highlighting the temporality and mediating effects of different risk, protective, systemic, and procedural factors. Finally, we draw caution against using casenotes in NLP-based systems by unpacking their limitations and biases embedded within them.
The increasing integration of AI and data-driven technologies in various sectors of society presents both opportunities and challenges, particularly regarding ethical, fair, and inclusive applications. This one-day workshop aims to convene researchers, practitioners, policymakers, and other community members to explore the complexities of community-driven AI with the aim to promote responsible and sustainable practices in the AI ecosystem. Drawing on interdisciplinary perspectives, participants will engage in discussions focused on themes such as understanding stakeholder relationships and positions in AI systems, evaluating community impact in AI ecosystems, examining policy and intervention implications, and devising strategies for effective partnership and collaboration with communities. The workshop will provide a platform for sharing experiences, exchanging ideas, and fostering long-lasting collaborations. By collectively addressing the challenges and opportunities in community-driven AI, we aspire to build an active community that supports ethical, inclusive, and sustainable AI-driven solutions, empowering communities in an increasingly data-driven world.
Algorithms in public services such as child welfare, criminal justice, and education are increasingly being used to make high-stakes decisions about human lives. Drawing upon findings from a two-year ethnography conducted at a child welfare agency, we highlight how algorithmic systems are embedded within a complex decision-making ecosystem at critical points of the child welfare process. Caseworkers interact with algorithms in their daily lives where they must collect information about families and feed it to algorithms to make critical decisions. We show how the interplay between systemic mechanics and algorithmic decision-making can adversely impact the fairness of the decision-making process itself. We show how functionality issues in algorithmic systems can lead to process-oriented harms where they adversely affect the nature of professional practice, and administration at the agency, and lead to inconsistent and unreliable decisions at the street level. In addition, caseworkers are compelled to undertake additional labor in the form of repair work to restore disrupted administrative processes and decision-making, all while facing organizational pressures and time and resource constraints. Finally, we share the case study of a simple algorithmic tool that centers caseworkers' decision-making within a trauma-informed framework and leads to better outcomes, however, required a significant amount of investments on the agency's part in creating the ecosystem for its proper use.
Download This Paper Open PDF in Browser Add Paper to My Library Share: Permalink Using these links will ensure access to this page indefinitely Copy URL Copy DOI
The U.S. Child Welfare System (CWS) is increasingly seeking to emulate business models of the private sector centered in efficiency, cost reduction, and innovation through the adoption of algorithms. These data-driven systems purportedly improve decision-making, however, the public sector poses its own set of challenges with respect to the technical, theoretical, cultural, and societal implications of algorithmic decision-making. To fill these gaps, my dissertation comprises four studies that examine: 1) how caseworkers interact with algorithms in their day-to-day discretionary work, 2) the impact of algorithmic decision-making on the nature of practice, organization, and street-level decision-making, 3) how casenotes can help unpack patterns of invisible labor and contextualize decision-making processes, and 4) how casenotes can help uncover deeper systemic constraints and risk factors that are hard to quantify but directly impact families and street-level decision-making. My goal for this research is to investigate systemic disparities and design and develop algorithmic systems that are centered in the theory of practice and improve the quality of human discretionary work. These studies have provided actionable steps for human-centered algorithm design in the public sector.
Child welfare (CW) agencies use risk assessment tools as a means to achieve evidence-based, consistent, and unbiased decision-making. These risk assessments act as data collection mechanisms and have been further developed into algorithmic systems in recent years. Moreover, several of these algorithms have reinforced biased theoretical constructs and predictors because of the easy availability of structured assessment data. In this study, we critically examine the Washington Assessment of Risk Model (WARM), a prominent risk assessment tool that has been adopted by over 30 states in the United States and has been repurposed into more complex algorithmic systems. We compared WARM against the narrative coding of casenotes written by caseworkers who used WARM. We found significant discrepancies between the casenotes and WARM data where WARM scores did not not mirror caseworkers' notes about family risk. We provide the SIGCHI community with some initial findings from the quantitative de-construction of a child-welfare risk assessment algorithm.
Caseworkers are trained to write detailed narratives about families in Child-Welfare (CW) which informs collaborative high-stakes decision-making. Unlike other administrative data, these narratives offer a more credible source of information with respect to workers' interactions with families as well as underscore the role of systemic factors in decision-making. SIGCHI researchers have emphasized the need to understand human discretion at the street-level to be able to design human-centered algorithms for the public sector. In this study, we conducted computational text analysis of casenotes at a child-welfare agency in the midwestern United States and highlight patterns of invisible street-level discretionary work and latent power structures that have direct implications for algorithm design. Casenotes offer a unique lens for policymakers and CW leadership towards understanding the experiences of on-the-ground caseworkers. As a result of this study, we highlight how street-level discretionary work needs to be supported by sociotechnical systems developed through worker-centered design. This study offers the first computational inspection of casenotes and introduces them to the SIGCHI community as a critical data source for studying complex sociotechnical systems.
Local governments use a wide array of software, algorithms, and data systems across domains such as policing, probation, child protective services, courts, education, public employment services, homelessness services, etc. A growing body of work in CSCW and HCI has emerged to study, design, or demonstrate the boundaries of these technologies, oftentimes working with local governments. Local governments ostensibly aim to serve the public. So, some prior work has collaborated with local governments in the name of the public interest. However, others argue that local governments primarily police poor, minoritized communities, especially with increasingly limited funding for public services such as education or housing. These tensions raise critical questions: (How) should researchers collaborate with local governments? When should we oppose governments? How do we ethically engage with communities without being extractive? In this one-day workshop, we will bring together researchers from academia, the public sector, and community organizations to first take stock of work around public interest technologies. We will reflect on critical questions to orient the future of public interest technology and how we can work with, around, or against local governments while centering impacted communities.
Computer-aided classification of breast cancer using histopathological images can play a significant role in clinical practice by detecting the distinct type of malignant and/or benign tumor. However, currently proposed deep learning models developed using the BreakHis dataset only conduct a binary classification between benign and malignant tumors, and are also scale-dependent. This study utilizes a ResNet-50 implementation to transform images from the four magnification factors such that all images can be used for training the deep neural network. This process yields a larger training set that is also scale-independent. For this paper, we utilized a dual step approach with the first pass being binary classification and the second pass being a multi-class classifier of malignant tumors that offers higher clinical utility.
Cristinel Ababei合作论文数ECE Department2