
We study how citizens perceive the fairness of using artificial intelligence (AI) in criminal sentencing, using a survey experiment with a representative sample in Norway ( N = 2222 ). Participants were randomly assigned to one of four experimental vignettes describing the use of AI in judicial decision-making that varied along two dimensions: human involvement (decision support vs. fully automated decision-making) and explainability (whether it is possible to determine which factors the algorithm gives the most weight to or not). In addition, respondents could be assigned to a fifth baseline condition representing the status quo, in which sentencing decisions were made solely by human judges. We find that even when AI is fully explainable and used only to support human judges, it significantly reduces perceived fairness compared to human-only decisions. This fairness gap widens with a lack of human involvement (i.e., fully automated AI system) and when decisions are non-explainable. Our results highlight that the mere involvement of AI in legal decision-making can undermine public fairness perceptions, regardless of its technical merits.
How can institutional rupture be identified in settings where the treated unit is singular, interventions are composite, and institutional deterioration unfolds through mutually reinforcing constitutional, political, and legal transformations? We develop a counterfactual framework for identifying systemic judicial rupture in single-unit settings using synthetic-control methods, generalized synthetic controls, synthetic difference-in-differences, matrix completion, placebo inference, and latent institutional measurement. We apply the framework to Venezuela following the 1999 constitutional overhaul under Hugo Chávez. Across nine dimensions of judicial integrity, we document a large, persistent, and system-wide post-1999 divergence relative to multiple counterfactual constructions, including regional and worldwide donor pools. To address concerns regarding outcome proliferation and measurement instability, we construct a latent judicial quality index using principal-components analysis and show that the estimated rupture remains exceptionally strong across all specifications. The article contributes a general empirical framework for studying institutional collapse, constitutional backsliding, and judicial capture in environments where traditional panel approaches are poorly suited to the object of inference.
With the rise of standard contracts, tensions persist between protecting consumers from unfair terms and preserving freedom of contract. A key concern is that consumers may waive legal rights, often unknowingly, particularly when dealing with trusted brands. To address this, we conducted a conjoint experiment with a representative US sample to evaluate how consumers value legal rights—such as warranties, liability, privacy, and conflict resolution—when they are fully aware, understand the terms, and must trade them off against product features like price and brand reputation. Our findings reveal that even when rights can be traded off against other features of the product, consumers value legal rights. Of the legal rights we tested, warranties and price were found to be the most decisive factors in their purchase decisions. Notably, we did not find evidence that brand reputation significantly alters the valuation of legal rights. These insights offer important empirical evidence to enrich the extensive policy debate.
Ideology is central to understanding the behavior of U.S. Supreme Court Justices. The two most widely used measures of judicial ideology—Martin–Quinn and Segal–Cover scores—each have limitations. We propose an alternative approach that leverages large language models (LLMs) to estimate ideology scores for Justices from 1925 to the present. Using ChatGPT-5, we generate static and dynamic scores for Justices identified as shifting ideologically. Our scores strongly correlate with existing measures but are not mere reproductions, and they exhibit high replicability across iterations. These findings suggest LLMs provide a promising new tool for generating valid and reliable measures of judicial ideology.
Scholars frequently argue that public perceptions of judicial nominees are shaped by two competing frames; the judiciousness frame focuses on positive aspects of a nominee that meet expectations of a good judge. The ideological frame depicts judges as political actors and is often deployed against nominees. Drawing on legal scholarship that argues that judiciousness may be easier to measure in its absence, we explore a third frame by focusing on departures from ethical expectations. This frame captures a politically potent way to portray nominees that is distinct from their legal qualifications or ideological identification. Using a conjoint experiment, we find evidence that ideological differences reduce support for judicial nominees and evidence that allegations of injudiciousness are an effective third frame that can be used in confirmation politics. Not only does the public consistently recognize and sanction injudicious behavior, but the severity of sanctions for specific injudicious behaviors varies depending on the ideology of those assessing the nominee. By bridging conceptualizations of judiciousness found in judicial politics and legal scholarship, our results highlight the importance of allegations of unethical behavior in how the public perceives individual nominees and possibly the judiciary as a whole.
Can artificial intelligence provide the reasonable justifications required in judicial opinions? This study uses a survey experiment to examine how the public perceives AI-generated judicial reasoning. Results show that participants are unable to distinguish between AI-generated and judge-authored opinions, with detection accuracy at near-perfect chance; this finding holds even among participants who work in law-related fields or who have received legal education. In blinded tests, AI-generated reasoning is rated as equal in adequacy and quality to human output. However, once an opinion is identified as AI-generated, its perceived quality declines. We also find mixed evidence of a contamination effect, where the potential involvement of AI diminishes the perceived quality of actual judge-written opinions. Exploratory analysis suggests that people’s skepticism towards AI’s participation in the judicial process stems from concerns about procedural fairness as well as doubts about AI’s capability. These findings underscore the potential risks that AI involvement could pose to judicial credibility.
Can large language models (LLMs) replace human judges? By replicating a prior 2 × 2 factorial experiment conducted on 31 U.S. federal judges, we evaluate the judicial ability of OpenAI’s GPT-4o. The experiment involves a simulated appeal in an international war crimes case, with two altered variables: the degree to which the defendant is sympathetically portrayed and the consistency of the lower court’s decision with precedent. We find that GPT-4o is a competent judge who applies precedent correctly. GPT-4o disregards the illegally irrelevant factor of sympathy, similar to students who were subjects in the same experiment but the opposite of the professional judges, who were influenced by sympathy.
We use staggered difference-in-differences and panel data methods to study the factors that predict medical malpractice (“med mal”) insurance premia, using national data on three specialties (internal medicine, general surgery, ob-gyn) from Medical Liability Monitor over 1992 to 2017. A difference-in-differences analysis of states that adopted caps during our sample period provides evidence supporting a causal link between cap adoption and higher premia, lower direct costs (payouts plus defense costs), and thus much higher profitability (proxied by the premium-to-direct-cost ratio). The savings to insurers from lower direct costs, following damage cap adoption, are at most partially reflected in premia even over long time periods. Instead, insurers in new-cap states have been able to charge apparently supra-competitive prices for a sustained period. In the panel data analysis, we estimate long run elasticities of premia to direct cost, allowing for lags of up to four years, of only around +0.40, when one might expect elasticities near one. Also, the premium-to-cost ratio, which one might expect in competitive markets to be fairly constant over time, varies widely both across states at a given time and within states across time.
While recent advances in machine learning and especially text analysis have already transformed empirical legal scholarship, previous work has mostly ignored some of the richest sources of legal data: images and audio. This article contributes to reducing this gap by introducing empirical legal scholars to computer-vision and computational-audio techniques that can be applied to the empirical legal domain. These techniques enable descriptive, causal, and predictive studies that were previously impossible due to scale and computational complexity. After reviewing general approaches to audio and visual machine learning, we illustrate the usefulness of these methods on a sample of videos from the United States Ninth Circuit Court of Appeals.
Generative AI is set to transform the legal profession, though its most promising uses and ultimate effects are still unclear. While AI models like GPT-4 improve efficiency, they can also “hallucinate” and may undermine legal judgment, particularly in complex tasks typically handled by skilled lawyers. This article examines two emerging AI innovations that may mitigate these concerns: Retrieval Augmented Generation (RAG), which grounds AI-powered analysis in legal sources, and AI reasoning models, which structure complex reasoning before generating output. We conduct the first randomized controlled trial assessing these technologies, assigning upper-level law students to complete legal tasks using a RAG-powered legal AI tool (Vincent AI, 2024), an AI reasoning model (OpenAI’s o1-preview), or no AI. We find that both AI tools significantly enhance legal work quality, a marked contrast with previous research examining older large language models like GPT-4. Moreover, these newer models appear to maintain the efficiency benefits associated with older AI technologies. Our findings also show that these AI tools significantly boost productivity in five out of six tested legal tasks, with statistically significant gains of anywhere from 50% to 130%. They perform exceptionally well in complex tasks like drafting persuasive letters and analyzing complaints. Notably, o1-preview improves the analytical depth of work product and Vincent AI avoids introducing more hallucinations, suggesting that integrating domain-specific RAG capabilities with reasoning models could yield even larger improvements.
Starting in January 2023, Louisiana and more than 20 other states passed laws requiring age verification for websites with substantial adult content. Using Google Trends data and a synthetic control design, we examine how these laws affect the public’s digital behavior across four dimensions: searches for compliant websites, non-compliant websites, VPNs, and adult content. Three months after the laws were passed, results show a 51% decrease in searches for the main compliant platform, while searches increased for both non-compliant platform (48.1%) and VPN services (23.6%). Through multiverse analyses, we demonstrate the robustness of these findings to numerous model specifications. Our findings reveal that while regulations reduce traffic to compliant sites and likely decrease overall consumption, users adapt by shifting to providers without verification requirements. This approach provides valuable insights for policymakers around the world considering similar legislative measures of digital content regulation. Our methodology also offers a framework for real-time policy evaluation in contexts with staggered implementation.
Deference by courts to democratically elected legislature is at the heart of our constitutional democracy. This paper constructs a novel database of 249 cases involving the judicial review of legislation in Canada from the inception of the Charter to present. Deference increased sharply as the Charter was introduced, but has been steadily decreasing since 2000 after the McLachlin and Wagner courts. Deference is rising for the right to free expression, but declining for penal statutes and the right to equality. The recent fall in judicial deference can largely be attributed to certain Liberal appointees striking down more penal statutes due to both construing criminal rights more broadly and, as predicted by Irwin Toy , finding that the least intrusive means have not been taken. The justice-level data also provide insights on differences (or lack thereof) in judicial behaviour across sex and politics. One could conclude that Canada does, in fact, have de facto tiered judicial review.
With the advent of remote-video technology and recent pushes to include video feeds in U.S. Supreme Court hearings, many are concerned about the effect that video and streaming might have on the behavior of U.S. judges and court participants. Previous research has shown that judges react to public sentiment, and anecdotal evidence suggests that introducing video recording might induce greater levels of “performative judging,” where behavior changes when there are cameras in the courtroom. We use AI-powered diarization to analyze the transcripts of oral arguments, leveraging the quasi-random adoption of video feeds in the U.S. Ninth Circuit Court. We find suggestive evidence that judges speak more and are more likely to interrupt attorneys when proceedings are video recorded rather than only audio recorded. However, these results are mixed and model dependent. When our estimation strategies account for interactive effects or include only judges that spoke in both audio-only and video hearings, we observe a camera effect for some outcomes of interest. By contrast, non-interactive models that include all judges result in null findings. These mixed findings provide novel but preliminary evidence exploring the performative judging hypothesis and suggest that courts and policy makers should consider the potential effects on judges’ behavior when deciding whether to introduce cameras into the courtroom.
Comparative research on law and legal institutions depends on high-quality data infrastructure. This article introduces the Australian High Court Database—a new resource that encodes structured information on all full judgments of the High Court of Australia between 1995 and 2020, and all leave applications (Australia’s equivalent to petitions for certiorari) from 2003 to 2018. The database is built in accordance with core principles that support comparative research: it is adaptable, and comparable. By attending to jurisdictional specificity while adhering to general standards, the database supports both within-country analysis and cross-national comparison. We illustrate how the Australian High Court Database can be used to study comparative judicial behavior by analyzing judicial dissent rates across apex courts, judicial ideology, and agenda setting.
This article proposes a new method of increasing public support for redistribution through tax-and-transfer law by harnessing the psychological phenomenon known as the identifiability effect . Enhancing such support is crucial, as implementing or sustaining redistributive measures is difficult in the face of unsympathetic public opinion. The article reports on the findings of three original, pre-registered experiments—conducted in Israel, England, and the United States—that substantiate the theoretical argument. The findings reveal that minimal and noninvasive identification of a single recipient can enhance people’s support for redistributive taxes—including direct income transfers—that would benefit the entire group of welfare beneficiaries. These results are relevant for the literatures on redistribution and the identifiability effect, as well as for policymakers.
In the past few years, large language models (LLMs) have achieved significant technical advances, enabling legal-advocacy organizations to adopt them as complements to—or substitutes for—lawyers and other human experts. The role of LLMs in legal education, however, is underexplored. While several studies have examined LLMs’ performance in taking law school exams, finding mixed results, there have been no published studies systematically analyzing LLMs’ competence at one of law professors’ chief responsibilities: grading law school exams. This paper presents results of an analysis of how LLMs perform in evaluating student responses to legal analysis questions of the kind typically contained in law school exams. The data come from exams in four subjects administered at top-30 U.S. law schools. Unlike some projects in computer or data science, our goal is not to design a new LLM that minimizes error or that maximizes agreement with human graders. Rather, we seek to determine whether existing models—which can be straightforwardly applied by most professors and students—are already suitable for law exam evaluation. We find that, when provided with a detailed rubric, the LLM grades correlate with the human grader at Pearson correlation coefficients of up to 0.93. Our findings suggest that, even if they do not fully replace humans in the near future, LLMs could soon be put to valuable tasks by law school professors, such as reviewing and validating professor grading, providing substantive feedback on ungraded midterms, and providing students feedback on self-administered practice exams.
The Supreme Court has recently delivered big wins for conservatives on issues such as guns, abortion, and campaign finance. In many well-known cases in the Court’s history, however, various justices have cast votes that split the difference between conservative and liberal positions. This paper develops the first systematic measure of “difference-splitting votes” at the Supreme Court and considers what predicts this voting behavior. I find that these votes have declined as the Court’s docket has shrunk. Not all justices are inclined to cast these votes: justices who are more conscientious and less extreme are more likely to engage in difference-splitting. Justices are also more likely to cast these votes in salient cases. By understanding these votes, we can learn more about how jurists craft nuanced legal outcomes and the ways justices can send signals to other actors through their voting behavior.
The growing emphasis on “fit” as a hiring criterion introduces the potential for a new, subtle form of discrimination (Bertrand & Duflo, 2017). Analysis of 1,901 U.S. Supreme Court oral arguments from 1998 to 2012 documents that voice-based snap judgments predict court outcomes. Male petitioners who rank below median in perceived masculinity are 7 percentage points more likely to win. This negative correlation between perceived masculinity and winning cases in the Supreme Court is more pronounced in masculine industries. Perceived femininity of women lawyers also predicts court outcomes. Democrats favor men with less masculine-sounding voices. Perceived masculinity explains additional variance in Supreme Court decisions beyond what is predicted by the best random forest prediction model. A de-biasing experiment using information and incentives in factorial design is consistent with misperceptions and taste for masculine-sounding lawyers explaining the negative correlation between perceived masculinity and Supreme Court wins.
How does partisan identity shape perceptions of guilt? In this paper, we examine whether a hypothetical defendant’s perceived political party identification influences jurors’ beliefs about guilt. Among both Democrats and Republicans, we find a striking pattern of bias in favor of copartisan defendants and against out-partisan defendants, with the greatest effects among those that are most affectively polarized. In explaining their decisions, both Democrat and Republican respondents utilize similar themes, but employ them differently by the perceived partisanship of the defendant. Our results shed new light on the significance of partisan bias in American society, demonstrating how partisan allegiances distort jurors’ evaluations of guilt and innocence.
Algorithms often outperform humans in making decisions, in large part because they are more consistent. Despite this, there remains widespread demand to keep a “human in the loop” to address concerns about fairness and transparency. Although evidence suggests that most human overrides are errors, we argue these errors can provide value: they generate new data from which algorithms can learn. To remain accurate, algorithms must be updated over time, but data generated solely from algorithmic decisions is biased, including only cases selected by the algorithm (e.g., individuals released on parole). Training on this algorithmically selected data can significantly reduce predictive accuracy. When a human overrides an algorithmic denial, it generates valuable training data for updating the algorithm. On the other hand, overriding a grant removes potentially useful data. Fortunately, demand for human oversight is strongest for algorithmic denials of benefits, where overrides add the most value. This alignment suggests a politically feasible and accuracy-enhancing reform: limiting human overrides to algorithmic denials. The article illustrates the accuracy-sustaining benefits of strategically keeping “error in the loop” with datasets on parole, credit, and law school admissions. In all three contexts, we demonstrate that simulated human overrides of algorithmic denials significantly improve the predictive value of newly generated data.