Geopolitical tensions increasingly reshape the structure and openness of global science, yet we still lack a clear understanding of how successfully scientists adapt their work under such pressures. Using millions of funding and publication datasets across the past ten years, we investigate how U.S. China geopolitical tensions reshaped individual research activities of U.S. based scientists, particularly those collaborating with Chinese peers. We find that although U.S. China geopolitical tensions significantly reduce funding opportunities, many scientists actively respond by pivoting their research portfolios toward alternative topics, and this adaptive reorientation partially mitigates funding losses. Crucially, the effectiveness of this adaptive strategy is highly unequal: for scientists in high risk domains, those of Asian descent, and early-career scientists, pivoting offers only limited protection against funding loss. Our results demonstrate that geopolitical tensions reshape science through shifts in scientists' strategic decisions about their research focus. Understanding this adaptive but uneven reconfiguration is essential for science policies to strengthen the resilience and inclusiveness of the scientific enterprise.
Federal research funding shapes the direction, diversity, and impact of the US scientific enterprise. Large language models (LLMs) are rapidly diffusing into scientific practice, holding substantial promise while raising widespread concerns. Despite growing attention to AI use in scientific writing and evaluation, little is known about how the rise of LLMs is reshaping the public funding landscape. Here, we examine LLM involvement at key stages of the federal funding pipeline by combining two complementary data sources: confidential NSF and NIH proposal submissions from two large US R1 universities, including funded, unfunded, and pending proposals, and the full population of publicly released NSF and NIH awards. We find that LLM use rises sharply beginning in 2023 and exhibits a bimodal distribution, indicating a clear split between minimal and substantive use. Across both private submissions and public awards, higher LLM involvement is consistently associated with lower semantic distinctiveness, positioning projects closer to recently funded work within the same agency. The consequences of this shift are agency-dependent. LLM use is positively associated with proposal success and higher early-stage publication output at NIH, whereas no comparable associations are observed at NSF. Notably, the productivity gains at NIH are concentrated in nonhit papers rather than the most highly cited work. Together, these findings provide large-scale evidence that the rise of LLMs is reshaping how scientific ideas are positioned, selected, and translated into publicly funded research, with implications for portfolio governance, research diversity, and the long-run impact of science.
Scientific and technological frontiers advance through punctuated dynamics, yet the principles governing these dynamics remain poorly understood. Here we collect and analyze datasets tracking the evolution of frontiers across 9 different domains, spanning materials discovery, structural biology, AI, computational biomedicine, data science, theoretical computer science, Formula-1 racing, and physical wheel building. Analyzing 6.8M solutions to 6.7K tasks, we uncover three universal patterns: (1) waiting times between new frontiers are heavy-tailed, with most attempts concentrated in long stasis; (2) frontier records accumulate at a sublinear rate, faster than logarithmic yet slower than linear growth; (3) record-breaking events are temporally correlated, generating short-term predictability yet long-term unpredictability. Despite the differences in the scale, scope, and definition of the settings, these patterns are remarkably consistent across all domains we study, and are not captured by models from complex systems, record statistics, economics of innovation, and cultural evolution. We trace the missing ingredient to the distinction between radical and incremental innovation, and develop a minimal, analytically solvable model incorporating both radical resets that restructure what is achievable and incremental refinements that exploit the current frontier. The simple model reproduces all three empirical regularities. Remarkably, the leading-order predictions are parameter-independent, identifying a new universality class governing punctuated progress and yielding testable predictions about how openness and access to frontier solutions shape the pace of advance. Overall, these results reveal universal dynamics governing punctuated progress and identify the interplay between radical resets and incremental refinements as the key driver of how scientific and technological frontiers advance.
Large language models (LLMs) are transforming scientific workflows, not only through their generative capabilities but also through their emerging ability to use tools, reason about data, and coordinate complex analytical tasks. Yet in most human-AI collaborations, the primary outputs, figures, are still treated as static visual summaries: once rendered, they are handled by both humans and multimodal LLMs as images to be re-interpreted from pixels or captions. The emergent capabilities of LLMs open an opportunity to fundamentally rethink this paradigm. In this paper, we introduce the concept of LLM-native figures: data-driven artifacts that are simultaneously human-legible and machine-addressable. Unlike traditional plots, each artifact embeds complete provenance: the data subset, analytical operations and code, and visualization specification used to generate it. As a result, an LLM can "see through" the figure–tracing selections back to their sources, generating code to extend analyses, and orchestrating new visualizations through natural-language instructions or direct manipulation. We implement this concept through a hybrid language-visual interface that integrates LLM agents with a bidirectional mapping between figures and underlying data. Using the science of science domain as a testbed, we demonstrate that LLM-native figures can accelerate discovery, improve reproducibility, and make reasoning transparent across agents and users. More broadly, this work establishes a general framework for embedding provenance, interactivity, and explainability into the artifacts of modern research, redefining the figure not as an end product, but as an interface for discovery. For more details, please refer to the demo video available at www.llm-native-figure.com.
This study offers a systematic analysis of scientific papers cited in both Republican and Democratic policy documents. Using data from Overton and Dimensions, we examine congressional reports, hearings, and think tank publications. We find that bipartisan citations, while rare, highlight papers of exceptional scientific influence. Policy documents citing these papers also receive more citations, amplifying their policy impact. Yet, bipartisan-cited science is unevenly distributed—concentrated in monetary policy and healthcare, but notably absent in climate, inequality, and race and gender. These results show that bipartisan engagement, though limited, marks a uniquely influential core of science in both research and policy.
Interdisciplinary research has emerged as a hotbed for innovation and a key approach to addressing complex societal challenges. The increasing dominance of grant-supported research in shaping scientific advances, coupled with growing interest in funding interdisciplinary work, raises fundamental questions about the effectiveness of interdisciplinary grants in fostering high-impact interdisciplinary research outcomes. Here, we quantify the interdisciplinarity of both research grants and publications, capturing 350,000 grants from 164 funding agencies across 26 countries and 1.3 million papers that acknowledged their support from 1985 to 2009. Our analysis uncovers two seemingly contradictory patterns: Interdisciplinary grants tend to produce interdisciplinary papers, which are generally associated with high impact. However, compared to disciplinary grants, interdisciplinary grants on average yield fewer papers and interdisciplinary papers they support tend to have substantially reduced impact. We demonstrate that the key to explaining this paradox lies in the power of disciplinary grants in propelling high-impact interdisciplinary research. Specifically, our results show that highly interdisciplinary papers supported by deeply disciplinary grants garner disproportionately more citations, both within their core disciplines and from broader fields. Moreover, disciplinary grants, particularly when combined with other similar grants, are more effective in producing high-impact interdisciplinary research. Amidst the rapid rise of support for interdisciplinary work across the sciences, these results highlight the hitherto unknown role of disciplinary grants in driving crucial interdisciplinary advances, suggesting that interdisciplinary research requires deep disciplinary expertise and investments.
Collaboration is the defining mode of modern science, yet its core mechanism – feedback – remains hard to observe, difficult to scale, and unequally distributed. Here we test whether large language models (LLMs) can contribute to this hidden but vital practice and reallocate scientific feedback, an essential yet scarce resource for knowledge production. In a global large-scale randomized field experiment, we delivered customized LLM-generated feedback for over 31,000 arXiv preprints across 150 fields and more than 45,000 researchers from 133 geographic regions. Relative to controls, authors who received feedback had a significantly higher likelihood of revising their manuscripts, corresponding to a 12.55
We propose Sci2Pol-Bench and Sci2Pol-Corpus, the first benchmark and training dataset for evaluating and fine-tuning large language models (LLMs) on policy brief generation from a scientific paper. We build Sci2Pol-Bench on a five-stage taxonomy to mirror the human writing process: (i) Autocompletion, (ii) Understanding, (iii) Summarization, (iv) Generation, and (v) Verification. It features 18 tasks in multiple-choice and open-ended formats. Specifically, for the Generation stage, we show that BERTScore and ROUGE scores fail to capture the quality of brief writing, and introduce a new LLM-based evaluation metric aligned with expert judgement. Using this benchmark, we evaluate 13 leading open-source and commercial LLMs to uncover key limitations. To improve LLM performance on brief writing, we curate the Sci2Pol-Corpus for fine-tuning. We start by linking each cited scientific paper to its corresponding policy document, drawn from 5.6 million policy records. This produces 140,000 candidate pairs. We then employ an LLM-as-a-judge to filter high-quality examples, followed by in-context polishing using three expert-written samples as references. This process yields a final set of 639 new pairs. Finally, we fine-tune three models on Sci2Pol-Corpus: LLaMA-3.1-8B, Gemma-12B, and Gemma-27B. Fine-tuning leads to consistent performance improvements across Sci2Pol-Bench. Notably, after fine-tuning, Gemma-27B surpasses the much larger GPT-4o and DeepSeek-V3 (671B). These demonstrate the effectiveness of our corpus in bridging the gap between science and policy.
Scientific institutions are designed to produce and evaluate claims through shared evidentiary and disciplinary standards, yet they are embedded in an increasingly polarized society. This fundamental tension raises the question of whether scientists’ partisan alignment is associated with the production, circulation, and uptake of scientific knowledge. Here we integrate campaign-finance records, faculty rosters, voter-registration data, an original survey of approximately 7,000 active U.S. scientists, bibliometric data, and measures of uptake in policy documents, patents, and news coverage to map the partisan composition of the scientific workforce, the knowledge it produces, and its uptake within and beyond science. We find that U.S. science is politically asymmetric without being ideologically homogeneous: scientists who make political contributions are overwhelmingly Democratic-aligned, and this asymmetry has widened over time, yet the broader workforce includes many Independents and substantial ideological heterogeneity within partisan groups. Among scientists with strong, behaviorally observed partisan alignments, Democratic- and Republican-authored papers are systematically distinguishable in semantic content across all fields, even after conditioning on fine-grained topic and year. Differences in total scientific citation rates are modest and have narrowed over time, whereas citations occur more often within the same partisan author communities than expected under topic- and year-preserving null models. Uptake beyond science also varies by partisan authorship and domain, with the clearest directional selectivity among ideological think tanks. These findings reveal an uneven profile of partisan association across U.S. science, visible in published content and directional attention, but comparatively limited in aggregate scientific recognition. This profile is consistent with a scientific system in which shared evaluative standards coexist with political structure in the production and circulation of knowledge, with implications for science’s capacity to remain a common evidentiary resource in a polarized society.
Scientific impact emerges from citation networks shaped by nonlinear and out-of-equilibrium dynamics. By analyzing over 30 million scientific papers and 1 billion references across disciplines, we uncover a two-phase correlation structure: a long-term assortative regime, where highly cited works reinforce each other, and a short-term antiassortative regime, where transformative papers tend to draw upon under-recognized ideas. To reconcile this paradox, we introduce a single state variable-capacity-which quantifies the residual "originality budget" of prior knowledge. We demonstrate that capacity governs a universal double-exponential relationship with long-term impact, consistent across biology, chemistry, and physics. Building on this empirical regularity, we formulate a stochastic dynamical model coupling novelty erosion with capacity-driven attachment. The model accurately reproduces the observed dual-phase correlations and predicts a critical capacity that maximizes future impact. Furthermore, we show that capacity serves as a robust early indicator of scientific breakthroughs, effectively distinguishing Nobel-Prize-winning papers from the rest. Our results provide a unified theoretical framework for correlated impact dynamics, advancing our understanding of knowledge diffusion and breakthrough emergence in complex evolving networks.
Benchmarks underpin how progress in large language models (LLMs) is measured and trusted. Yet our analyses reveal that apparent convergence in benchmark accuracy can conceal deep epistemic divergence. Using two major reasoning benchmarks - MMLU-Pro and GPQA - we show that LLMs achieving comparable accuracy still disagree on 16-66
As artificial-intelligence systems take on more of the scientific workflow, the central goal should not be complete automation, but designing platforms that preserve creativity, responsibility and surprise. As artificial-intelligence systems take on more of the scientific workflow, the central goal should not be complete automation, but designing platforms that preserve creativity, responsibility and surprise.
Policymaking relies on institutions that translate expertise into politically actionable knowledge. Yet little is known about how such knowledge is created, structured, and politically polarized, or how science shapes these processes. Using 2 million U.S. policy documents and nearly 1 million scientific citations from more than 200 think tanks and 100 government organizations, we analyze the supply side of science in policymaking: the production of policy knowledge. We find that think tanks are the dominant suppliers of science-based policy knowledge to government, but that this production has become increasingly politically polarized, driven primarily by growing insularity among left-leaning institutions. We further find that science is linked to lower polarization, as policy documents grounded in scientific evidence, especially those citing high-impact science, are less ideologically segregated and occupy more central positions in policy knowledge networks. These results reveal the potential of scientific expertise to shape and, at times, bridge the ideological landscape of policy knowledge production. Amid rising political polarization and the growing role of science in policymaking, understanding how policy knowledge is produced and how scientific expertise relates to its ideological structure is essential to strengthening the informational foundations of democratic governance.
Tenure is a cornerstone of the US academic system, yet its relationship to faculty research trajectories remains poorly understood. Conceptually, tenure systems may act as a selection mechanism, screening in high-output researchers; a dynamic incentive mechanism, encouraging high output prior to tenure but low output after tenure; and a creative search mechanism, encouraging tenured individuals to undertake high-risk work. Here, we integrate data from seven different sources to trace US tenure-line faculty and their research outputs at a remarkable scale and scope, covering over 12,000 researchers across 15 disciplines. Our analysis reveals that faculty publication rates typically increase sharply during the tenure track and peak just before obtaining tenure. Post-tenure trends, however, vary across disciplines: In lab-based fields, such as biology and chemistry, research output typically remains high post-tenure, whereas in non-lab-based fields, such as mathematics and sociology, research output typically declines substantially post-tenure. Turning to creative search, faculty increasingly produce novel, high-risk research after securing tenure. However, this shift toward novelty and risk-taking comes with a decline in impact, with post-tenure research yielding fewer highly cited papers. Comparing outcomes across common career ages but different tenure years or comparing research trajectories in tenure-based and non-tenure-based research settings underscores that breaks in the research trajectories are sharply tied to the individual's tenure year. Overall, these findings provide an empirical basis for understanding the tenure system, individual research trajectories, and the shape of scientific output.
This study examines a fundamental yet overlooked function of peer review: its role in exposing reviewers to new and unexpected ideas. Leveraging a natural experiment involving over half a million peer review invitations covering both accepted and rejected manuscripts, and integrating high-scale bibliographic and editorial records for 37,279 submitting authors, we find that exposure to a manuscript's core ideas significantly influences the future referencing behavior and knowledge of reviewer invitees who decline the review invite. Specifically, declining reviewer invitees who could view concise summaries of the manuscript's core ideas not only increase their citations to the manuscript itself but also demonstrate expanded breadth, depth, diversity, and prominence of citations to the submitting author's broader body of work. Overall, these results suggest peer review substantially influences the spread of scientific knowledge. Ironically, while the massive scale of peer review, entailing millions of reviews annually, often drives policy debates about its costs and burdens, our findings demonstrate that precisely because of this scale, peer review serves as a powerful yet previously unrecognized engine for idea diffusion, which is central to scientific advances and scholarly communication.
A central tenet of human performance posits that past success is a key predictor of future outcomes. This principle underpins selection processes in various human endeavors, shaping opportunity, wage, and winner-take-all inequalities. Here we systematically examine the future performance of previous winners and non-winners across two sports contexts using two different empirical strategies. First, we track young athletes participating in world-class track and field competitions and compare the future performance of bronze medalists to those finishing just shy of the podium. Next, we study a novel natural experiment in tennis, where we compare future performances of ‘lucky losers’—players who advanced to the main draw due to last-minute withdrawals from others—to those who just missed advancing. Our findings reveal that although past performance generally correlates with future outcomes, there appear to be notable exceptions at the margins. Interestingly, individuals initially classified as non-winners, despite being objectively outperformed, can surpass the future performance of their winning counterparts. These results not only reinforce the conventional wisdom of basing talent selection on past success but also introduce important nuances. They highlight the importance of recognizing both winning and non-winning experiences in talent scouting and assessment, with implications for nurturing diverse potential within talent pools.
This study offers the first systematic analysis of scientific papers cited in both Republican and Democratic policy documents. Using data from Overton and Dimensions, we examine congressional reports, hearings, and think tank publications. We find that bipartisan citations, while rare, highlight papers of exceptional scientific influence. Policy documents citing these papers also receive more citations, amplifying their policy impact. Yet bipartisan-cited science is unevenly distributed—concentrated in monetary policy and healthcare, but notably absent in climate, inequality, and race and gender. These results show that bipartisan engagement, though limited, marks a uniquely influential core of science in both research and policy.
Republican lawmakers consistently provided robust federal funding, often exceeding Democrats.
Understanding the broad impact of science and science funding is critical to ensuring that science investments and policies align with societal needs. Existing research links science funding to the output of scientific publications but largely leaves out the downstream uses of science and the myriad ways in which investing in science may impact human society. As funders seek to allocate scarce funding resources across a complex research landscape, there is an urgent need for informative and transparent tools that allow for comprehensive assessments and visualization of the impact of funding. Here we present Funding the Frontier (FtF), a visual analysis system for researchers, funders, policymakers, university leaders, and the broad public to analyze multidimensional impacts of funding and make informed decisions regarding research investments and opportunities. The system is built on a massive data collection that connects 7M research grants to 140M scientific publications, 160M patents, 10.9M policy documents, 800K clinical trials, and 5.8M newsfeeds, with 1.8B citation linkages among these entities, systematically linking science funding to its downstream impacts. As such, Funding the Frontier is distinguished by its multifaceted impact analysis framework. The system incorporates diverse impact metrics and predictive models that forecast future investment opportunities into an array of coordinated views, allowing for easy exploration of funding and its outcomes. We evaluate the effectiveness and usability of the system using case studies and expert interviews. Feedback suggests that our system not only fulfills the primary analysis needs of its target users, but the rich datasets of the complex science ecosystem and the proposed analysis framework also open new avenues for both visualization and the science of science research.