Artificial Intelligence in Science: Returns, Reallocation, and Reorganization
Moh Hosseinioun, Brian Uzzi, Henrik Barslund Fosse · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Large-scale observational study using a comprehensive dataset of research proposals (both funded and unfunded) submitted to an international funding agency.
Sample
N = 18896, 7 groups
Primary method
Logistic regressions to model funding likelihood; ordinary least squares (OLS) regression models with proposal-level controls including applicant demographics (age, gender, prior experience), project length, team size, textual similarity to contemporaneous proposals, year and domain fixed effects; semantic matching procedure to construct comparable pairs of AI-enabled and non-AI proposals; Mann-Whitney U tests and Kolmogorov-Smirnov tests for task-count distribution comparisons; Jaccard similarity and Cohen's kappa for assessing LLM classification reliability; two-stage LLM classification pipeline; sentence-embedding models (paraphrase-multilingual-mpnet-base-v2) for budget categorization.
Main result
The study found that "in the short run, AI adoption is associated with modest improvements in scientific outcomes concentrated in the upper tail. Instead, its primary effects arise in the organization of research: AI-enabled projects reallocate resources toward human capital, involve larger teams, and undertake a broader set of tasks." Additionally, "AI-enabled projects reallocate funds away from equipment and operational expenses toward human capital—particularly salaries—and are undertaken by larger teams."
Reports effect sizes and confidence intervals.
Research paradigm
Positivist empirical research using large-scale observational data and computational methods
Author conclusions
The authors conclude that "the adoption of modern AI is associated with first-order gains concentrated in the upper tail of outcomes" and that "the primary effects of AI appear in the process of scientific production. AI-enabled projects reallocate inputs away from equipment and operational expenditures toward human capital, expand the scope of tasks undertaken, and rely on larger and more diverse teams." They further note that "these findings offer evidence consistent with a broader view of AI as a general-purpose technology that reorganizes, rather than immediately enhancing, the knowledge production process," and that "the limited first-order gains we observe may therefore reflect a transitional phase characterized by reallocation and reorganization rather than immediate increases in output."
Risk of bias
Selection bias from using only funded and unfunded proposals from one funding agency; Publication bias in outcome measurement (using publications as proxy for research success); Classification uncertainty for AI detection and functional role assignment (Jaccard similarity = 0.41, Cohen's κ = 0.31); Potential confounding from unobserved proposal characteristics; Differential reliability across classification stages (ideation and experimentation showed κ = 0.06-0.12); Unmeasured factors influencing funding decisions that may correlate with AI adoption; Selection bias: Only proposals submitted to one major funding agency studied; results may not generalize to other funding contexts; Publication bias: Analysis of proposals captures ex ante choices but outcomes measured through publications which may have their own selection mechanisms; Confounding bias: Unobserved confounding acknowledged by authors; proposal characteristics (quality, innovativeness) not fully captured by textual similarity matching; Measurement error: LLM classification of AI methods and functional roles shows moderate agreement (Jaccard similarity=0.41, Cohen's κ=0.31); manual review applied to high-disagreement categories (ideation, experimentation); Survivor bias: Funded proposals may differ systematically from unfunded proposals in ways related to AI adoption; Selection bias: Study uses proposals from a single large international funding agency, limiting generalizability to other funding contexts; Confounding: Unobserved confounders may influence both AI adoption and outcomes despite multiple control specifications; LLM classification reliability: Agreement between two independent LLMs showed moderate Jaccard similarity (0.41) and moderate Cohen's κ (0.31 average), with near-chance agreement for ideation and experimentation categories; Temporal bias: Cross-sectional analysis may not capture long-term effects or causal relationships; Publication bias: Outcomes measured via publications, which may be systematically different between AI and non-AI projects
Limitations
- The authors note that "although unobserved confounding cannot be ruled out, the consistency of results across these specifications supports the robustness of the empirical patterns." They also state that "the absence of clear short-term effect on scientific output" means "adoption decision by scientists may rely closely to risks of shifting to new methods, which likely vary by domain." Additionally, the study has moderate classification reliability for interpretive categories: "experimentation and ideation showed near-chance agreement (κ = 0.06 and 0.12)."
Open questions raised
- Long-run implications of AI adoption on scientific productivity and skill requirements
- How organizational changes (larger teams, expanded task scope) affect scientific collaboration structure and knowledge production
- Evolution of expertise and skill acquisition required of scientists as AI capabilities advance
- Whether future AI advances will alter project length patterns observed in current period
- Domain-specific variation in adoption risks and organizational responses
- Authors identify several future research directions: (1) understanding how advances in AI may alter project length and productivity patterns; (2) clarifying the future role of scientists as LLMs expand automation capabilities; (3) examining long-run implications for skill acquisition and expertise requirements; (4) investigating how institutional decisions around funding, evaluation, and workforce development will shape the organization of science; (5) understanding whether current patterns differ across research domains.
Explore related topics
Related papers
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- To use or not to use ChatGPT in higher education? A study of students’ acceptance and use of technologyArtur Strzelecki · 2023 · 691 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- Do AI chatbots improve students learning outcomes? Evidence from a meta‐analysisRong Wu · 2023 · 469 citations
- Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performanceYizhou Fan · 2024 · 419 citations