12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

LLM4SR: A Survey on Large Language Models for Scientific Research

Zonglin Yang, Zheng Xu, Wei Yang, Xinya Du · arXiv (Cornell University) · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

10/10
Relevance
2/4
Quality (LMQS)
I
Evidence
5
Citations

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.48550/arxiv.2501.04306

Methodology & findings

Study design

Comprehensive narrative literature review synthesizing prior research across LLM applications in scientific research.

Main result

The survey identifies four general tasks where LLMs have demonstrated notable potential: "scientific hypothesis discovery, where LLMs leverage existing knowledge and experimental observations to suggest novel research ideas; experiment planning and implementation, where LLMs aid in optimizing experimental design, automating workflows, and analyzing data; scientific writing, including the generation of citations, related work sections, and even drafting entire papers; and peer review, where LLMs support the evaluation of scientific papers by offering automated reviews and identifying errors or inconsistencies." Major achievements include Yang et al. [174] being "the first to demonstrate that LLMs are capable of generating novel and valid scientific hypotheses, as confirmed through expert evaluation."

Research paradigm

Interpretivist/Constructivist

Author conclusions

The authors conclude that "LLMs represent advanced productivity tools, offering new methods across all stages of modern scientific research. Despite being constrained by inherent limitations, technical barriers, and ethical considerations in domain-specific tasks, the continued advancement of LLM capabilities promises to revolutionize research practices. As these systems evolve, their integration into scientific workflows will not only accelerate discoveries but also foster unprecedented innovation and collaboration in the scientific community."

Risk of bias

Publication bias toward computer science venues (explicitly acknowledged); Potential selection bias toward published works in English; Possible geographic bias toward research published in major venues; Discipline-specific coverage limitations (focuses on documented LLM applications); Publication bias favoring computer science venues over domain-specific discipline venues; Language bias toward English-language publications; Potential selection bias in reviewing primarily cited works within computer science literature; Limited coverage of research from non-computer-science researchers; Publication bias: Likely focused on computer science venues, potentially missing works from domain scientists; Geographic bias: Emphasis on English-language publications in major conferences; Recency bias: Survey from 2025 may not capture earlier foundational work in related domains; Disciplinary bias: Heavy focus on computer science perspectives rather than integration with domain expertise

Limitations

  • The authors explicitly state: "The general concept 'AI for Science' is a huge topic, and this survey only focuses on the LLMs for scientific research aspect
  • In addition, many researchers from 'science' background but not computer science background might also have conducted works in this domain, but might not published in a computer science venues
  • We might have missed some of these works in this survey."

Open questions raised

  • Automated experimental execution - "The first line of future work is to enhance automated experimental execution, as it remains the most reliable way to test the validity of a hypothesis."
  • Enhanced LLM capabilities for hypothesis generation - "The second line of future work is to enhance the LLM's ability in hypothesis generation. Currently, it is still not very clear how to increase this ability."
  • Internal reasoning structures for scientific discovery - "The third line of future work is to investigate other internal reasoning structures of the scientific discovery process."
  • Automatic benchmark collection - "The fourth line of future work is to investigate how to leverage LLMs to automatically collect accurate and well-structured benchmark."
  • Planning capability limitations - "One fundamental limitation is their planning capability... LLMs in autonomous modes often fail to generate executable plans."
  • Prompt robustness in multi-stage contexts - "Prompt robustness poses another critical challenge in multi-stage experimental contexts."
Data: SciMON dataset - Biomedical data from 1952 to June 2022 (67,408 papers); Tomato dataset - Social Science (50 papers from January 2023); Qi et al. [119] dataset - Biomedical (2,900 papers from August 2023 test set); Kumar et al. [68] dataset - Five disciplines (100 papers from January 2022); Tomato-Chem dataset - Chemistry & Material Science (51 papers from January 2024); DiscoveryBench [108] - 264 discovery tasks from 20+ papers plus 903 synthetic tasks; DiscoveryWorld [57] - 120 different challenge tasks in virtual environment; MOPRD [94], NLPeer [33], ASAP-Review [183], PeerRead [65] - Peer review datasets; AAN, SciSummNet, S2ORC, CORWA - Related work generation datasets; SciMON dataset - NLP and Biomedical (67,408 papers, June 2022); Tomato dataset - Social Science (50 papers, January 2023); Qi et al. dataset - Biomedical (2,900 papers, August 2023); Kumar et al. dataset - Five disciplines (100 papers, January 2022); Tomato-Chem dataset - Chemistry & Material Science (51 papers, January 2024); DiscoveryBench - 264 discovery tasks from papers + 903 synthetic tasks; DiscoveryWorld - 120 challenge tasks in virtual environment; ALCE benchmark - Diverse domains with Wikipedia to web-scale corpora; CiteBench - unified benchmark for citation generation; AAN Corpus - ACL Anthology Network; SciSummNet, Delve, S2ORC, CORWA - related work generation datasets; SciGen, SciXGen - drafting and writing benchmarks; MOPRD, NLPeer, ASAP-Review, PeerRead, ReviewCritique - peer review datasets; SciMON benchmark (NLP & Biomedical, up to June 2022); Tomato dataset (Social Science, 50 papers from January 2023); Qi et al. dataset (Biomedical, 2900 papers from August 2023); Kumar et al. dataset (5 disciplines, 100 papers from January 2022); Tomato-Chem dataset (Chemistry & Material Science, 51 papers from January 2024); DiscoveryBench (264 manual discovery tasks + 903 synthetic tasks); DiscoveryWorld (120 virtual environment challenge tasks); ALEC benchmark (citation text generation); CiteBench (citation text generation); MOPRD (peer review dataset); NLPeer (peer review dataset); PeerRead (peer review dataset); PEERSUM (peer review dataset)Extracted from: pdfAgreement 78%

Explore related topics

Related papers