LLM4SR: A Survey on Large Language Models for Scientific Research
Zonglin Yang, Zheng Xu, Wei Yang, Xinya Du · arXiv (Cornell University) · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.48550/arxiv.2501.04306
Methodology & findings
Study design
Comprehensive narrative literature review synthesizing prior research across LLM applications in scientific research.
Main result
The survey identifies four general tasks where LLMs have demonstrated notable potential: "scientific hypothesis discovery, where LLMs leverage existing knowledge and experimental observations to suggest novel research ideas; experiment planning and implementation, where LLMs aid in optimizing experimental design, automating workflows, and analyzing data; scientific writing, including the generation of citations, related work sections, and even drafting entire papers; and peer review, where LLMs support the evaluation of scientific papers by offering automated reviews and identifying errors or inconsistencies." Major achievements include Yang et al. [174] being "the first to demonstrate that LLMs are capable of generating novel and valid scientific hypotheses, as confirmed through expert evaluation."
Research paradigm
Interpretivist/Constructivist
Author conclusions
The authors conclude that "LLMs represent advanced productivity tools, offering new methods across all stages of modern scientific research. Despite being constrained by inherent limitations, technical barriers, and ethical considerations in domain-specific tasks, the continued advancement of LLM capabilities promises to revolutionize research practices. As these systems evolve, their integration into scientific workflows will not only accelerate discoveries but also foster unprecedented innovation and collaboration in the scientific community."
Risk of bias
Publication bias toward computer science venues (explicitly acknowledged); Potential selection bias toward published works in English; Possible geographic bias toward research published in major venues; Discipline-specific coverage limitations (focuses on documented LLM applications); Publication bias favoring computer science venues over domain-specific discipline venues; Language bias toward English-language publications; Potential selection bias in reviewing primarily cited works within computer science literature; Limited coverage of research from non-computer-science researchers; Publication bias: Likely focused on computer science venues, potentially missing works from domain scientists; Geographic bias: Emphasis on English-language publications in major conferences; Recency bias: Survey from 2025 may not capture earlier foundational work in related domains; Disciplinary bias: Heavy focus on computer science perspectives rather than integration with domain expertise
Limitations
- The authors explicitly state: "The general concept 'AI for Science' is a huge topic, and this survey only focuses on the LLMs for scientific research aspect
- In addition, many researchers from 'science' background but not computer science background might also have conducted works in this domain, but might not published in a computer science venues
- We might have missed some of these works in this survey."
Open questions raised
- Automated experimental execution - "The first line of future work is to enhance automated experimental execution, as it remains the most reliable way to test the validity of a hypothesis."
- Enhanced LLM capabilities for hypothesis generation - "The second line of future work is to enhance the LLM's ability in hypothesis generation. Currently, it is still not very clear how to increase this ability."
- Internal reasoning structures for scientific discovery - "The third line of future work is to investigate other internal reasoning structures of the scientific discovery process."
- Automatic benchmark collection - "The fourth line of future work is to investigate how to leverage LLMs to automatically collect accurate and well-structured benchmark."
- Planning capability limitations - "One fundamental limitation is their planning capability... LLMs in autonomous modes often fail to generate executable plans."
- Prompt robustness in multi-stage contexts - "Prompt robustness poses another critical challenge in multi-stage experimental contexts."
Explore related topics
Related papers
- Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statementDavid Moher · 2009 · 83,271 citations
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviewsMatthew J. Page · 2021 · 10,956 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations