Transforming Science with Large Language Models: A Survey on AI-assisted Scientific Discovery, Experimentation, Content Generation, and Evaluation
Steffen Eger, Cao Yong, Jennifer D’Souza, Andreas Geiger, Christian Greisinger, Stephanie Groß et al. · ArXiv.org · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.48550/arxiv.2502.05151
Methodology & findings
Study design
Narrative survey methodology.
Primary method
The survey does not employ empirical statistical testing. It uses narrative synthesis and qualitative analysis of literature. Methods include systematic analysis of datasets, tool features (documented in Tables 1, 3, 4, 5, 6, 7), and categorization of AI approaches across five research cycle stages.
Main result
The survey identifies that "AI-powered tools are transforming these tasks by leveraging NLP, machine learning (ML), LLMs, citation and knowledge graphs (KGs) to automate the retrieval, extraction, and summarization of scientific information." Key findings include that AI tools show promise across five aspects of the research cycle (search, experimentation, content generation, multimodal content, and peer review), yet "these tools exhibit various limitations, including (i) hallucinating and fabricating content, (ii) exhibiting bias, (iii) having limited reasoning abilities, (iv) lacking proper evaluation mechanisms, and (v) posing significant environmental costs."
Reports effect sizes.
Research paradigm
Mixed (interpretive/systematic analysis of AI applications across empirical and theoretical domains)
Author conclusions
The authors conclude: "In this paper, we surveyed approaches in the area of AI4Science, with a particular focus on recent large language model-based methods. We examined five key aspects of the research cycle: (1) search, (2) experimentation and research idea generation, (3) text-based content production, (4) multimodal content production, and (5) peer review... we hope that this survey inspires new initiatives in AI4Science, driving faster, more efficient, and more inclusive scientific discovery, experimentation, reporting and content synthesis-while upholding the highest ethical standards. As the ultimate goal of science is to serve humanity, we hope these advancements will accelerate knowledge creation and enhance the accessibility and reliability of research, leading to improved healthcare, medical treatments, economic processes, among a myriad of other societal benefits."
Risk of bias
Matthew effect bias in recommendation systems (well-known researchers receiving disproportionate attention); Training data bias in LLMs affecting hypothesis generation; Affiliation bias in peer review systems (von Wedel et al. 2024 showed LLMs exhibit affiliation biases); Selection bias in datasets (mostly from arXiv, ICLR, ACL communities); Domain-specific limitations in generalization of methods; Matthew effect reinforcement: "Existing dynamics such as the Matthew effect, where well-known researchers receive disproportionate attention, might be reinforced by the AI algorithms, intensifying inequalities."; Training data bias in search algorithms; Affiliation bias in LLM peer review: "von Wedel et al. recently showed that LLMs exhibit affiliation biases when reviewing abstracts."; Limited domain representation in training datasets; Hallucination bias in citation generation; Matthew effect in AI search systems (well-known researchers receiving disproportionate attention); Training data bias in embedding and retrieval models; Reinforcement of established research paradigms by AI systems trained on existing literature; Affiliation biases in LLMs when reviewing abstracts; Limited diversity in domains studied (predominantly computer science and STEM)
Limitations
- The authors state: "One of the primary challenges is data quality and coverage gaps, as these systems often struggle with handling incomplete, non-standardized, or outdated data sources, which can lead to inaccuracies and inconsistencies in retrieved information." Additional limitations include "bias in AI models remains a critical concern, where search and ranking algorithms may introduce biases based on training data, potentially influencing the visibility of certain research areas." For peer review specifically: "the variety of scientific domains that have been studied is still limited
- As most of the works rely on data from OpenReview, most studies focus on peer review within the ICLR and ACL communities."
Open questions raised
- Future work should focus on "improving feasibility and diversity of ideas and hypotheses, incorporating real-time scientific papers, refining ideation and hypothesis generation through data inspection"
- Need for enhanced personalization in search engines adapted to user preferences
- Fostering interdisciplinary collaboration through integration of AI-powered search systems with other digital tools
- Limited exploration of scientific rigor assessment - most existing studies rely on predefined rigor checklists not easily scalable across domains
- Need for domain adaptation approaches for peer review systems
- Lack of approaches for generating slides from multiple documents
Explore related topics
Related papers
- Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statementDavid Moher · 2009 · 83,271 citations
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviewsMatthew J. Page · 2021 · 10,956 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations