From Text to Discovery: How Large Language Models Are Accelerating and Complicating Research Across Scientific and Humanistic Disciplines
Saleh Afroogh, Yasser Pouresmaeil, Yiming Xu, Kevin Chen, Abhejay Murali, Junfeng Jiao (7040318) · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Systematic scoping review following PRISMA guidelines.
Main result
The analysis reveals a consistent pattern: "LLMs meaningfully accelerate research workflows — from hypothesis generation and literature synthesis to data analysis and scientific writing — while introducing serious challenges related to hallucination, reproducibility, dataset bias, and model opacity." Additionally, the authors identify that "LLMs offer significant potential to revolutionize research in materials science and chemistry by enhancing productivity, improving data extraction, and accelerating scientific discovery. However, challenges related to hallucination, reproducibility, and bias must be addressed before LLMs can be reliably integrated into research workflows."
Research paradigm
Interpretive/Critical; mixed-methods qualitative synthesis across disciplinary contexts
Author conclusions
The authors conclude: "By carefully considering the ethical implications and ensuring proper oversight, the scientific community can responsibly harness the potential of LLMs to advance research." They further state that for healthcare specifically, "LLMs in medicine and healthcare hold considerable potential for advancing research, improving diagnostics, and supporting clinical decision-making. In order for the integration of LLMs to be effective, several critical challenges that need to be met: data quality, privacy, ethical concerns, refinement of models, increasing transparency, and rigorous validation processes that ensure robust results for the field." Most broadly, they emphasize that their analysis "provides a foundation for informed, responsible integration of LLMs into scientific research, with implications for both current and future practices across disciplines."
Risk of bias
Language bias: inclusion restricted to English-language publications; Database bias: search limited to Google Scholar; may underrepresent non-indexed sources; Publication bias: restriction to peer-reviewed journals and credible databases may exclude gray literature and critical perspectives; Selection bias: multi-stage screening process may exclude relevant papers during title/abstract review; Disciplinary bias: explicit exclusion of formal and applied sciences narrows scope; Potential confirmation bias in thematic coding despite stated dual-coding procedures; Author screening bias: Though mitigated by independent review and consensus resolution; Selection bias: Inclusion criteria requiring explicit discussion of LLM impacts may exclude neutral descriptions; Coder bias: Although mitigated through independent coding and consensus, interpretive synthesis remains subject to reviewer interpretation
Limitations
- The authors state: "To maintain a focused and rigorous analysis, this review concentrates on the humanities, social sciences, and natural sciences, where text-based methodologies are foundational to research practices
- The exclusion of formal sciences (e.g., mathematics, computer science) and applied sciences (e.g., engineering, agriculture) is intentional, as these fields often emphasize algorithmic development, technical implementation, or creative applications that diverge from the textual and interpretive challenges examined here." Additionally, they acknowledge that "existing studies predominantly focus on specific applications, such as their role in accelerating discoveries in biology and chemistry, or their utility in assisting with academic writing and publication processes
- However, these works do not address the broader challenges or cross-disciplinary implications of LLM integration."
Open questions raised
- lack of comprehensive cross-disciplinary review of LLM integration in scientific research
- insufficient investigation of ethical dilemmas and systemic risks
- need for interdisciplinary governance frameworks
- requirement for robust validation standards and expanded explainability research
- underexplored challenges including erosion of researcher autonomy, AI-driven confirmation bias, authorship ambiguity, and unequal access to technologies
- absence of standard benchmarks for model interpretability and performance in materials science
Explore related topics
Related papers
- Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statementDavid Moher · 2009 · 83,271 citations
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviewsMatthew J. Page · 2021 · 10,956 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations