From Text to Discovery: How Large Language Models Are Accelerating and Complicating Research Across Scientific and Humanistic Disciplines
Saleh Afroogh, Yasser Pouresmaeil, Yiming Xu, Kevin Chen, Abhejay Murali, Junfeng Jiao (7040318) · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Systematic scoping review following PRISMA guidelines.
Main result
The analysis reveals a consistent pattern: "LLMs meaningfully accelerate research workflows — from hypothesis generation and literature synthesis to data analysis and scientific writing — while introducing serious challenges related to hallucination, reproducibility, dataset bias, and model opacity." Additionally, the authors identify that "LLMs offer significant potential to revolutionize research in materials science and chemistry by enhancing productivity, improving data extraction, and accelerating scientific discovery. However, challenges related to hallucination, reproducibility, and bias must be addressed before LLMs can be reliably integrated into research workflows."
Research paradigm
Interpretive/Critical; mixed-methods qualitative synthesis across disciplinary contexts
Author conclusions
The authors conclude: "By carefully considering the ethical implications and ensuring proper oversight, the scientific community can responsibly harness the potential of LLMs to advance research." They further state that for healthcare specifically, "LLMs in medicine and healthcare hold considerable potential for advancing research, improving diagnostics, and supporting clinical decision-making. In order for the integration of LLMs to be effective, several critical challenges that need to be met: data quality, privacy, ethical concerns, refinement of models, increasing transparency, and rigorous validation processes that ensure robust results for the field." Most broadly, they emphasize that their analysis "provides a foundation for informed, responsible integration of LLMs into scientific research, with implications for both current and future practices across disciplines."
Risk of bias
Language bias: inclusion restricted to English-language publications; Database bias: search limited to Google Scholar; may underrepresent non-indexed sources; Publication bias: restriction to peer-reviewed journals and credible databases may exclude gray literature and critical perspectives; Selection bias: multi-stage screening process may exclude relevant papers during title/abstract review; Disciplinary bias: explicit exclusion of formal and applied sciences narrows scope; Potential confirmation bias in thematic coding despite stated dual-coding procedures; Selection bias: Inclusion limited to peer-reviewed journals and English-language publications only; Database bias: Search conducted only on Google Scholar rather than multiple databases (Scopus, PubMed, Web of Science); Language bias: Exclusion of non-English publications; Publication bias: Focus on published literature may exclude gray literature and negative results; Author screening bias: Though mitigated by independent review and consensus resolution; Language bias: Only English-language publications included; Publication bias: Reliance on peer-reviewed journals and credible databases may exclude negative findings; Selection bias: Inclusion criteria requiring explicit discussion of LLM impacts may exclude neutral descriptions; Coder bias: Although mitigated through independent coding and consensus, interpretive synthesis remains subject to reviewer interpretation; Disciplinary scope bias: Exclusion of formal and applied sciences narrows the evidence base
Limitations
- The authors state: "To maintain a focused and rigorous analysis, this review concentrates on the humanities, social sciences, and natural sciences, where text-based methodologies are foundational to research practices
- The exclusion of formal sciences (e.g., mathematics, computer science) and applied sciences (e.g., engineering, agriculture) is intentional, as these fields often emphasize algorithmic development, technical implementation, or creative applications that diverge from the textual and interpretive challenges examined here." Additionally, they acknowledge that "existing studies predominantly focus on specific applications, such as their role in accelerating discoveries in biology and chemistry, or their utility in assisting with academic writing and publication processes
- However, these works do not address the broader challenges or cross-disciplinary implications of LLM integration."
Open questions raised
- The authors identify: (1) lack of comprehensive cross-disciplinary review of LLM integration in scientific research; (2) insufficient investigation of ethical dilemmas and systemic risks; (3) need for interdisciplinary governance frameworks; (4) requirement for robust validation standards and expanded explainability research; (5) underexplored challenges including erosion of researcher autonomy, AI-driven confirmation bias, authorship ambiguity, and unequal access to technologies; (6) absence of standard benchmarks for model interpretability and performance in materials science; (7) need for improved citation integrity mechanisms; (8) limited understanding of privacy preservation in healthcare contexts; (9) requirement for standardized evaluation methods in psychology; (10) insufficient research on responsible integration protocols across disciplines.
- Lack of comprehensive cross-disciplinary review of LLM integration across humanities, social sciences, and natural sciences
- Need for governance frameworks addressing AI-driven confirmation bias, authorship ambiguity, and researcher autonomy erosion
- Insufficient standardized benchmarks for LLM development and application in specialized domains like materials science
- Limited research on unequal access to advanced LLM technologies between high-income and low-income research settings
- Need for expanded explainability research and formal verification tools for LLM-assisted scientific discovery
Explore related topics
Related papers
- Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statementDavid Moher · 2009 · 83,271 citations
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviewsMatthew J. Page · 2021 · 10,956 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations