12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Friend or foe? Exploring the implications of large language models on the science system

Benedikt Fecher, Marcel Hebing, Melissa Laufer, Jörg Pohle, Fabian Sofsky · AI & Society · 2023

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
E
Evidence
66
Citations
2.34
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s00146-023-01791-1

Methodology & findings

Study design

Two-round Delphi study.

Sample

N = 72, 6 groups

Primary method

Frequency distributions reported for ranking questions with two scoring approaches: (1) simple sum of frequencies, (2) weighted rank score (factor 4 for rank 1, factor 2 for rank 2). Likert-scale responses (5-point) aggregated by combining 'agree' and 'strongly agree' categories for analysis. Inductive and deductive coding of open-ended responses, with inter-rater reliability established through dual coding of initial subset and consensus discussion by all authors.

Main result

Our findings indicate that experts anticipate that the utilization of LLMs will have a transformative and largely positive impact on science and scientific practice. "In LLMs, they recognize significant potential for administrative, creative, and analytical tasks. The main risks associated with LLMs pertain to issues of bias, misinformation, and overburdening of the scientific quality assurance system." The study found that text improvement is considered the most important application of LLMs, followed by text summary and code writing, with 59.6% of experts either already using or expressing intention to use LLMs in their work.

Reports effect sizes.

Research paradigm

Mixed methods (inductive and deductive coding with quantitative ranking)

Author conclusions

"Our study highlights the great potential generative AI has for transforming the science system." The authors emphasize that "striking a balance between embracing the benefits of LLMs and upholding scientific principles is crucial" and that "we must withstand any attempt to compromise the quality standards that we as a science community have established." They conclude: "while the transformative scenario holds great promise for the positive impact of LLMs on the science system and society, it is imperative to proactively address the potential risks and challenges to ensure that the integration of generative AI in science is guided by ethical considerations, scientific integrity, and a commitment to societal benefit."

Risk of bias

Selection bias: Convenience sampling through professional networks and institutional mailing lists; Technology proficiency bias: Authors acknowledge sampling likely involved 'technology-proficient experts'; Self-selection bias: Experts interested in AI/digitization may overrepresent positive views; Attrition: 72% response rate in second round (52 of 72 participants); Selection bias: Convenience sampling via professional networks may preferentially recruit technology-proficient experts and AI specialists; Response bias: Round 2 response rate of 72% (52/72) with potential non-response bias from those not responding to second survey; Sampling strategy bias: Authors acknowledge 'possibly involving technology-proficient experts' which may bias toward positive assessment of LLM benefits; Limited temporal scope: Short interval between data collection rounds prevents assessment of long-term implications; Convenience sampling via professional networks may introduce selection bias favoring technology-aware researchers; Technology-proficient expert sample may overstate benefits and minimize risks; 72% response rate in round 2 may introduce attrition bias; Short interval between survey rounds limits ability to detect long-term effects; Open questions in round 1 may have elicited self-selection in responses

Limitations

  • The authors acknowledge that "Our Delphi approach allowed us to identify and refine various implications of LLMs on the science system
  • however, it was not without its limitations
  • For example, we were unable to track long-term implications as the interval between the data collections were relatively short." Additionally, the sampling strategy "possibly involving technology-proficient experts" may have biased results toward more positive assessments of LLM benefits.

Open questions raised

  • The authors identify the gap that 'there is much less—especially empirical research—on the implications of LLMs as well as LLM-based chatbots or prompts on scholarly practices and the science system.' They suggest future research directions including: the need to reconsider conventional understandings of authorship in light of LLM impacts; exploration of how reputation systems and citation-based metrics may evolve; investigation of quality control system strain; and public discussions on ethical considerations and suitable proactive regulation approaches.
  • Limited empirical research on effects of LLMs and LLM-based chatbots on science and scientific practice (stated as motivation)
  • Need for longitudinal studies tracking long-term implications (authors acknowledge inability to assess due to short interval between data collections)
  • Unclear dynamics of how scholarly evaluation metrics and reputation systems will evolve with automated content generation
  • Insufficient clarity on how quality assurance systems will adapt to increased volume of LLM-assisted submissions
  • Need for further exploration of how to balance advancement of scientific norms around authorship and creativity with emerging AI capabilities
Data: https://doi.org/10.5281/zenodo.8009429 (Zenodo repository with survey instruments and response data); Full dataset available on Zenodo repository at https://doi.org/10.5281/zenodo.8009429; Data including survey instruments published under CC-BY-license, available via Zenodo repository at https://doi.org/10.5281/zenodo.8009429Code: Not mentionedExtracted from: pdfAgreement 56%

Explore related topics

Related papers