12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

LLMs as Tools for Low-Priority Scientific Intelligence

Nemi Pelgrom · Nordic Machine Intelligence · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.5617/nmi.13230

Methodology & findings

Study design

Empirical annotation study using automated classification by two LLMs (GPT-4o and Ovis2.5-9B).

Sample

N = 6000, 9 groups

Primary method

Manual verification sampling (20 abstracts per field), descriptive frequency analysis, percentage comparisons across models and fields. Analysis conducted in Excel. No formal inferential statistics reported.

Main result

The study found that "Out of the 120 abstracts, with 360 annotations, only 6 faulty annotations were identified," demonstrating 98% accuracy. Additionally, "The large majority, 92-95% (GPT-4o & Ovis2.5-9B), of all abstracts from the field of physics mention which method is used in the paper, along with 83-88% of them stating the results of the paper," revealing systematic differences in abstract writing conventions across academic fields.

Reports effect sizes.

Research paradigm

Empiricist/pragmatist

Author conclusions

The authors conclude that "The findings in this and other similar studies, support the possibility of using LLMs models to extract metadata from research level texts. This is valuable as a support for meta-research, and further to both support and challenge policy makers on what standards exist and can be expected to be present in academic texts." They further state "Human annotation cannot be replaced with LLMs annotation, but it can be supplemented with it. When there are datasets that are too large to invest in human annotation of them, LLMs trained on scientific data may fill in the gap between our capabilities and insight desires."

Risk of bias

Selection bias: Abstracts sourced from OpenAlex, which may not represent all academic sources equally; Model bias: LLMs may have inherent biases based on training data and prompt formulation; Sycophancy: Authors acknowledge this concern during prompt design; False negatives: All 6 inaccurate annotations consisted of false negatives rather than false positives; Varying model interpretations: Different definitions of technical terms between GPT-4o and Ovis2.5-9B; Selection bias: Abstracts sourced from OpenAlex only, which while diverse may not represent all academic sources; Model hallucination: GPT-4o extracted technical terms not present in abstracts (false positives); Prompt sensitivity: Models susceptible to prompt formulation, temperature settings, and training-data bias; False negatives bias: All 6 inaccurate annotations were false negatives (annotating N where Y accurate); Limited ground truth verification: Manual verification only on GPT-4o responses, not replicated for Ovis2.5-9B; Sycophancy bias: Paper mentions this was considered in prompt design but limitations remain; Selection bias: Abstracts gathered from OpenAlex only, which may not represent all academic sources despite stated diversity goals; Variance in abstract collections across fields and within fields on multiple unmeasured variables; Model hallucination: GPT-4o extracted technical terms not present in source abstracts; Verification sample not replicated across both models—only GPT-4o verified; No inter-rater reliability testing between human annotators; Potential sycophancy bias in model responses despite prompt design attempts to mitigate

Limitations

  • "Only a representative amount of abstracts were used as basis for this study, which adds another dimension of uncertainty on top of the uncertainty of the correctness of the models used
  • the uncertainty of whether the abstracts are a representative collection or not." Additionally, "However, since there is no ground truth to compare the results with, we cannot know whether the models are similarly right, or if they are both making the same kind of mistakes in their classification." The authors also note that "A manual check of a small sample of the extracted terms indicate that the additional technical terms provided by GPT-4o were added by the model, rather than extracted from the text."

Open questions raised

  • Limited understanding of normative standards for and between academic fields
  • Need for systematic extraction of metadata from research texts across disciplines
  • Lack of large-scale meta-research data on abstract writing conventions
  • Opportunities to map research topics, track terminology evolution, and support automated indexing of literature within fields
  • Limited exploration of normative differences between academic fields regarding abstract writing standards
  • Need for more rigorous and dependable methods for meta-research data gathering
Data: OpenAlex API - source of 6,000 scientific abstracts (1,000 from each of 6 disciplines); 6000 scientific paper abstracts collected from OpenAlex API (1000 from each of 6 fields: Biology, Physics, Computer Science, Sociology, Philosophy, Economics); 6,000 scientific paper abstractsExtracted from: pdfAgreement 60%

Explore related topics

Related papers