Language agents achieve superhuman synthesis of scientific knowledge
Michael Skarlinski, Sam Cox, Jon M. Laurent, James D. Braza, Michaela M. Hinks, Michael J. Hammerling et al. · arXiv (Cornell University) · 2024
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.48550/arxiv.2409.13740
Methodology & findings
Study design
Rigorous human-AI comparison methodology involving three tasks: (1) multiple-choice question answering on LitQA2 benchmark with 248 questions requiring retrieval from scientific literature main text; (2) summarization task producing cited Wikipedia-style articles (WikiCrow) on human protein-coding genes with blind expert evaluation; (3) contradiction detection task (ContraCrow) extracting claims and checking for contradictions.
Main result
PaperQA2 achieved superhuman precision on literature retrieval tasks, with the study showing that "PaperQA2 thus achieved superhuman precision on this task (t(8.6) = 3.49, p = 0.0036) and did not differ significantly from humans in accuracy (t(8.5) = −0.42, p = 0.66)." Additionally, "WikiCrow had significantly fewer 'cited and unsupported' statements than the paired Wikipedia articles (13.5% vs. 24.9%) (p = 0.0075, χ2(1), N = 375)" and "PaperQA2 identifies 2.34 ± 1.99 (mean ± SD, N = 93 papers) contradictions per paper in a random subset of biology papers, of which 70% are validated by human experts."
Research paradigm
Empirical-computational with comparative human-AI methodology
Author conclusions
"We developed a methodology to compare or validate AI systems against human performance in realistic tasks for scientific research. PaperQA2 outperforms human experts on answering questions across all scientific literature; produces summaries that are, on average, more factual than Wikipedia summaries; and can be deployed to identify contradictions in scientific literature at scale." The authors emphasize that "Although PaperQA2 is expensive compared to lower accuracy commercial systems, it is inexpensive in absolute terms, costing $1 to $3 per query. Scaling up PaperQA2 and other literature-enabled agents like WikiCrow and ContraCrow empowers us to take advantage of the latent knowledge in literature at much greater scale than is possible today."
Risk of bias
Selection bias in human evaluators: Only 9 PhD students/postdocs evaluated LitQA2, small sample size with potential self-selection bias; Financial incentive bias: Humans received $3-12 per question plus performance bonuses, potentially influencing effort differently than AI systems; Question design bias: Questions generated by authors and contractors may contain implicit biases favorable to retrieval-based systems; Evaluator bias in WikiCrow validation: Only 4 evaluators graded WikiCrow vs Wikipedia statements; potential correlation with model outputs despite blinding; ContraCrow validation bias: Concern acknowledged by authors that contradiction validation annotators may be influenced by model reasoning and sources; Data contamination risk: Initial 147 LitQA2 questions were indexed by Google, requiring exclusion of third round human evaluations on those questions; Commercial system evaluation limitations: Perplexity and Elicit evaluated by manual author entry, not automated; Elicit only evaluated on 147/248 questions due to algorithm change; Human evaluators on LitQA2 may have been biased by financial incentives ($3-12 per question); Potential selection bias in WikiCrow evaluation: only non-stub Wikipedia articles were included (n=3,639 out of 19,255 genes); Human annotators for contradiction validation may have been influenced by ContraCrow's reasoning and chosen sources; Limited human sample size for some comparisons (n=9 PhD-level annotators); ContraCrow shows 'overconfidence' bias, disagreeing with human experts on 40% of claims; Selection bias in LitQA2 question generation (manually created by experts and contractors with calibration); Potential annotator bias toward model outputs in contradiction validation task (mitigated by separate contradiction detection task); Human evaluators may have used AI tools despite explicit instruction not to; Data leakage risk: initial 147 LitQA2 questions were indexed by Google, affecting validity of third evaluation round; Potential funding bias (FutureHouse Inc. funded work on PaperQA2); Limited human sample size (n=9 for LitQA2, n=4-5 for WikiCrow/ContraCrow evaluations)
Limitations
- The paper acknowledges several limitations: "Science requires extreme attention to detail, and LLMs can overlook or misuse details when faced with challenging reasoning problems." Additionally, regarding contradiction detection validation, "humans are significantly more correlated with each other than they are with ContraCrow (p = 0.015)," and the authors note "overconfidence on ContraCrow's part is the primary driver of its lack of agreement with human annotators." The paper also states that benchmarks "are not suitable as performance proxies for real scientific research tasks." Furthermore, human evaluators were asked but enforcement was not possible: "They were also allowed to use tools such as internet search or journal collection search provided via their institutions
- They were asked to explicitly refrain from using AI-based tools such as ChatGPT or Claude, though we did not have any method of enforcement of this request."
Open questions raised
- The paper identifies that existing benchmarks for retrieval and reasoning across scientific literature are underdeveloped, restricted to abstracts, limited to fixed corpora, or provide relevant papers directly without considering entire literature and lack direct human performance comparison. The authors also identify the need for improved contradiction detection systems, noting that overconfidence in model predictions should be addressed in future work.
- The authors identify that existing benchmarks for retrieval and reasoning across scientific literature are underdeveloped, restricted to abstracts, fixed corpora, or provide relevant papers directly, and lack direct human-performance comparisons. They note that ContraCrow overconfidence indicates 'directions for future improvement' and suggest that more research is needed on automated question generation (citing emerging ideas about automation).
- Authors identify the need for improved model confidence calibration in contradiction detection (noting overconfidence as primary limitation), improved distinction between contradictions and lack of support, and development of more comprehensive benchmarks for scientific literature tasks that compare to human performance.
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations
- Artificial intelligence and the conduct of literature reviewsGerit Wagner · 2021 · 275 citations