12,637 papers · updated 18 Sept 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Expert evaluation of LLM world models: A high-T c superconductivity case study

Haoyu Guo, Maria Tikhanovskaya, Paul Raccuglia, Alexey Vlaskin, Chris Co, Daniel J. Liebling et al. · Proceedings of the National Academy of Sciences · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1073/pnas.2533676123

Methodology & findings

Study design

Expert-curated evaluation study using a panel of 12 high-temperature superconductivity experts who created 67 questions and answers probing deep understanding of a curated database of 1,726 experimental papers.

Sample

N = 67, 5 groups

Primary method

Mann-Whitney U test (non-parametric comparison between systems); Mean and standard deviation (SD) calculation across questions and experts for each aspect and system; Three-point Likert scale (0, 1, 2) for ordinal outcome scoring; Distribution analysis of grades across models and aspects

Main result

The study found that "systems utilizing curated literature databases generally demonstrate superior efficacy compared to those sourcing information from unfiltered Internet data when addressing inquiries pertaining to advanced research on high-Tc cuprate superconductors." More specifically, "the NotebookLM system, which utilizes a curated literature database, surpasses closed LLM-based search engines that source unfiltered data from the Internet in terms of providing a balanced perspective, factual thoroughness, and supporting evidence."

Reports effect sizes.

Research paradigm

Empirical-pragmatic (evaluative benchmarking of AI systems using expert judgment)

Author conclusions

The authors conclude that "A major conclusion of this work is that grounding answers in the experimental literature improved their quality. When models were provided with context of the entire relevant literature, and asked to answer with support from these sources, the quality definitively improved in a blind test. This is reassuring as a conclusion, as it points the way toward more capable expert systems." They further state that "the results showed that current AI systems fall significantly short on this task. While for foundational or introductory purposes, LLM systems may serve as a useful springboard, they currently lack the ability to distinguish central theoretical frameworks from peripheral ideas."

Risk of bias

Expert panel composition bias: 12 experts with potentially overlapping perspectives; may not represent all viewpoints in the field; Temporal bias: Evaluations conducted December 2024–early 2025; rapid LLM development makes findings quickly outdated; Question design bias: Questions formulated by same expert panel conducting evaluation; potential for questions to align with panelists' perspectives; Evaluator fatigue/inconsistency: Multiple experts grading subsets of questions; potential for inter-rater reliability issues not explicitly reported; Literature curation bias: Initial selection based on 15 review articles recommended by experts; may exclude non-mainstream perspectives; Closed model access bias: ChatGPT-4o, Perplexity, Claude, Gemini trained on internet data; potential for systematic biases in their training corpora; Rubric subjectivity: Three-point scale (0, 1, 2) evaluation by human experts introduces inherent subjectivity; System selection bias: Only 6 systems tested; proprietary systems and newer models not included; may miss relevant papers outside standard references; Limited generalizability: Findings specific to high-temperature superconductivity domain

Limitations

  • The authors state: "One major limitation of this study is how difficult it is to put together this type of evaluation
  • One needs a finite field and a set of world experts that are able to pose questions and grade answers on topics that correspond to this expertise
  • Getting bandwidth from such a group is highly nontrivial
  • Grading responses across our rubric requires expert evaluation, and does not scale either to other fields or even to newer models." Additionally, "the evaluation used in this study is not up to date
  • The evaluations shown in this paper were carried out in early 2025, and so the LLM systems producing the answers were those available at late 2024."

Open questions raised

  • Limited visual reasoning: LLMs lack ability to meaningfully extract quantitative information from scientific data visualization
  • identified as "major direction of improvement for next-generation LLMs"
  • Temporal understanding: Systems fail to recognize relationships between conflicting or outdated claims across literature
  • Conceptual link identification: Models struggle with implicit conceptual connections
  • rely on surface-level textual similarity rather than deeper conceptual relevance
  • Multiturn interaction: Only initial responses analyzed
Data: "The LLM responses of system 5 and system 6 cannot be shared based on legal counsel since there may be possibilities that systems might reproduce content from papers that have copyright restrictions. All other data are included in the manuscript and/or supporting information." The curated literature database of 1,726 experimental papers is referenced but availability not explicitly stated for external use.Code: No code repositories explicitly mentionedExtracted from: pdf

Explore related topics

Related papers