Expert evaluation of LLM world models: A high-T c superconductivity case study
Haoyu Guo, Maria Tikhanovskaya, Paul Raccuglia, Alexey Vlaskin, Chris Co, Daniel J. Liebling et al. · Proceedings of the National Academy of Sciences · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1073/pnas.2533676123
Methodology & findings
Study design
Expert-curated evaluation study using a panel of 12 high-temperature superconductivity experts who created 67 questions and answers probing deep understanding of a curated database of 1,726 experimental papers.
Sample
N = 67, 10 groups
Primary method
Mann-Whitney U test (non-parametric comparison between systems); Mean and standard deviation (SD) calculation across questions and experts for each aspect and system; Three-point Likert scale (0, 1, 2) for ordinal outcome scoring; Distribution analysis of grades across models and aspects
Main result
The study found that "systems utilizing curated literature databases generally demonstrate superior efficacy compared to those sourcing information from unfiltered Internet data when addressing inquiries pertaining to advanced research on high-Tc cuprate superconductors." More specifically, "the NotebookLM system, which utilizes a curated literature database, surpasses closed LLM-based search engines that source unfiltered data from the Internet in terms of providing a balanced perspective, factual thoroughness, and supporting evidence."
Reports effect sizes.
Research paradigm
Empirical-pragmatic (evaluative benchmarking of AI systems using expert judgment)
Author conclusions
The authors conclude that "A major conclusion of this work is that grounding answers in the experimental literature improved their quality. When models were provided with context of the entire relevant literature, and asked to answer with support from these sources, the quality definitively improved in a blind test. This is reassuring as a conclusion, as it points the way toward more capable expert systems." They further state that "the results showed that current AI systems fall significantly short on this task. While for foundational or introductory purposes, LLM systems may serve as a useful springboard, they currently lack the ability to distinguish central theoretical frameworks from peripheral ideas."
Risk of bias
Expert panel composition bias: 12 experts with potentially overlapping perspectives; may not represent all viewpoints in the field; Temporal bias: Evaluations conducted December 2024–early 2025; rapid LLM development makes findings quickly outdated; Question design bias: Questions formulated by same expert panel conducting evaluation; potential for questions to align with panelists' perspectives; Evaluator fatigue/inconsistency: Multiple experts grading subsets of questions; potential for inter-rater reliability issues not explicitly reported; Literature curation bias: Initial selection based on 15 review articles recommended by experts; may exclude non-mainstream perspectives; Closed model access bias: ChatGPT-4o, Perplexity, Claude, Gemini trained on internet data; potential for systematic biases in their training corpora; Expert selection bias: Only 12 experts in high-temperature superconductivity, potentially not representative of all perspectives in the field; Temporal bias: Evaluation conducted in early 2025 using late 2024 LLM versions, limiting generalizability; Question formulation bias: Questions designed by the same expert panel that evaluated responses, though experts were blinded to system identity; Rubric subjectivity: Three-point scale (0, 1, 2) evaluation by human experts introduces inherent subjectivity; System selection bias: Only 6 systems tested; proprietary systems and newer models not included; Literature bias: 1,726 papers curated based on review articles and expert recommendations; may miss relevant papers outside standard references; Expert panel selection bias: Only 12 experts selected, may not represent all perspectives in field; Question formulation bias: Questions formulated by same experts who then evaluated answers; Evaluation timing bias: Evaluations conducted in early 2025 but responses generated from late 2024 models, creating knowledge gap; Limited generalizability: Findings specific to high-temperature superconductivity domain
Limitations
- The authors state: "One major limitation of this study is how difficult it is to put together this type of evaluation
- One needs a finite field and a set of world experts that are able to pose questions and grade answers on topics that correspond to this expertise
- Getting bandwidth from such a group is highly nontrivial
- Grading responses across our rubric requires expert evaluation, and does not scale either to other fields or even to newer models." Additionally, "the evaluation used in this study is not up to date
- The evaluations shown in this paper were carried out in early 2025, and so the LLM systems producing the answers were those available at late 2024."
Open questions raised
- Limited visual reasoning: LLMs lack ability to meaningfully extract quantitative information from scientific data visualization; identified as "major direction of improvement for next-generation LLMs"
- Temporal understanding: Systems fail to recognize relationships between conflicting or outdated claims across literature
- Conceptual link identification: Models struggle with implicit conceptual connections; rely on surface-level textual similarity rather than deeper conceptual relevance
- Multiturn interaction: Only initial responses analyzed; several experts reported improved quality in follow-up exchanges, suggesting "iterative dialogue may help LLMs refine their reasoning and outputs"
- Domain adaptation: Future directions include exploring few-shot prompting, LoRA-based fine-tuning on smaller domain-specific datasets, and agentic workflows
- Generalization beyond superconductivity: Study limited to single domain; unclear how findings generalize to other specialized scientific fields
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- To use or not to use ChatGPT in higher education? A study of students’ acceptance and use of technologyArtur Strzelecki · 2023 · 691 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performanceYizhou Fan · 2024 · 419 citations
- Impact of AI assistance on student agencyAli Darvishi · 2023 · 384 citations