12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Extending AI for Research to the Humanities: A Multi-Agent Framework for Evidence-Grounded Scholarship

Yating Pan, Jiajun Zhang, Jun Wang, Qi Su · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Computational system design and evaluation study using: (1) multi-agent workflow orchestration over a typed evidence pool; (2) multi-scale close-reading infrastructure combining passages, entity-relation graph communities, and semantic clusters; (3) peer-reviewed-paper benchmark construction with manual extraction of research questions, findings, and cited primary-source evidence from 406 papers; (4) automatic evaluation via evidence recall at multiple granularities (work, section, sentence level); (5) blind scholarly evaluation by human experts and LLM judges on four dimensions (accuracy, depth, coverage, evidence quality); (6) ablation studies removing individual agents and retrieval tiers..

Sample

N = 406, 16 groups

Primary method

Automatic metrics: IR recall (macro-averaged across papers at multiple granularities), Mean Reciprocal Rank (MRR@k), normalized Discounted Cumulative Gain (nDCG@k), BGE-M3 cosine similarity, BERTScore, ROUGE-L. Human and LLM judging on ordinal 1–5 scale with inter-rater agreement measured via Gwet's AC2 (quadratic weights) and quadratic-weighted Cohen's κ. Ablation studies removing individual agents and retrieval tiers, with component contributions isolated. Threshold sensitivity analysis sweeping cosine similarity thresholds (τs, τc) across five settings.

Main result

On a peer-reviewed-paper benchmark over classical Chinese and Greco-Roman Latin scholarship, "SPIRE recovers cited primary-source evidence more reliably than Naive LLM, Text RAG, and GraphRAG, and receives higher blind-judge scores on answer accuracy, depth, coverage, and evidence quality." Specifically, "SPIRE attains 44.3% evidence recall versus ≤22.4% for the strongest baseline (roughly double)" and "leads the scholarly judge on every aspect, for both LLM raters and both human experts."

Reports effect sizes.

Research paradigm

Empirical-computational with interpretive scholarship foundations

Author conclusions

"We introduce SPIRE, a multi-agent framework for evidence-grounded scholarship in the classical humanities. It operationalizes Scholarly Primitives as the unit of computation through seven cooperating agents over a shared EvidencePool and a multi-scale substrate of passages, graph communities, and semantic clusters, making retrieval, contextualization, comparison, and writing inspectable. On a peer-reviewed-paper benchmark over classical Chinese and Latin scholarship, SPIRE achieves higher primary-source evidence retrieval than Naive LLM, Text RAG, and GraphRAG, and receives higher human and LLM judge scores. Its contribution is an architecture making agentic humanities research assistance operational and inspectable, together with a peer-reviewed-paper benchmark and evaluation protocol for further humanities AI research."

Risk of bias

Manual evidence extraction by two graduate students introduces human coding bias and potential inconsistency; Limited to two classical traditions (Chinese and Latin); generalizability to other humanities domains unclear; Benchmark drawn from peer-reviewed papers, which may not represent all scholarly evidence-use patterns; LLM judges may exhibit systematic biases aligned with LLM-generated text; Generator LLM (DeepSeek-V4-Flash) choice may favor certain response patterns; No blinding of developers to benchmark during system design; Threshold values for evidence matching (τs=0.80, τc=0.70) set a priori but not validated on independent data; Manual extraction by two graduate students introduces potential annotation bias; Selection of papers from specific databases (JSTOR, Academia, Google Scholar, CNKI) may not represent full corpus of humanities scholarship; LLM judge may exhibit biases inherent to foundation models; Cross-language evaluation (Chinese-Latin) may introduce language-specific biases in evaluation metrics; Manual benchmark construction by two graduate students (potential annotation bias/disagreement not reported); Evaluation using LLM judges which may inherit generator biases; Single corpus instantiation (classical texts only) limits generalizability; Retrieval threshold selection (τs=0.80, τc=0.70) fixed a priori but robustness tested across sweep

Limitations

  • "We instantiate and evaluate SPIRE on classical Chinese and Greco-Roman Latin corpora, two rich and demanding reservoirs of humanistic scholarship..
  • Future work will extend the same framework to a broader range of humanities domains, including manuscript studies, vernacular and modern literatures, religious studies, archival history, and scholarship in additional languages and historiographic traditions." Additionally, "This paper focuses on textual evidence—passages, works, chapters, quotations, and relations among texts..
  • As humanities research increasingly brings together texts with images, archaeological objects, performance records, maps, and archival materiality, an important next step is to extend SPIRE with multimodal retrieval, provenance tracking, and agent analysis over non-textual evidence."

Open questions raised

  • Lack of multi-agent research frameworks designed for evidence-grounded interpretive scholarship over primary sources (addressed by SPIRE)
  • Gap between RAG/GraphRAG optimization for retrieval vs. humanistic need for faithful citation, verifiable provenance, and close reading
  • Limited evaluation methods for humanities AI that prioritize evidence-groundedness and defensible interpretation over quantitative reproduction
  • Absence of benchmark for classical humanities scholarship with primary-source evidence annotation
  • Need to extend beyond classical Chinese and Latin to broader humanities domains (manuscript studies, vernacular literature, religious studies, archival history)
  • Limitation of current work to textual evidence; need for multimodal extensions to images, archaeological objects, maps, and archival materiality
Data: Classical Chinese corpus: Zhongguo Xueshu Mingzhu Tiyao (Zhou, 1992); 226 documents, 51,726 text chunks, 8c BCE–19c CE time span; Greco-Roman Latin corpus: The Latin Library (2026); 710 documents, 35,117 text chunks, 5c BCE–20c CE time span; Peer-reviewed-paper benchmark: 406 papers (286 English, 120 Chinese) from JSTOR, Academia.com, Google Scholar, and CNKI with manually extracted research questions, findings, and cited primary-source evidence; Classical Chinese corpus (Zhongguo Xueshu Mingzhu Tiyao): 226 documents, 51,726 text chunks; Greco-Roman Latin corpus (The Latin Library): 710 documents, 35,117 text chunks; Peer-reviewed-paper benchmark: 406 papers with extracted research questions, findings, and cited primary-source evidence; Classical Chinese corpus: Zhongguo Xueshu Mingzhu Tiyao (Zhou, 1992); Greco-Roman Latin corpus: The Latin Library (https://www.thelatinlibrary.com/); Peer-reviewed-paper benchmark: 406 papers from JSTOR, Academia, Google Scholar, and CNKI with manually extracted research questions and primary-source evidence; Code, data catalogues, and reproduction scripts: https://github.com/YatingPan/SPIRECode: https://github.com/YatingPan/SPIRE; https://github.com/YatingPan/SPIRE - Code, data catalogues, and reproduction scriptsExtracted from: pdfAgreement 56%

Explore related topics

Related papers