12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Paper Reconstruction Evaluation: Evaluating Presentation and Hallucination in AI-written Papers

Atsuyuki Miyai, Mashiro Toyooka, Zaiying Zhao, Kenta Watanabe, Toshihiko Yamasaki, Kiyoharu Aizawa · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Benchmark evaluation framework (PaperRecon) with dual-axis assessment: (1) rubric-based presentation evaluation using LLM judges scoring key elements on 1-5 scale across section categories (Abstract, Introduction, Method, Benchmark Construction, Experiment, Related Work, Conclusion); (2) agentic hallucination evaluation via two-stage claim extraction and verification against ground-truth papers.

Sample

N = 51, 12 groups

Primary method

Kendall's τb correlation analysis for human validation; precision calculation for hallucination detection (96% of 97 flagged claims verified as genuine); F1-based metrics for citation evaluation (precision, recall, F1 harmonic mean); Mann-Whitney U or similar non-parametric tests implied for model comparisons (not explicitly stated); LLM-based classification with binary/multi-class outcomes

Main result

The study found that "Claude Code achieves higher presentation quality than Codex" with an average rubric score of 3.86 vs. 3.59, but "Codex produces fewer hallucinations than Claude Code," with Claude Code exhibiting "more than 10 hallucinations per paper on average" while Codex limits this to "around 3." These results reveal "a clear trade-off: while both ClaudeCode and Codex improve with model advances, ClaudeCode achieves higher presentation quality at the cost of more than 10 hallucinations per paper on average, whereas Codex produces fewer hallucinations but lower presentation quality."

Reports effect sizes and confidence intervals.

Research paradigm

Empirical-positivist (quantitative benchmark evaluation)

Author conclusions

"This work takes a first step toward establishing evaluation frameworks for AI-driven paper writing and improving the understanding of its risks within the research community." The authors conclude that "Paper Reconstruction Evaluation, the first systematic evaluation framework for AI-generated scientific papers" along with "PaperWrite-Bench, a benchmark designed for PaperRecon," enables "comprehensive evaluation of the capabilities and risks of modern writing agents." Furthermore, they state that "the importance of evaluating both dimensions [presentation quality and hallucination] to accurately assess model performance" is critical, as their findings "reveal a clear trade-off between presentation quality and hallucination, highlighting the importance of evaluating both dimensions."

Risk of bias

Selection bias in benchmark curation: only 51 papers manually selected from top-tier venues; Potential LLM-as-judge bias in rubric evaluation; Limited human validation (only 3 reviewers, 72 paper pairs); Benchmark limited to papers published after 2025, may not reflect older research styles; Rubric generation and refinement by authors introduces potential subjective bias; False positive/negative detection in hallucination evaluation despite two-stage verification; Selection bias: Papers manually curated from top-tier venues; may not represent broader distribution of AI-written papers; Evaluator bias: LLM judges (GPT-5.4) used for rubric and hallucination scoring; LLM biases may propagate; Human validation limited to 3 reviewers; small sample for generalizing human preference alignment; Temporal bias: Benchmark consists of papers published after 2025; future generalization unknown; Model capability bias: Evaluation limited to specific agent architectures (Claude Code, Codex, Claude Agent Teams); results may not generalize to other agents; Prompt engineering bias: Writing pipeline and agent prompts explicitly designed; different prompt formulations could yield different results; Rubric generation and refinement by authors may introduce subjective bias; Human validation conducted by authors themselves for hallucination verification; Limited to papers published after 2025, potentially biased toward recent trends; Paper selection from top-tier venues may not represent broader publication landscape; LLM-based classification and scoring may have inherent biases

Limitations

  • "Our framework provides structured resources, including figures, tables, and references, to the agent
  • This design reduces external dependencies such as retrieval and reference collection, and allows us to focus on evaluating core writing ability
  • Evaluating writing performance under more limited resources, including settings where models rely on external systems, is an important direction for future work." Additionally, "Evaluating scientific papers is inherently challenging, as human writing is diverse and not fully captured by current LLMs
  • As a result, section-wise evaluation may not fully reflect overall quality." The overlap penalty "operates on root (pelvis) distances and does not model the full 3D body mesh
  • in highly contorted poses, slight limb penetrations can still occur."

Open questions raised

  • Evaluating writing performance under more limited resources, including settings where models rely on external systems for retrieval and reference collection
  • Developing more robust methods for capturing diverse human writing styles
  • Exploring alternative graph topologies (ring, tree structures) for pivot assignment in multi-person generation
  • Combining inference-time optimization with lightweight fine-tuning for improved semantic fidelity
  • Extending to interactions requiring simultaneous coupling among three or more persons
  • Limited prior work on systematic evaluation of AI-written papers; existing approaches using AI reviewers tend to assign higher scores to papers with severe fabrications
Data: PaperWrite-Bench: 51 papers from top-tier venues (NeurIPS, ICLR, CVPR, ICCV, ACL, ACMMM) published after 2025; Associated resources mentioned: tables, figures, references (bib files), and code repositories where available; PaperWrite-Bench: 51 papers from top-tier venues (ACL 2025, EMNLP 2025, CVPR 2025, CVPR 2026, ICCV 2025, ICLR 2025, NeurIPS 2025, ICLR 2026, ACMMM 2025); InterHuman dataset: ~107M frames across 23,337 text-annotated two-person interactions (mentioned in PINO method section); EgoLife dataset: 266-hour week-long multimodal egocentric data with 6 participants (example in appendix); PaperWrite-Bench: 51 papers from ACL 2025, EMNLP 2025, CVPR 2025, CVPR 2026, ICCV 2025, ICLR 2025, NeurIPS 2025, ICLR 2026, ACMMM 2025Code: Not explicitly listed; paper mentions providing code when available from original papers but does not specify a dedicated repository for PaperRecon or PaperWrite-Bench; Not explicitly provided; paper states agents have access to associated codebase when available from arXiv sourcesExtracted from: pdfAgreement 46%

Explore related topics

Related papers