Paper Reconstruction Evaluation: Evaluating Presentation and Hallucination in AI-written Papers
Atsuyuki Miyai, Mashiro Toyooka, Zaiying Zhao, Kenta Watanabe, Toshihiko Yamasaki, Kiyoharu Aizawa · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Benchmark evaluation framework (PaperRecon) with dual-axis assessment: (1) rubric-based presentation evaluation using LLM judges scoring key elements on 1-5 scale across section categories (Abstract, Introduction, Method, Benchmark Construction, Experiment, Related Work, Conclusion); (2) agentic hallucination evaluation via two-stage claim extraction and verification against ground-truth papers.
Sample
N = 51, 12 groups
Primary method
Kendall's τb correlation analysis for human validation; precision calculation for hallucination detection (96% of 97 flagged claims verified as genuine); F1-based metrics for citation evaluation (precision, recall, F1 harmonic mean); Mann-Whitney U or similar non-parametric tests implied for model comparisons (not explicitly stated); LLM-based classification with binary/multi-class outcomes
Main result
The study found that "Claude Code achieves higher presentation quality than Codex" with an average rubric score of 3.86 vs. 3.59, but "Codex produces fewer hallucinations than Claude Code," with Claude Code exhibiting "more than 10 hallucinations per paper on average" while Codex limits this to "around 3." These results reveal "a clear trade-off: while both ClaudeCode and Codex improve with model advances, ClaudeCode achieves higher presentation quality at the cost of more than 10 hallucinations per paper on average, whereas Codex produces fewer hallucinations but lower presentation quality."
Reports effect sizes and confidence intervals.
Research paradigm
Empirical-positivist (quantitative benchmark evaluation)
Author conclusions
"This work takes a first step toward establishing evaluation frameworks for AI-driven paper writing and improving the understanding of its risks within the research community." The authors conclude that "Paper Reconstruction Evaluation, the first systematic evaluation framework for AI-generated scientific papers" along with "PaperWrite-Bench, a benchmark designed for PaperRecon," enables "comprehensive evaluation of the capabilities and risks of modern writing agents." Furthermore, they state that "the importance of evaluating both dimensions [presentation quality and hallucination] to accurately assess model performance" is critical, as their findings "reveal a clear trade-off between presentation quality and hallucination, highlighting the importance of evaluating both dimensions."
Risk of bias
Selection bias in benchmark curation: only 51 papers manually selected from top-tier venues; Potential LLM-as-judge bias in rubric evaluation; Limited human validation (only 3 reviewers, 72 paper pairs); Benchmark limited to papers published after 2025, may not reflect older research styles; Rubric generation and refinement by authors introduces potential subjective bias; False positive/negative detection in hallucination evaluation despite two-stage verification; Selection bias: Papers manually curated from top-tier venues; may not represent broader distribution of AI-written papers; Evaluator bias: LLM judges (GPT-5.4) used for rubric and hallucination scoring; LLM biases may propagate; Human validation limited to 3 reviewers; small sample for generalizing human preference alignment; Temporal bias: Benchmark consists of papers published after 2025; future generalization unknown; Model capability bias: Evaluation limited to specific agent architectures (Claude Code, Codex, Claude Agent Teams); results may not generalize to other agents; Prompt engineering bias: Writing pipeline and agent prompts explicitly designed; different prompt formulations could yield different results; Rubric generation and refinement by authors may introduce subjective bias; Human validation conducted by authors themselves for hallucination verification; Limited to papers published after 2025, potentially biased toward recent trends; Paper selection from top-tier venues may not represent broader publication landscape; LLM-based classification and scoring may have inherent biases
Limitations
- "Our framework provides structured resources, including figures, tables, and references, to the agent
- This design reduces external dependencies such as retrieval and reference collection, and allows us to focus on evaluating core writing ability
- Evaluating writing performance under more limited resources, including settings where models rely on external systems, is an important direction for future work." Additionally, "Evaluating scientific papers is inherently challenging, as human writing is diverse and not fully captured by current LLMs
- As a result, section-wise evaluation may not fully reflect overall quality." The overlap penalty "operates on root (pelvis) distances and does not model the full 3D body mesh
- in highly contorted poses, slight limb penetrations can still occur."
Open questions raised
- Evaluating writing performance under more limited resources, including settings where models rely on external systems for retrieval and reference collection
- Developing more robust methods for capturing diverse human writing styles
- Exploring alternative graph topologies (ring, tree structures) for pivot assignment in multi-person generation
- Combining inference-time optimization with lightweight fine-tuning for improved semantic fidelity
- Extending to interactions requiring simultaneous coupling among three or more persons
- Limited prior work on systematic evaluation of AI-written papers; existing approaches using AI reviewers tend to assign higher scores to papers with severe fabrications
Explore related topics
Related papers
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations
- ChatGPT: Bullshit spewer or the end of traditional assessments in higher education?Jürgen Rudolph · 2023 · 1,674 citations
- ChatGPT for Education and Research: Opportunities, Threats, and StrategiesMd. Mostafizer Rahman · 2023 · 904 citations
- ChatGPT and a new academic reality: Artificial Intelligence‐written research papers and the ethics of the large language models in scholarly publishingBrady Lund · 2023 · 769 citations
- ChatGPT in higher education: Considerations for academic integrity and student learningMiriam Sullivan · 2023 · 740 citations