12,637 papers · updated 18 Sept 2026livingmeta.ai
← Browse all papers
AI evidence extraction

PaperMentor: A Human-Centered Multi-Agent Writing Tutor for AI Research Papers on Overleaf

Jiarui Liu, Terry Jingchen Zhang, Ryan Faulkner, X. Angelo Huang, Vilém Zouhar, Dominik Glandorf et al. · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
D
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Mixed-methods study combining artifact design with user evaluation.

Primary method

Design science research with human-centered design principles; participatory input from expert researchers; iterative system development with user evaluation

Main result

PaperMentor significantly outperforms a direct prompting baseline in both validity and actionability. "PaperMentor significantly outperforms the direct prompting baseline in both validity and actionability. In contrast, baseline comments achieve higher conciseness on average. Overall, incorporating the skill library enables PaperMentor to generate feedback that is more accurate and more actionable." Specifically, "PaperMentor improves validity by 6.5 percentage points and actionability by 4.1 percentage points." In the user study with 14 AI researchers, "90.6% of the generated comments were rated actionable and 67.5% were rated valid."

Research paradigm

Design science; human-centered AI systems

Author conclusions

The authors conclude: "By grounding specialized review agents in a curated skill library distilled from senior researchers' guidance, the system significantly improves the validity and actionability of comments over a direct prompting baseline. More broadly, our results suggest that AI writing support for research papers should move beyond generic rewriting toward structured, mentor-like feedback that helps authors revise their own work while preserving authorship and judgment." They also state that "PaperMentor introduces a human-centered, multi-agent writing assistant that delivers expert-guided, actionable feedback directly within the Overleaf drafting workflow."

Risk of bias

Small sample size (n=14 annotators, 80 papers) may limit generalizability across writing styles, venues, and disciplines; Selection bias: papers intentionally sampled from all ICLR 2026 submissions rather than accepted papers, but still limited to LaTeX-available papers; Annotator bias: while comments were blinded to source (PaperMentor vs. baseline), annotators may have implicit preferences for longer or more detailed feedback; Limited diversity in paper types and research backgrounds represented in sample; System performance depends on underlying LLM reliability, which may vary by task; Selection bias: Papers sampled from ICLR submissions may not represent broader AI research writing practices; Annotator bias: Only 14 AI researchers with academic backgrounds (undergrad to PhD); Social desirability bias: Annotators knew they were evaluating AI-generated feedback; Skill library bias: Authors acknowledge that "the skill library, though distilled from expert guidance, reflects the norms and conventions of predominantly English-language, Western AI venues, and may not generalize equitably to researchers writing from different cultural or disciplinary backgrounds"; LLM baseline confound: Single baseline model (GPT-5.2); unclear if improvements generalize to other LLMs; Selection bias in baseline comparison: Direct comparison uses same LLM without skill library, which is a limited baseline

Limitations

  • The authors state that "PaperMentor currently operates primarily over LaTeX source and may therefore miss issues that depend on rendered PDF output, visual figure quality, or numerical verification." Additionally, "Our evaluation includes 80 papers and 14 annotators, which is sufficient to demonstrate statistically significant improvements over the baseline, but does not capture the full diversity of writing styles, venues, disciplines, and researcher backgrounds." They also note "the system depends on both the coverage of the skill library and the reliability of the underlying LLM." A key technical limitation is that "section-specific agents operate on limited portions of the manuscript rather than the entire paper" which causes "some validity errors occur when agents identify terms, definitions, or experimental details as missing even though they are introduced elsewhere in the document."

Open questions raised

  • Need for direct comparison against comments written by experienced researchers (expert-authored Overleaf comments) rather than only ablation comparison
  • Tradeoff between specialization and global document awareness: section-specific agents miss definitions and experimental details introduced elsewhere in the document
  • need for lightweight mechanisms for document-wide grounding
  • Limited diversity in evaluation: paper sample and annotator backgrounds do not capture full spectrum of writing styles, venues, disciplines, and researcher backgrounds
  • Need to evaluate on rendered PDF output, visual figure quality, and numerical verification
  • System currently depends on skill library coverage and underlying LLM reliability
Data: 80 papers used in evaluation: 10 internal student submissions and 70 from ICLR 2026 submissions with downloadable LaTeX source files from arXivCode: https://github.com/jiarui-liu/overleaf (main codebase, AGPL-3.0 license); https://github.com/overleaf/overleaf (Overleaf Community Edition, which PaperMentor is built upon)Extracted from: pdf

Explore related topics

Related papers