12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

PaperMentor: A Human-Centered Multi-Agent Writing Tutor for AI Research Papers on Overleaf

Jiarui Liu, Terry Jingchen Zhang, Ryan Faulkner, X. Angelo Huang, Vilém Zouhar, Dominik Glandorf et al. · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Mixed-methods study combining artifact design with user evaluation.

Primary method

Design science research with human-centered design principles; participatory input from expert researchers; iterative system development with user evaluation

Main result

PaperMentor significantly outperforms a direct prompting baseline in both validity and actionability. "PaperMentor significantly outperforms the direct prompting baseline in both validity and actionability. In contrast, baseline comments achieve higher conciseness on average. Overall, incorporating the skill library enables PaperMentor to generate feedback that is more accurate and more actionable." Specifically, "PaperMentor improves validity by 6.5 percentage points and actionability by 4.1 percentage points." In the user study with 14 AI researchers, "90.6% of the generated comments were rated actionable and 67.5% were rated valid."

Research paradigm

Design science; human-centered AI systems

Author conclusions

The authors conclude: "By grounding specialized review agents in a curated skill library distilled from senior researchers' guidance, the system significantly improves the validity and actionability of comments over a direct prompting baseline. More broadly, our results suggest that AI writing support for research papers should move beyond generic rewriting toward structured, mentor-like feedback that helps authors revise their own work while preserving authorship and judgment." They also state that "PaperMentor introduces a human-centered, multi-agent writing assistant that delivers expert-guided, actionable feedback directly within the Overleaf drafting workflow."

Risk of bias

Small sample size (n=14 annotators, 80 papers) may limit generalizability across writing styles, venues, and disciplines; Selection bias: papers intentionally sampled from all ICLR 2026 submissions rather than accepted papers, but still limited to LaTeX-available papers; Annotator bias: while comments were blinded to source (PaperMentor vs. baseline), annotators may have implicit preferences for longer or more detailed feedback; Limited diversity in paper types and research backgrounds represented in sample; System performance depends on underlying LLM reliability, which may vary by task; Selection bias: Papers sampled from ICLR submissions may not represent broader AI research writing practices; Annotator bias: Only 14 AI researchers with academic backgrounds (undergrad to PhD); limited demographic diversity; Social desirability bias: Annotators knew they were evaluating AI-generated feedback; Skill library bias: Authors acknowledge that "the skill library, though distilled from expert guidance, reflects the norms and conventions of predominantly English-language, Western AI venues, and may not generalize equitably to researchers writing from different cultural or disciplinary backgrounds"; LLM baseline confound: Single baseline model (GPT-5.2); unclear if improvements generalize to other LLMs; Limited diversity in evaluation: 80 papers from ICLR 2026 and 10 internal student submissions may not represent full diversity of writing styles, venues, disciplines, and researcher backgrounds; Annotator bias: 14 AI researchers with academic backgrounds ranging from undergraduate to PhD may have similar perspectives; Skill library bias: Reflects norms and conventions of predominantly English-language, Western AI venues, may not generalize equitably to researchers from different cultural or disciplinary backgrounds; Selection bias in baseline comparison: Direct comparison uses same LLM without skill library, which is a limited baseline

Limitations

  • The authors state that "PaperMentor currently operates primarily over LaTeX source and may therefore miss issues that depend on rendered PDF output, visual figure quality, or numerical verification." Additionally, "Our evaluation includes 80 papers and 14 annotators, which is sufficient to demonstrate statistically significant improvements over the baseline, but does not capture the full diversity of writing styles, venues, disciplines, and researcher backgrounds." They also note "the system depends on both the coverage of the skill library and the reliability of the underlying LLM." A key technical limitation is that "section-specific agents operate on limited portions of the manuscript rather than the entire paper" which causes "some validity errors occur when agents identify terms, definitions, or experimental details as missing even though they are introduced elsewhere in the document."

Open questions raised

  • Need for direct comparison against comments written by experienced researchers (expert-authored Overleaf comments) rather than only ablation comparison
  • Tradeoff between specialization and global document awareness: section-specific agents miss definitions and experimental details introduced elsewhere in the document; need for lightweight mechanisms for document-wide grounding
  • Limited diversity in evaluation: paper sample and annotator backgrounds do not capture full spectrum of writing styles, venues, disciplines, and researcher backgrounds
  • Need to evaluate on rendered PDF output, visual figure quality, and numerical verification
  • System currently depends on skill library coverage and underlying LLM reliability; opportunity to extend skill library with community contributions across diverse subfields and cultural/disciplinary backgrounds
  • Direct comparison against human expert-authored feedback: "our evaluation focuses on an ablation study that isolates the contribution of the skill library by comparing PaperMentor against the same LLM without access to expert writing skills. While this design allows us to measure the effect of the skill library, it does not directly compare system-generated feedback against comments written by experienced researchers. Collecting and benchmarking against expert authored Overleaf comments would provide a stronger reference point."
Data: 80 papers used in evaluation: 10 internal student submissions and 70 from ICLR 2026 submissions with downloadable LaTeX source files from arXiv; 80 papers: 10 from internal student submissions and 70 randomly sampled from ICLR 2026 submissions with publicly available LaTeX sources on arXiv; 80 papers: 10 from internal student submissions and 70 randomly sampled from ICLR 2026 submissions with downloadable LaTeX source files from arXivCode: https://github.com/jiarui-liu/overleaf (main codebase, AGPL-3.0 license); https://github.com/overleaf/overleaf (Overleaf Community Edition, which PaperMentor is built upon); https://github.com/jiarui-liu/overleaf (main codebase under AGPL-3.0 license); Overleaf Community Edition: https://github.com/overleaf/overleaf; https://github.com/jiarui-liu/overleaf (code released under AGPL-3.0 license); https://github.com/overleaf/overleaf (Overleaf Community Edition)Extracted from: pdfAgreement 47%

Explore related topics

Related papers