GoodPoint: Learning Constructive Scientific Paper Feedback from Author Responses
Jimin Mun, Chani Jung, Xuhui Zhou, Hyunwoo Kim, Maarten Sap · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
The study employs a multi-faceted computational approach combining supervised fine-tuning (SFT) and direct preference optimization (DPO) on a curated dataset of 19,534 ICLR papers.
Main result
The study found that "GOODPOINT-trained Qwen3-8B achieves 83.7% improvement in predicted success rate compared to the base model" and "GOODPOINT-SFT improves F1 by 58.8% over the base model and even exceeds Gemini-3-flash and GPT-5.2 in precision." Additionally, "in an expert human study, GOODPOINT-DPO outperforms Qwen3-8B across validity, actionability, specificity, and helpfulness, meaningfully reducing the gap to Gemini-3-flash."
Research paradigm
Empirical/pragmatist - leveraging real-world author responses as training signals
Author conclusions
"By defining feedback quality along the axes of validity and actionability, we developed both a large-scale dataset and a training strategy that aligns LLM outputs with feedback that authors acknowledge and act upon. Experiment results on both automatic metrics and human evaluation show that our approach delivers high-quality, precise feedback, even enabling small models to rival much larger ones. Our findings highlight the importance of grounding LLMs in human-centric signals." The authors argue that "AI should instead empower researchers—particularly those who are traditionally disadvantaged (junior and/or non-native English speaking researchers)—rather than replace human scientific judgment."
Risk of bias
Time constraints and community norms shape author responses, affecting supervision signal quality; Rebuttal incentives may bias author acknowledgment of feedback validity; Focus on ICLR may not generalize to other venues with different feedback norms; Potential dataset contamination from papers with knowledge cutoff dates (mitigated through held-out 2026 test set); Subsampling of 5 feedback units per paper may introduce coverage bias; Author response bias: Replies shaped by time constraints, community norms, and rebuttal incentives; Domain specificity bias: Dataset limited to ICLR papers only; Temporal contamination risk: Mitigated by holding out ICLR 2026 as separate test set; Annotation bias: Inter-annotator agreement PABAK=0.747 (moderate) and 0.837 (strong) for validity and action parsing; Selection bias in human evaluation: Participants primarily PhD students (N=13) with research experience in NLP (13/15); Venue-specific bias: ICLR-only training may not generalize to other conferences or disciplines; Temporal data contamination: ICLR 2026 papers held out as test set to avoid baseline knowledge cutoff contamination; Selection bias: Feedback quality assessment relies on author agreement, which may not capture objectively correct feedback that authors disagree with; Subsampling bias: Random subsampling of 5 feedback units per paper may introduce sampling variance
Limitations
- The authors state that "our approach uses author responses as a supervision signal, but these are an imperfect proxy for feedback quality, as replies can be shaped by time constraints, community norms, and rebuttal incentives
- while we mitigate this with DPO and quality filtering, future work could explore expert audits, finer-grained actionability labels, or revision-based outcome signals
- Additionally, our focus on ICLR limits generalizability, as feedback norms and criteria may differ across venues and disciplines, warranting future validation across additional fields."
Open questions raised
- Expert audits of author responses to improve supervision signal quality
- Finer-grained actionability labels for feedback
- Revision-based outcome signals as alternatives to author response parsing
- Validation across additional venues and disciplines beyond ICLR
- Generalization of constructive feedback frameworks to different scientific communities
- Finer-grained actionability labels needed beyond current six-action classification
Explore related topics
Related papers
- A comprehensive AI policy education framework for university teaching and learningCecilia Ka Yuk Chan · 2023 · 1,160 citations
- ChatGPT for Education and Research: Opportunities, Threats, and StrategiesMd. Mostafizer Rahman · 2023 · 904 citations
- Exploring Students’ Perceptions of ChatGPT: Thematic Analysis and Follow-Up SurveyAbdulhadi Shoufan · 2023 · 464 citations
- AI-generated feedback on writing: insights into efficacy and ENL student preferenceJuan Escalante · 2023 · 461 citations
- Is it harmful or helpful? Examining the causes and consequences of generative AI usage among university studentsMuhammad Abbas · 2024 · 372 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations