12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

GoodPoint: Learning Constructive Scientific Paper Feedback from Author Responses

Jimin Mun, Chani Jung, Xuhui Zhou, Hyunwoo Kim, Maarten Sap · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
C
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

The study employs a multi-faceted computational approach combining supervised fine-tuning (SFT) and direct preference optimization (DPO) on a curated dataset of 19,534 ICLR papers.

Main result

The study found that "GOODPOINT-trained Qwen3-8B achieves 83.7% improvement in predicted success rate compared to the base model" and "GOODPOINT-SFT improves F1 by 58.8% over the base model and even exceeds Gemini-3-flash and GPT-5.2 in precision." Additionally, "in an expert human study, GOODPOINT-DPO outperforms Qwen3-8B across validity, actionability, specificity, and helpfulness, meaningfully reducing the gap to Gemini-3-flash."

Research paradigm

Empirical/pragmatist - leveraging real-world author responses as training signals

Author conclusions

"By defining feedback quality along the axes of validity and actionability, we developed both a large-scale dataset and a training strategy that aligns LLM outputs with feedback that authors acknowledge and act upon. Experiment results on both automatic metrics and human evaluation show that our approach delivers high-quality, precise feedback, even enabling small models to rival much larger ones. Our findings highlight the importance of grounding LLMs in human-centric signals." The authors argue that "AI should instead empower researchers—particularly those who are traditionally disadvantaged (junior and/or non-native English speaking researchers)—rather than replace human scientific judgment."

Risk of bias

Time constraints and community norms shape author responses, affecting supervision signal quality; Rebuttal incentives may bias author acknowledgment of feedback validity; Focus on ICLR may not generalize to other venues with different feedback norms; Potential dataset contamination from papers with knowledge cutoff dates (mitigated through held-out 2026 test set); Subsampling of 5 feedback units per paper may introduce coverage bias; Author response bias: Replies shaped by time constraints, community norms, and rebuttal incentives; Domain specificity bias: Dataset limited to ICLR papers only; Temporal contamination risk: Mitigated by holding out ICLR 2026 as separate test set; Annotation bias: Inter-annotator agreement PABAK=0.747 (moderate) and 0.837 (strong) for validity and action parsing; Selection bias in human evaluation: Participants primarily PhD students (N=13) with research experience in NLP (13/15); Venue-specific bias: ICLR-only training may not generalize to other conferences or disciplines; Temporal data contamination: ICLR 2026 papers held out as test set to avoid baseline knowledge cutoff contamination; Selection bias: Feedback quality assessment relies on author agreement, which may not capture objectively correct feedback that authors disagree with; Subsampling bias: Random subsampling of 5 feedback units per paper may introduce sampling variance

Limitations

  • The authors state that "our approach uses author responses as a supervision signal, but these are an imperfect proxy for feedback quality, as replies can be shaped by time constraints, community norms, and rebuttal incentives
  • while we mitigate this with DPO and quality filtering, future work could explore expert audits, finer-grained actionability labels, or revision-based outcome signals
  • Additionally, our focus on ICLR limits generalizability, as feedback norms and criteria may differ across venues and disciplines, warranting future validation across additional fields."

Open questions raised

  • Expert audits of author responses to improve supervision signal quality
  • Finer-grained actionability labels for feedback
  • Revision-based outcome signals as alternatives to author response parsing
  • Validation across additional venues and disciplines beyond ICLR
  • Generalization of constructive feedback frameworks to different scientific communities
  • Finer-grained actionability labels needed beyond current six-action classification
Data: GOODPOINT-ICLR dataset with 19,534 ICLR papers (2020-2026) will be released upon acceptance (as stated: "We will release code, dataset, and trained models upon acceptance."); GOODPOINT-ICLR: 19,534 ICLR papers (2020–2026) with reviewer feedback annotated for validity and author action; sources include Re2 (2020–2023), arXiv (2024–2025), and OpenReview (2026); to be released upon acceptance; GOODPOINT-ICLR dataset (19,534 ICLR papers from 2020-2026 with reviewer feedback annotated for validity and actionability). Authors state: "We will release code, dataset, and trained models upon acceptance."Code: Code and trained models will be released upon acceptance according to the paper. Currently not publicly available.; Code and trained models to be released upon acceptance (stated in footnote 1); Code, dataset, and trained models to be released upon paper acceptance (not yet available at submission time)Extracted from: pdfAgreement 48%

Explore related topics

Related papers