12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

AstroReview: An LLM-driven Multi-Agent Framework for Telescope Proposal Peer Review and Refinement

Yutong Wang, Yunxiang Xiao, Yonglin Tian, Junyong Li, Jing Wang, Yisheng Lv · ArXiv.org · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Multi-stage framework design with empirical evaluation through controlled experiments.

Primary method

Design science methodology with iterative refinement; task decomposition into isolated stages based on human expert review workflows

Main result

The study found that "without any domain specific fine tuning, AstroReview used in our experiments only for the last stage, correctly identifies genuinely accepted proposals with an accuracy of 87%" and that "with its integrated Proposal Authoring Agent, the acceptance rate of revised drafts increases by 66% after two iterations, showing that iterative feedback combined with automated meta-review and reliability verification delivers measurable quality gains."

Research paradigm

Design science and applied computer science

Author conclusions

The authors conclude that "By uniting open-source large language models within a purpose-built, multi-agent architecture, our work demonstrates a viable path toward alleviating the twin bottlenecks of proposal preparation and review that presently limit the scientific return of time-dominated astronomical facilities." They further state that "Empirical tests show that proposals iteratively polished by the system score higher on clarity, technical soundness, and scientific merit, while the Review Agent reliably reproduces human acceptance decisions—offering an automated, expert-level second opinion. Taken together, these results indicate that domain-tuned LLM agents can both enhance the quality of investigator submissions and lighten reviewer workloads, paving the way for more efficient, equitable, and impactful transient-focused observing programs in the next generation of astrophysical research."

Risk of bias

Selection bias: Only HST accepted proposals available; no genuine rejected proposals in dataset; Synthetic negative samples may not reflect actual rejected proposals; Label leakage mitigated through de-identification, but metadata removal incomplete verification unclear; Model selection bias: Qwen-2.5-72B chosen after benchmarking three models; Stochastic variation in LLM outputs acknowledged but not fully quantified; Single reviewer per round in refinement experiments to limit computational cost may reduce diversity; Selection bias: Dataset limited to accepted HST proposal abstracts only; no genuine rejected proposals available; Synthetic negative sample bias: Artificially generated negatives may not reflect real rejection patterns; Label leakage risk: Initial inclusion of metadata (Prop. Type, Category, ID, Cycle, Title, PI) mitigated through de-identification; Model bias: Single model selection (Qwen-2.5-72B) may introduce model-specific biases; Stochastic variation: LLM sampling temperature effects addressed through reliability verification; Synthetic dataset bias: Negative samples generated by controlled perturbations of accepted proposals rather than using genuine rejected proposals, potentially not capturing real rejection characteristics; Class label leakage mitigated through de-identification, but inherent biases in original HST acceptance decisions remain in positive samples; Model selection bias: Qwen-2.5-72B chosen based on subjective qualitative assessment against two other models with limited objective comparison metrics; Ceiling effect observed in Round 3 (99% acceptance rate), limiting ability to assess performance on borderline cases; Single-model dependency: All experiments use one primary LLM (Qwen-2.5-72B); generalizability to other models limited

Limitations

  • The authors acknowledge that "Publicly available corpora of complete observing proposals are exceedingly scarce
  • Consequently, our experiments rely exclusively on the Proposal Abstracts Catalog, which contains only the abstracts of successful Hubble Space Telescope (HST) submissions." They further note that "An abstract, however, is merely a concise excerpt of the full proposal, and the catalog's entries vary widely in writing style, formatting, length, etc
  • This heterogeneity yields a fragmented linguistic signal that prevents the semantic- and parameter-parsing modules in Stages 1 and 2 from reliably extracting the structured information needed for novelty and feasibility assessment
  • Accordingly, the current implementation of the Review Agent bypasses these two stages and operates solely at Stage 3." Additionally, the negative samples are synthetically generated: "the 'negative' drafts we synthesize still retain features that appeal to reviewers, so...our strongest configuration...raises the accuracy on true positives yet drives accuracy on negatives down to only 46%, actually worse than random guessing (50%)
  • The so-called rejected cases are intentionally degraded versions of proposals that were once accepted
  • even after the prose is diluted, the drafts still present attractive targets and significant scientific value, making them hard for the agent to dismiss."

Open questions raised

  • Future work should: (1) Collaborate with observatories to obtain full-text proposals (both accepted and rejected) with corresponding reviewer comments for domain-specific fine-tuning; (2) Incorporate observation strategy simulation into the workflow for feasibility assessment; (3) Develop multimodal review mechanism combining textual and structural assessment with parameter-optimized simulation results for fully automated end-to-end proposal refinement.
  • Future work requires: (1) collaboration with observatories to obtain full-text proposals (both accepted and rejected) along with corresponding reviewer comments for domain-specific fine-tuning; (2) integration of observation strategy simulation into the workflow to enable feasibility and scientific value assessment beyond textual analysis; (3) incorporation of parameter-optimized simulations into the Review Agent's evaluation process for multimodal review mechanism.
  • Lack of publicly available full-text proposals (both accepted and rejected) with corresponding reviewer comments for domain-specific fine-tuning
  • Need for integration of observation strategy simulation into the workflow for feasibility and scientific value evaluation
  • Collaboration with observatories needed to obtain real proposal data and validate framework effectiveness
Data: HST Proposal Abstracts Catalog (13,411 accepted proposals used as base dataset); synthetically generated balanced dataset of 26,822 proposals (1:1 positive-negative ratio). Dataset construction methodology provided but no URL or public repository link specified.; HST Proposal Abstracts Catalog (13,411 accepted proposals) - publicly available via Hubble Space Telescope archives; Constructed balanced dataset: 26,822 proposals (13,411 positive samples + 13,411 synthetically generated negative samples); Proposal Abstracts Catalog (HST) - 13,411 accepted proposals publicly available from Hubble Space TelescopeCode: AstroReview described as 'open-source' but no specific GitHub URL provided in the paperExtracted from: pdfAgreement 65%

Explore related topics

Related papers