12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Paper Circle: An Open-source Multi-agent Research Discovery and Analysis Framework

Komal Kumar, Aman Chadha, Salman Khan, Fahad Shahbaz Khan, Hisham Cholakkal · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Design science approach with artifact development, evaluation through multi-query benchmarking (50-500 queries), ablation studies comparing retrieval baselines and pipeline configurations, and qualitative assessment through 81 real-world discovery sessions conducted by researchers.

Primary method

Design science with iterative development of multi-agent architecture based on smolagents library, incorporating user feedback through real-world evaluation sessions

Main result

The study found that "Two agent models achieve the highest retrieval effectiveness with an 80% hit rate, qwen3-coder-30b-Q3KM (quantized) and qwen3-coder:30b—with qwen3-coder-30b-Q3KM also delivering the best ranking quality (MRR = 0.627)" and demonstrated that "Paper Circle provides built-in evaluation metrics" with "Recall@K, Precision@K, and hit rates" computed across multiple configurations. Additionally, "Preliminary user feedback indicates minimal cognitive load when using PaperCircle. NASA-LLX assessment yields an overall workload of 1.2/7, with five of six dimensions scoring the minimum (1/7) and effort at 2/7."

Research paradigm

Design Science Research (DSR) / engineering paradigm

Author conclusions

"Paper Circle shows how multi-agent workflows can streamline research literature management. Its discovery pipeline unifies heterogeneous search sources and multi-criteria scoring into a reproducible tool, using a simple agent–tool interface with shared state, deterministic ranking, and synchronized multi-format outputs. Its analysis pipeline converts papers into structured knowledge graphs that enable graph-aware QA, coverage checks, and human-in-the-loop verification."

Risk of bias

Limited evaluation dataset (50 ICLR 2024 papers for review analysis, 50 queries for semantic benchmark); Potential selection bias in user study participants (81 sessions, 78 unique queries from researchers across 9 domains); Model-dependent performance variations (substantial differences between agent models tested); Recency bias in scoring framework (more recent papers receive higher scores); Limited ground truth validation for knowledge graph extraction quality; Limited ground-truth dataset for evaluation (50 papers for main benchmark, 500 for extended evaluation); Synthetic query generation via LLM perturbation may bias toward certain query patterns; Potential model-selection bias in choice of Qwen3 models for primary benchmarking; Evaluation on computer science/ML conferences only (may not generalize to other domains); Selection bias in query dataset (50 queries for semantic benchmark, 500 for RA-Bench); Potential model-specific biases in LLM-based ranking agents; Citation count bias favoring established papers (mitigated by making it optional); Venue-specific bias in knowledge graph construction; Limited evaluation of review generation against human judgments (r < 0.25 correlation)

Limitations

  • The authors state that "Our review agent shows weak alignment with human judgments: across models, the correlation with human reviewer scores remains low (r < 0.25), and several metrics can even exhibit negative correlations, indicating that the system may rank papers in the opposite order of human preference
  • As a result, even the best-performing configurations do not reliably distinguish strong from weak submissions, and the system should not be used as a trusted mechanism for comparing or ranking papers."

Open questions raised

  • "Future work will focus on the optimization of the unification of the pipeline." Additionally, the paper identifies that "review quality improves with larger models, suggesting that capacity and instruction-following are particularly important for end-to-end reviewing" and notes that weak correlation between automated reviews and human judgments (r < 0.25) indicates this limitation can potentially "be overcome by large open/closed source models."
  • "Future work will focus on the optimization of the unification of the pipeline." The authors identify that larger open/closed source models could overcome alignment challenges with human reviewer judgments.
  • Future work will focus on optimization of pipeline unification. The weak correlation between automated review scores and human judgments (r < 0.25) suggests need for improvement in review agent alignment. The authors suggest this problem can be overcome by using larger open/closed source models.
Data: Database corpus: 297 papers from leading CS and ML conferences (Table 2) sourced from OpenReview; ICLR 2024 reviews dataset used for paper review analysis (50 papers); ICLR 2024 reviews (50 papers randomly selected for review analysis); Paper database curated from OpenReview and conferences (ICLR, NeurIPS, ICML, CVPR, IROS, ICRA, AAAI, ACL, ICCV, EMNLP) - 21,115 papers total; Curated corpus of research papers from OpenReview, augmented with metadata and peer-review information (Table 2 shows papers from ICLR, NeurIPS, ICML, CVPR, IROS, ICRA, AAAI, ACL, ICCV, EMNLP and other venues); ICLR 2024 reviews dataset (50 randomly selected papers for review evaluation)Code: GitHub: github.com/MAXNORM8650/papercircle; Website: papercircle.vercel.app/; github.com/MAXNORM8650/papercircle; papercircle.vercel.app/ (website)Extracted from: pdfAgreement 64%

Explore related topics

Related papers