Paper Circle: An Open-source Multi-agent Research Discovery and Analysis Framework
Komal Kumar, Aman Chadha, Salman Khan, Fahad Shahbaz Khan, Hisham Cholakkal · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Design science approach with artifact development, evaluation through multi-query benchmarking (50-500 queries), ablation studies comparing retrieval baselines and pipeline configurations, and qualitative assessment through 81 real-world discovery sessions conducted by researchers.
Primary method
Design science with iterative development of multi-agent architecture based on smolagents library, incorporating user feedback through real-world evaluation sessions
Main result
The study found that "Two agent models achieve the highest retrieval effectiveness with an 80% hit rate, qwen3-coder-30b-Q3KM (quantized) and qwen3-coder:30b—with qwen3-coder-30b-Q3KM also delivering the best ranking quality (MRR = 0.627)" and demonstrated that "Paper Circle provides built-in evaluation metrics" with "Recall@K, Precision@K, and hit rates" computed across multiple configurations. Additionally, "Preliminary user feedback indicates minimal cognitive load when using PaperCircle. NASA-LLX assessment yields an overall workload of 1.2/7, with five of six dimensions scoring the minimum (1/7) and effort at 2/7."
Research paradigm
Design Science Research (DSR) / engineering paradigm
Author conclusions
"Paper Circle shows how multi-agent workflows can streamline research literature management. Its discovery pipeline unifies heterogeneous search sources and multi-criteria scoring into a reproducible tool, using a simple agent–tool interface with shared state, deterministic ranking, and synchronized multi-format outputs. Its analysis pipeline converts papers into structured knowledge graphs that enable graph-aware QA, coverage checks, and human-in-the-loop verification."
Risk of bias
Limited evaluation dataset (50 ICLR 2024 papers for review analysis, 50 queries for semantic benchmark); Potential selection bias in user study participants (81 sessions, 78 unique queries from researchers across 9 domains); Model-dependent performance variations (substantial differences between agent models tested); Recency bias in scoring framework (more recent papers receive higher scores); Limited ground truth validation for knowledge graph extraction quality; Limited ground-truth dataset for evaluation (50 papers for main benchmark, 500 for extended evaluation); Synthetic query generation via LLM perturbation may bias toward certain query patterns; Potential model-selection bias in choice of Qwen3 models for primary benchmarking; Evaluation on computer science/ML conferences only (may not generalize to other domains); Selection bias in query dataset (50 queries for semantic benchmark, 500 for RA-Bench); Potential model-specific biases in LLM-based ranking agents; Citation count bias favoring established papers (mitigated by making it optional); Venue-specific bias in knowledge graph construction; Limited evaluation of review generation against human judgments (r < 0.25 correlation)
Limitations
- The authors state that "Our review agent shows weak alignment with human judgments: across models, the correlation with human reviewer scores remains low (r < 0.25), and several metrics can even exhibit negative correlations, indicating that the system may rank papers in the opposite order of human preference
- As a result, even the best-performing configurations do not reliably distinguish strong from weak submissions, and the system should not be used as a trusted mechanism for comparing or ranking papers."
Open questions raised
- "Future work will focus on the optimization of the unification of the pipeline." Additionally, the paper identifies that "review quality improves with larger models, suggesting that capacity and instruction-following are particularly important for end-to-end reviewing" and notes that weak correlation between automated reviews and human judgments (r < 0.25) indicates this limitation can potentially "be overcome by large open/closed source models."
- "Future work will focus on the optimization of the unification of the pipeline." The authors identify that larger open/closed source models could overcome alignment challenges with human reviewer judgments.
- Future work will focus on optimization of pipeline unification. The weak correlation between automated review scores and human judgments (r < 0.25) suggests need for improvement in review agent alignment. The authors suggest this problem can be overcome by using larger open/closed source models.
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations