Queryome: Orchestrating Retrieval, Reasoning, and Synthesis across Biomedical Literature
Pranav Punuru, Nabil Ibtehaz, Swagarika Jaharlal Giri, Harsha Srirangam, Emilia A Tugolukova, Daisuke Kihara · bioRxiv (Cold Spring Harbor Laboratory) · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.64898/2025.12.22.696019
Methodology & findings
Study design
Computational system design and evaluation.
Main result
On the MIRAGE benchmark, "Queryome achieved 88.98% accuracy, surpassing prior systems by up to 14 points," and "improved reasoning accuracy on the biomedical Human's Last Exam (HLE) subset from 15.8% to 19.3%." Additionally, "in a task for constructing a review article, it earned the highest composite score in comparison with Deep Research from OpenAI, Google, Perplexity, and Scite.AI, reflecting its strong literature retrieval and synthesis capabilities."
Research paradigm
Computational/empirical; systems engineering approach to biomedical information retrieval and synthesis
Author conclusions
"Queryome represents a step toward AI systems that engage directly with the empirical foundations of biomedicine. The long-term vision is not automation but collaboration: an ecosystem of reasoning agents that extend human scientific inquiry, preserving rigor while amplifying reach." The authors note that "the improvement arises when retrieved evidence is interpreted through reasoning, not when it is simply appended to a prompt. In tasks demanding causal inference, temporal reasoning, or integration across biological domains, Queryome consistently constructs explanations that trace why a conclusion follows from literature."
Risk of bias
Model selection bias: o3 model used for PI agent may be optimized for specific benchmarks (MedQA, MedMCQA); Benchmark annotation bias: PubMedQA ground-truth labels shown to be based on simplistic single-paper extraction rather than comprehensive evidence synthesis; Retrieval noise: Ambiguous, conflicting, or spurious retrievals can amplify errors, particularly when distractors in questions are partially supported by retrieved text; Training data bias: General-domain LLMs have limited grasp of biomedical language; biomedical training data is scarce and fragmented; Potential overfitting to MIRAGE benchmark due to high performance (88.98%); Model overfitting to benchmark datasets (especially MMLU, MedQA, MedMCQA where improvements were marginal or negative relative to base model); Noise from ambiguous or conflicting retrievals amplifying errors; Handcrafted search policies not data-driven; Lack of access to full-text articles (abstract-only limitation); Temperature parameter of o3 model not exposed, requiring multiple runs to quantify variance; Potential optimization bias toward specific exam heuristics in licensing exam datasets; Benchmark selection bias: System performance on MIRAGE may not generalize to other medical QA tasks not represented in the benchmark; Model memorization bias: Comparisons with base models (o3, GPT-5) may conflate agentic orchestration improvements with inherent model capabilities; Training data bias: PubMed corpus may underrepresent certain medical specialties, geographic regions, or publication types; Evaluation bias: Modified prompt for PubMedQA to match benchmark construction methodology may artificially inflate reported accuracy; Retrieval bias: Abstract-only retrieval may miss context available in full-text articles, potentially biasing toward certain types of medical knowledge
Limitations
- "Queryome remains constrained by its design choices
- Limiting retrieval to PubMed ensures quality control but omits key knowledge sources such as clinical guidelines, preprints, and multimodal data
- Its search policy, though effective, is still handcrafted rather than being data driven." Additionally, "The limitation of retrieving only from abstracts rather than full-text articles contributes to this noise
- Access to the full texts would likely enable more grounded knowledge discovery and deep reasoning."
Open questions raised
- Integration of reinforcement or meta-learning to refine investigative heuristics
- Expansion to multilingual and multimodal retrieval
- Incorporation of evidence quality weighting based on study design and bias
- Introduction of temporal tracking to monitor shifts in scientific consensus
- Development of interfaces for clinical or research use that are safe, transparent, and trustworthy
- Access to broader range of biological databases, wikis, and scientific literature beyond PubMed
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations
- Artificial intelligence and the conduct of literature reviewsGerit Wagner · 2021 · 275 citations
- Artificial intelligence for literature reviews: opportunities and challengesF. J. Bolaños · 2024 · 188 citations