12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Queryome: Orchestrating Retrieval, Reasoning, and Synthesis across Biomedical Literature

Pranav Punuru, Nabil Ibtehaz, Swagarika Jaharlal Giri, Harsha Srirangam, Emilia A Tugolukova, Daisuke Kihara · bioRxiv (Cold Spring Harbor Laboratory) · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
C
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.64898/2025.12.22.696019

Methodology & findings

Study design

Computational system design and evaluation.

Main result

On the MIRAGE benchmark, "Queryome achieved 88.98% accuracy, surpassing prior systems by up to 14 points," and "improved reasoning accuracy on the biomedical Human's Last Exam (HLE) subset from 15.8% to 19.3%." Additionally, "in a task for constructing a review article, it earned the highest composite score in comparison with Deep Research from OpenAI, Google, Perplexity, and Scite.AI, reflecting its strong literature retrieval and synthesis capabilities."

Research paradigm

Computational/empirical; systems engineering approach to biomedical information retrieval and synthesis

Author conclusions

"Queryome represents a step toward AI systems that engage directly with the empirical foundations of biomedicine. The long-term vision is not automation but collaboration: an ecosystem of reasoning agents that extend human scientific inquiry, preserving rigor while amplifying reach." The authors note that "the improvement arises when retrieved evidence is interpreted through reasoning, not when it is simply appended to a prompt. In tasks demanding causal inference, temporal reasoning, or integration across biological domains, Queryome consistently constructs explanations that trace why a conclusion follows from literature."

Risk of bias

Model selection bias: o3 model used for PI agent may be optimized for specific benchmarks (MedQA, MedMCQA); Benchmark annotation bias: PubMedQA ground-truth labels shown to be based on simplistic single-paper extraction rather than comprehensive evidence synthesis; Retrieval noise: Ambiguous, conflicting, or spurious retrievals can amplify errors, particularly when distractors in questions are partially supported by retrieved text; Training data bias: General-domain LLMs have limited grasp of biomedical language; biomedical training data is scarce and fragmented; Potential overfitting to MIRAGE benchmark due to high performance (88.98%); Model overfitting to benchmark datasets (especially MMLU, MedQA, MedMCQA where improvements were marginal or negative relative to base model); Noise from ambiguous or conflicting retrievals amplifying errors; Handcrafted search policies not data-driven; Lack of access to full-text articles (abstract-only limitation); Temperature parameter of o3 model not exposed, requiring multiple runs to quantify variance; Potential optimization bias toward specific exam heuristics in licensing exam datasets; Benchmark selection bias: System performance on MIRAGE may not generalize to other medical QA tasks not represented in the benchmark; Model memorization bias: Comparisons with base models (o3, GPT-5) may conflate agentic orchestration improvements with inherent model capabilities; Training data bias: PubMed corpus may underrepresent certain medical specialties, geographic regions, or publication types; Evaluation bias: Modified prompt for PubMedQA to match benchmark construction methodology may artificially inflate reported accuracy; Retrieval bias: Abstract-only retrieval may miss context available in full-text articles, potentially biasing toward certain types of medical knowledge

Limitations

  • "Queryome remains constrained by its design choices
  • Limiting retrieval to PubMed ensures quality control but omits key knowledge sources such as clinical guidelines, preprints, and multimodal data
  • Its search policy, though effective, is still handcrafted rather than being data driven." Additionally, "The limitation of retrieving only from abstracts rather than full-text articles contributes to this noise
  • Access to the full texts would likely enable more grounded knowledge discovery and deep reasoning."

Open questions raised

  • Integration of reinforcement or meta-learning to refine investigative heuristics
  • Expansion to multilingual and multimodal retrieval
  • Incorporation of evidence quality weighting based on study design and bias
  • Introduction of temporal tracking to monitor shifts in scientific consensus
  • Development of interfaces for clinical or research use that are safe, transparent, and trustworthy
  • Access to broader range of biological databases, wikis, and scientific literature beyond PubMed
Data: MIRAGE benchmark (7,663 questions from 5 datasets: MMLU-Med, MedQA-US, MedMCQA, PubMedQA, BioASQ-Y/N); Humanity's Last Exam (HLE) biomedical subset (222 questions); PubMed Baseline collection (28.3 million abstracts with full metadata); PubMedQA dataset; BioASQ-Y/N dataset; MIRAGE benchmark (7,663 questions from five datasets: MMLU-Med, MedQA-US, MedMCQA, PubMedQA, BioASQ-Y/N); PubMed Baseline collection (28.3 million articles with abstracts, as of September 2025); PubMed baseline collection (28.3 million articles with abstracts as of September 2025); Humanity's Last Exam biomedical subset (222 questions)Code: Not explicitly provided in the paper. System prompts mentioned as provided in 'Supplementary Information 1'; Not explicitly mentioned; full system prompts provided in Supplementary InformationExtracted from: pdfAgreement 46%

Explore related topics

Related papers