12,637 papers · updated 18 Sept 2026livingmeta.ai
← Browse all papers
AI evidence extraction

OpenResearcher: Unleashing AI for Accelerated Scientific Research

Yu-Xiang Zheng, Shichao Sun, Lin Qiu, Dongyu Ru, Xuefeng Li, Jifan Lin et al. · arXiv (Cornell University) · 2024

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
D
Evidence
1
Citations

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.48550/arxiv.2408.06941

Methodology & findings

Study design

Comparative evaluation using human preference assessment (pairwise comparison with 12 student evaluators) and LLM-based preference evaluation using GPT-4o.

Primary method

Design science research with iterative development of system components (query tools, retrieval tools, post-processing tools, generation tools, refinement tools) and evaluation through human preference studies and LLM-based comparison.

Main result

OpenResearcher achieves superior information correctness, relevance, and richness compared to all other applications. Specifically, "Our OpenResearcher achieves superior information correctness, relevance, and richness compared to all other applications. OpenResearcher significantly outperforms Perplexity AI with more "Win" than "Lose"." In human evaluation across 30 questions, OpenResearcher achieved 10 wins versus 7 losses in correctness, 25 wins versus 1 loss in richness, and 15 wins versus 2 losses in relevance compared to Perplexity AI.

Research paradigm

pragmatist/design-science

Author conclusions

"We introduce OpenResearcher, an active AI assistant to accelerate the research process, catering to a broad spectrum of inquiries from researchers. OpenResearcher employs Retrieval-Augmented Generation (RAG) to enhance LLMs with the latest, verified, and domain-specific knowledge... OpenResearcher can use these tools flexibly to build a pipeline that delivers accurate and comprehensive answers, outperforming those from industry applications, as judged by human and GPT-4o."

Risk of bias

Selection bias: Only 30 of 109 questions were evaluated by humans, not necessarily representative of all question types; Annotator bias: Only 12 students with 'good research experience' conducted evaluation, potentially non-representative of broader user population; Comparison bias: Perplexity AI was used as the baseline for comparison without justification for why this specific application was chosen; Limited dataset scope: Only arXiv papers from Jan 2023-Jun 2024 were included, excluding earlier and potentially more established literature; LLM evaluation limitations: GPT-4o cannot verify factual accuracy without external knowledge, limiting reliability of LLM-based preference evaluation; Annotator bias: Evaluation by 12 graduate students who may have domain expertise biases; Limited inter-annotator agreement resolution: Third annotators only involved for disagreements; LLM evaluation bias: GPT-4o has known preferences that differ from human judgment; Comparison bias: Baseline selection favors newer industry applications; older academic systems not included in detailed comparison; Selection bias in question collection (from graduate students only, specific research areas); potential preference bias in human annotators toward the OpenResearcher system being evaluated.

Limitations

  • The paper acknowledges that "Despite being instructed to ground the generated responses in retrieved knowledge from scientific publications, LLMs may still generate hallucinations." Additionally, the evaluation was limited to 30 questions selected from 109 total questions for human evaluation, and ground truth annotations were not provided due to "the considerable effort and cost of annotating ground truth answers." The system's databases only encompass arXiv publications from January 2023 to June 2024.

Open questions raised

  • The paper identifies gaps in existing academic and industry applications: academic applications focus on single tasks without unified solutions for diverse inquiries, while industry applications are proprietary and hinder academic research development. Both academic and industry applications serve as passive assistants rather than active communicators.
Data: arXiv publications from January 2023 to June 2024 with metadata (used in system but not explicitly stated as publicly available); 109 research questions collected from graduate students (38 on scientific paper recommendation, 38 on scientific text summarization, 33 on others); 109 research questions collected from graduate students available implicitly through the evaluation section.Code: OpenResearcher is described as an 'open-source project' but no specific GitHub or repository URL is provided in the paperExtracted from: pdf

Explore related topics

Related papers