DeepResearch_Eco: A Recursive AgenticWorkflow for Complex Scientific Question Answering in Ecology
Jennifer D’Souza, Endres Keno Sander, Andrei Aioanei · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.32942/x2m06g
Methodology & findings
Study design
Artifact evaluation study using computational experiments.
Primary method
Design science research with computational system development and empirical evaluation
Main result
The study found that "a high-depth configuration can automatically integrate information from 111 sources-nearly 6× more than a shallow setting-and increase coverage of key concepts by 25%" and demonstrates that "carefully configured, high-parameter runs can approach expert-level integration in ecology, achieving an order-of-magnitude higher information density in outputs without loss of rigor or specificity." Maximum parameter configurations achieved a 21.2-fold increase in source utilization (from 9.1 to 192.9 sources) while word count expanded only 41.5%, yielding a 14.9-fold enhancement in information density.
Research paradigm
Empiricist/Pragmatist - testing a computational artifact through experiments and qualitative evaluation
Author conclusions
The authors conclude that "Increasing depth and breadth parameters improves analytical rigor, evidence diversity, and ecological specificity" and that "DeepResearch enables structured, transparent, and expert-like synthesis with tunable analytical control." They state that "DeepResearch Eco, a recursive, agentic workflow for controllable scientific synthesis, validated on 49 ecological research questions" demonstrates that the system successfully achieves the goal of balancing breadth versus depth in automated literature synthesis.
Risk of bias
Selection bias in ecological research questions (sourced from 9 fellows of specific interdisciplinary group); Source availability bias in retrieval systems (noted that low-breadth syntheses focus on 'well-studied European or North American temperate systems'); API-dependent biases in both Firecrawl and ORKG Ask search providers; Domain specificity: Evaluation limited to ecology may not generalize to other fields; LLM source selection bias: ORKG Ask corpus composition may favor certain publication types; Evaluation metric limitations: Automated quality metrics (substring matching, keyword detection) may not capture full semantic quality; Model-specific effects: Results tied to OpenAI models (o3, o3-mini); generalizability to other LLM providers unknown; Limited to 49 ecological research questions from single expert group (Mapping Evidence to Theory in Ecology fellows) - potential geographic/institutional bias; Evaluation metrics are primarily automated (substring matching, keyword detection) rather than human expert assessment; No independent validation of quality assessment framework - weights and thresholds appear author-determined; Comparison limited to two OpenAI models only; no comparison with other LLM providers or open-source alternatives; No inter-rater reliability reported for qualitative evaluation; Potential circular evaluation: metrics designed around observed paper characteristics
Limitations
- The paper states "Future work will also address current limitations by implementing an interactive agent for researcher feedback integration, enabling guided refinement across recursive steps" and notes that "Support for multimodal synthesis-including figures and tables-will be explored to enhance utility in data-rich fields." The evaluation is limited to ecology
- "Future work will focus on evaluating DeepResearch across additional domains beyond ecology, such as materials science and social science, to further demonstrate its generality and adaptability."
Open questions raised
- Evaluation across domains beyond ecology (materials science, social science)
- Interactive agent for researcher feedback integration across recursive steps
- Multimodal synthesis support including figures and tables
- Collaborative agentic workflows enabling distributed synthesis across teams or disciplines
- Generalizability beyond ecology: Need to evaluate across materials science and social science domains
- Interactive refinement: Implement agent for researcher feedback integration across recursive steps
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations
- Artificial intelligence and the conduct of literature reviewsGerit Wagner · 2021 · 275 citations