12,637 papers · updated 18 Sept 2026livingmeta.ai
← Browse all papers
AI evidence extraction

DeepResearch_Eco: A Recursive AgenticWorkflow for Complex Scientific Question Answering in Ecology

Jennifer D’Souza, Endres Keno Sander, Andrei Aioanei · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
D
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.32942/x2m06g

Methodology & findings

Study design

Artifact evaluation study using computational experiments.

Primary method

Design science research with computational system development and empirical evaluation

Main result

The study found that "a high-depth configuration can automatically integrate information from 111 sources-nearly 6× more than a shallow setting-and increase coverage of key concepts by 25%" and demonstrates that "carefully configured, high-parameter runs can approach expert-level integration in ecology, achieving an order-of-magnitude higher information density in outputs without loss of rigor or specificity." Maximum parameter configurations achieved a 21.2-fold increase in source utilization (from 9.1 to 192.9 sources) while word count expanded only 41.5%, yielding a 14.9-fold enhancement in information density.

Research paradigm

Empiricist/Pragmatist - testing a computational artifact through experiments and qualitative evaluation

Author conclusions

The authors conclude that "Increasing depth and breadth parameters improves analytical rigor, evidence diversity, and ecological specificity" and that "DeepResearch enables structured, transparent, and expert-like synthesis with tunable analytical control." They state that "DeepResearch Eco, a recursive, agentic workflow for controllable scientific synthesis, validated on 49 ecological research questions" demonstrates that the system successfully achieves the goal of balancing breadth versus depth in automated literature synthesis.

Risk of bias

Selection bias in ecological research questions (sourced from 9 fellows of specific interdisciplinary group); Source availability bias in retrieval systems (noted that low-breadth syntheses focus on 'well-studied European or North American temperate systems'); API-dependent biases in both Firecrawl and ORKG Ask search providers; Domain specificity: Evaluation limited to ecology may not generalize to other fields; LLM source selection bias: ORKG Ask corpus composition may favor certain publication types; Evaluation metric limitations: Automated quality metrics (substring matching, keyword detection) may not capture full semantic quality; Model-specific effects: Results tied to OpenAI models (o3, o3-mini); generalizability to other LLM providers unknown; Evaluation metrics are primarily automated (substring matching, keyword detection) rather than human expert assessment; No independent validation of quality assessment framework - weights and thresholds appear author-determined; Comparison limited to two OpenAI models only; No inter-rater reliability reported for qualitative evaluation; Potential circular evaluation: metrics designed around observed paper characteristics

Limitations

  • The paper states "Future work will also address current limitations by implementing an interactive agent for researcher feedback integration, enabling guided refinement across recursive steps" and notes that "Support for multimodal synthesis-including figures and tables-will be explored to enhance utility in data-rich fields." The evaluation is limited to ecology
  • "Future work will focus on evaluating DeepResearch across additional domains beyond ecology, such as materials science and social science, to further demonstrate its generality and adaptability."

Open questions raised

  • Evaluation across domains beyond ecology (materials science, social science)
  • Interactive agent for researcher feedback integration across recursive steps
  • Multimodal synthesis support including figures and tables
  • Collaborative agentic workflows enabling distributed synthesis across teams or disciplines
  • Model generality: Evaluation with non-OpenAI LLM providers to test robustness beyond proprietary models
  • Future work will focus on:
Data: https://github.com/sciknoworg/deep-research/blob/main/data/49-questions.csvCode: https://github.com/sciknoworg/deep-research (MIT license)Extracted from: pdf

Explore related topics

Related papers