DRACULA: Hunting for the Actions Users Want Deep Research Agents to Execute
Nishant Balepur, Malachi Hamada, Varsha Kishore, Sergey Feldman, Amanpreet Singh, Pao Siangliulue et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Mixed-methods study combining: (1) User annotation study over five weeks with 19 CS researchers selecting actions and judging execution quality (450 annotation hours); (2) Simulation experiments using LLMs (GPT-5, Claude-4.6 Opus, Gemini-3 Pro, DeepSeek-R1, Qwen-3) to predict user action selections; (3) Follow-up user study testing personalized action generation based on different user contexts..
Main result
The study reveals three key findings: (1) "LLM judges initially struggle to predict action selections, but improve most when using a user's full selection history, rather than self-reported or extrapolated user context signals"; (2) Users' selections for the same query differ based on unstated goals, bottlenecking simulation and motivating affordances that let users steer reports; and (3) actions generated from revealed preferences (past selections) were selected most often, with selection rates of 0.857 compared to 0.682 for research papers and 0.741 for stated preferences.
Research paradigm
Empirical, mixed-methods (simulation + user studies)
Author conclusions
The authors conclude that "while researchers extensively study how well DR systems can execute actions, we show a key bottleneck is predicting which actions users want in the first place—a difficult task that largely benefits from user-specific modeling." They state that "as execution horizons scale, agents must take increasingly complex actions to support users, making intermediate feedback beyond the final output key for improving their utility." Furthermore, "it is promising that gains from user-specific context in action prediction correspond to improvements in generated actions" and recommend that "future work explore the extent to which simulation accuracy on real user feedback can inform online interventions to improve agent utility."
Risk of bias
Selection bias: Only 19 CS researchers recruited from Upwork, self-selected as active DR users via pilot survey; Annotation bias: One annotator omitted for reusing rationales; quality control via IRB approval but limited inter-rater reliability reporting; Model selection bias: Five specific LLM judges tested; results may not generalize to other models; Attrition: One annotator removed for low quality; unclear if other dropouts occurred; Cold-start bias: Initial action generation conditioned on user-selected papers showed lower selection rates than generic actions, suggesting preference for less personalized suggestions initially; Selection bias: Annotators were recruited from Upwork and verified as active DR users, potentially excluding less engaged users; Annotator bias: 19 CS researchers from specific fields (Security, CV, NLP, CompBio, AI Ethics) may not represent broader research populations; Attrition: One annotator excluded for reusing rationales; ten queries omitted as out of scope; Social desirability bias: Researchers may tailor selections to perceived expectations; Model-dependent bias: LLM judges used for de-duplication and coding procedures may introduce systematic biases; Selection bias: Only 19 CS researchers from Upwork verified as active DR users; may not represent broader scientific community; Annotation quality: One annotator excluded for reusing rationales; Temporal bias: User preference stability study conducted 5 months later, subject to memory effects despite attempt to limit them; Model-centric bias: Heavy reliance on frontier LLMs (GPT-5, Claude-4.6, etc.) for both action generation and evaluation; Limited domain: Restricted to CS research queries from deployed ScholarQA system; Annotator fatigue: 450 hours across 19 researchers may introduce fatigue bias
Limitations
- The authors acknowledge several limitations: (1) "User stability—predicting initial user selections from their new ones—has F1=0.672, near GPT-5 on the same subset using user-specific history (0.664)", indicating inherent unpredictability in user preferences
- (2) "Only 19 cases" where users cited system limitations for rejecting actions, but the sample of 19 researchers may limit generalizability
- (3) The paper states "we omit noisy test set items where GPT-5 cannot predict the true label via the user's rationale or where we flag low-quality queries (§2)
- this impacts <5% of the data", indicating data cleaning that may affect representativeness
- (4) The study involved only CS researchers, limiting applicability to other scientific domains.
Open questions raised
- Action feedback for long-horizon agents remains understudied compared to execution-focused work
- User-specific modeling for action prediction and generation needs further development
- Affordances and mechanisms to clarify user goals in interactive DR systems are underexplored
- Cold-start personalization signals for DR (beyond research papers) require investigation
- Scaling intermediate feedback collection and simulation-based evaluation for long-form agent tasks
- Most DR feedback protocols only score final reports; intermediate action feedback is rarely collected
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations