12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps

Tanmay Asthana, Aman Saksena, Divyansh Sahu · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Benchmark evaluation with 42 SME-authored prompts in the Management Consulting domain tested across 3 deep research agents (126 total responses).

Sample

N = 126, 6 groups

Primary method

Paired McNemar tests on agent comparisons (McNemar 1947, Dietterich 1998) with multiple-testing correction via Holm method (Holm 1979, Demšar 2006). Cohen's d computed as (r1 - r2) / s_pooled on reasoning average r with conventional magnitude thresholds (0.2 small, 0.5 medium, 0.8 large; Sawilowsky 2009). Pearson and Spearman rank correlations reported for criterion validation (Pearson correlations; Spearman for ordinal rubric scores). Sole-cause failure analysis using logical filtering on (V ≥ 80%, min_i r_i > 0, r ≥ 2.5) decision thresholds. No formal software packages named; evaluation infrastructure implemented in Python with LLM-based classification (Claude Opus 4.7) for failure-mode tagging.

Main result

The study found that "Claude produces deliverable artifacts most reliably (4.5× the file-output rate of either other agent; 9 of 10 file-required tasks vs. 3 and 1), yet carries the highest fabrication signature when graded for content correctness. o3 attains the highest mean rubric score among the three but is caught by the verifier layer on dropped required sections and on cascading arithmetic errors that propagate across multi-step calculations. Gemini swings between high-quality responses and outright zero-scored ones more often than the other two combined (41 zero-scored rubric cells vs. 30 for Claude and 10 for o3)." The agents rank differently by metric: by strict VRS, o3 leads at 62.6; by per-prompt VRS argmax, Gemini leads with 19 of 42 prompts; by binary ACCEPT rate, Gemini leads at 21.4%.

Reports effect sizes and confidence intervals.

Research paradigm

Empirical-positivist (benchmark evaluation with structured measurement)

Author conclusions

The authors conclude: "What our benchmark actually measures is whether a frontier deep research agent can do the kind of structured, multi-document, decision-grade research a management consultant gets paid to do. Across the 42 graded prompts and 126 responses, the answer is: not yet, not reliably, and not in a way that any single performance metric captures." They further state: "The orderings are not in conflict: o3 builds its VRS lead on a heavier mass of non-zero-but-mediocre responses that fail the ACCEPT bar, while Gemini's higher zero count is offset by its non-failures more often clearing the production-quality threshold. Reporting any one of these views in isolation would mislead." Regarding rubric dimensionality, they note: "Across these views, the failure modes are agent-specific (Section 4.6), the prompt taxonomy probes capabilities the agents differ meaningfully on (Cohen's d > 1.0 on two of five classes), and the five rubric criteria correlate at mean Pearson 0.61, consistent with two to three latent factors rather than five orthogonal traits."

Risk of bias

Small per-class sample sizes (n=6-11) reduce inferential precision and inflate Type I/II error risk; Single SME grader per prompt (no parallel re-annotation) introduces potential single-rater drift; QC layer asymmetric (verification not re-grading), does not yield inter-rater reliability coefficient; Agent architectural differences (code execution environments, API surfaces, file handling) not fully controlled; Claude has sandbox access advantage, Gemini requires local subprocess execution; Closed-corpus enforcement intrinsically limited (agents are 'partial black boxes'; authors note enforcement is something 'measured rather than guaranteed'); Selection bias in SME pool (recruited from 'former MBB and Big Four consultants'); may bias rubric toward consultant-domain norms; Verifier suite task-specific; domain knowledge required to validate correctness of verifiers themselves; No blinding of SMEs to agent identity (rubric-grading may be influenced by knowledge of agent source); Single-domain evaluation (Management Consulting only); results may not generalize to other professional domains; Small sample size per prompt class (n = 6-11) limits inferential precision on effect sizes; Single primary SME grader per response (no inter-rater reliability coefficient reported; QC is verification not re-annotation); Subjective ordinal scoring (0-3) despite rubric guidelines; potential for scorer drift; Architectural differences between agents (API surface, code-execution environment, file handling) may confound capability assessment; Cognitive traps embedded in prompts may advantage certain agent architectures over others; Single-domain evaluation (Management Consulting only); generalization to other professional domains uncertain; Single SME grading per prompt (no parallel double-blind annotation); potential for individual rater bias; Small sample size per prompt class (n=6-11) limits statistical power and precision of effect estimates; Selection bias in SME pool recruitment (former MBB and Big Four consultants may have domain-specific calibrations); Agent API architectural differences (Claude sandbox, o3 containers, Gemini local subprocess) could confound code-generation performance with model capability; Web search enforcement variability across agents (Claude injectable, o3 soft enforcement, Gemini always-on) affects CRP performance comparability

Limitations

  • The authors explicitly state several limitations: "Sample size, inter-rater reliability, and significance (P0)
  • The evaluation comprises 42 graded prompts × 3 agents (126 attempts), with per-class sample sizes of n = 6-11
  • At these sizes Cohen's d > 0.8 thresholds describe magnitude rather than inferential precision, and paired-comparison tests on per-agent ACCEPT outcomes do not reach p < 0.05 (Section 4.10.4)." Additionally, "Each (prompt × agent) cell was graded by one primary SME and independently reviewed by a second SME via the QC protocol of Appendix C
  • This provides an error-correction pass against rubric-evidence mismatches and surfaces fabricated citations, but does not yield a Cohen's κ inter-rater reliability statistic since QC is an asymmetric defensibility check rather than a parallel re-annotation." The authors further note: "The evaluation is single-domain (MC)
  • generalization to other professional domains requires parallel datasets" and "The mean off-diagonal Pearson correlation of 0.61 across the five reasoning criteria (Section 4.9) suggests 2-3 effective latent factors rather than five orthogonal traits, motivating a factor analysis."

Open questions raised

  • Planned v2 release will (i) roughly double prompt count and add Investment Banking (IB) domain alongside Management Consulting; (ii) conduct formal inter-rater reliability (IRR) study with parallel double-grading on held-out subset; (iii) report bootstrap confidence intervals on all headline metrics; (iv) perform factor analysis to determine true latent structure of five rubric criteria (mean off-diagonal ρ≈0.61 suggests 2-3 underlying factors); (v) add cross-validated weight recalibration for VRS composite score; (vi) extend evaluation to additional professional domains beyond MC and IB.
  • Formal inter-rater reliability (Cohen's κ) study needed; v2 will include parallel double-grading on held-out subset
  • Rubric dimensionality analysis required; mean off-diagonal Pearson correlation of 0.61 suggests 2-3 latent factors rather than five orthogonal traits; formal factor analysis listed as P1 priority
  • Single-domain evaluation (MC only); generalization to other professional domains requires parallel datasets including Investment Banking (IB) and additional sectors planned for v2
  • Larger corpus needed for cross-validated weight calibration of VRS aggregate
  • Bootstrap confidence intervals and paired-comparison adjustments recommended before external publication
Data: DRA Benchmark (42-prompt corpus); 42-prompt benchmark corpus with verifier specifications and cognitive-trap definitions released at https://huggingface.co/datasets/deccan-ai/dra-bench; 42-prompt corpus with verifier specifications, authorized-source lists, and cognitive-trap definitions released at: https://huggingface.co/datasets/deccan-ai/dra-benchCode: GitHub: https://github.com/tanm-ast-deccan/dra-response-gen; Full evaluation infrastructure (agent adapters, dispatch loop, result-store API, diagnostic tooling) available at https://github.com/tanm-ast-deccan/dra-response-gen; Full evaluation infrastructure (agent adapters, dispatch loop, result-store API, diagnostic tooling) available at: https://github.com/tanm-ast-deccan/dra-response-genExtracted from: pdfAgreement 44%

Explore related topics

Related papers