12,637 papers · updated 18 Sept 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Learning from AVA: Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research

Nimisha Karnatak, Mohamad Chatila, Daniel Alejandro Pinzon Hernandez, Reza Yazdanfar, Michelle Dugas, Renos Vakis · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
E
Evidence
1
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1145/3772318.3791062

Methodology & findings

Study design

Mixed-methods, randomized-invitation field evaluation combining: (1) longitudinal platform logs (timestamps, session identifiers, query text, response metadata); (2) baseline and endline surveys (2,259 registrants matched n=1,029); (3) in-app pop-up surveys during deployment; and (4) 20 semi-structured qualitative interviews.

Sample

N = 2764, 6 groups

Primary method

Intention-to-treat (ITT) analysis comparing treatment vs. control groups using endline survey outcomes; Difference-in-Differences (DiD) estimation leveraging baseline and endline responses; Linear regression with robust standard errors; Natural language processing (NLP) for query classification using deterministic rule-based taxonomy system; Forward-filling imputation for uncategorized queries (one-hour inactivity window); Session-based user segmentation (single-session vs. multi-session cohorts)

Main result

The study found that sustained engagement with AVA was associated with substantial time savings, with Difference-in-Differences estimates associating "sustained engagement with 2.4–3.9 hours saved weekly." Qualitatively, participants used AVA as a specialized "evidence engine"; reasoned abstention clarified scope boundaries, and trust was calibrated through institutional provenance and page-anchored citations.

Reports effect sizes.

Research paradigm

Mixed-methods pragmatism (quantitative + qualitative)

Author conclusions

The authors conclude that "We derive generalisable lessons for building trustworthy knowledge systems that can inform deployments in other high-stakes domains, including the need for an end-to-end trust pipeline (from corpus curation through abstention to verification), strategies for managing corpus quality–coverage trade-offs, and interface patterns that prioritise verification over disclosure. We also articulate a vision for ecosystem-aware specialised AI systems that prioritise collaborative interoperability with general-purpose models."

Risk of bias

Selection bias: Voluntary participation and self-selection into treatment/control groups despite randomization; Attrition bias: 243 treatment and 121 control participants completed endline survey from 2,764 initial registrants (8.8% treatment, 21.9% control completion rates); Language bias: 83.8% of interactions in English despite 60+ language support, limiting non-English sample conclusions; Institutional affiliation bias: 95.4% external participants may skew findings toward non-embedded users; Social desirability bias in qualitative interviews (n=20) and in-app pop-up surveys; Hawthorne effect: Being observed may alter user behavior in-the-wild deployment; voluntary participation via recruitment through multilateral development bank may introduce professional/institutional bias; Self-selection: Participants who sustained engagement may differ systematically from those who disengaged; Interview sampling: 20 semi-structured interviews via purposive sampling may not represent full user population; Open-label design: Participants and researchers aware of treatment assignment; Self-reported outcomes: Productivity and efficiency measures rely on participant perception rather than objective measurement

Open questions raised

  • Limited understanding of how Humble AI mechanisms function in generative systems (versus predictive systems that classify or rank data)
  • Lack of longitudinal, in-the-wild evidence on how such systems integrate into everyday professional practice
  • Few multi-month, mixed-methods deployments in evidence-dependent professional domains connecting log-level behaviors with self-reported outcomes
  • Limited methods for evaluating multilingual coverage and workflow integration in domain-bounded tools used by globally distributed professionals
  • Need for ecosystem-aware AI systems that prioritize collaborative interoperability with general-purpose models
  • Absence of evidence on how verification-oriented affordances (page-level citations, reasoned refusals) are actually used by professionals in high-stakes workflows
Data: No public datasets mentioned as available in the provided text. The paper notes that operational logs were pseudonymized, but no data repository URL or access instructions are provided.; Operational logs were "pseudonymized with hashed identifiers." No explicit statement that data will be made publicly available. The paper mentions analysis of "usage logs linked to baseline and endline surveys" but does not state availability.Code: No code repositories mentioned in the provided textExtracted from: pdf

Explore related topics

Related papers