12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Learning from AVA: Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research

Nimisha Karnatak, Mohamad Chatila, Daniel Alejandro Pinzon Hernandez, Reza Yazdanfar, Michelle Dugas, Renos Vakis · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
E
Evidence
1
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1145/3772318.3791062

Methodology & findings

Study design

Mixed-methods, randomized-invitation field evaluation combining: (1) longitudinal platform logs (timestamps, session identifiers, query text, response metadata); (2) baseline and endline surveys (2,259 registrants matched n=1,029); (3) in-app pop-up surveys during deployment; and (4) 20 semi-structured qualitative interviews.

Sample

N = 2764, 17 groups

Primary method

Intention-to-treat (ITT) analysis comparing treatment vs. control groups using endline survey outcomes; Difference-in-Differences (DiD) estimation leveraging baseline and endline responses; Linear regression with robust standard errors; Natural language processing (NLP) for query classification using deterministic rule-based taxonomy system; Forward-filling imputation for uncategorized queries (one-hour inactivity window); Session-based user segmentation (single-session vs. multi-session cohorts)

Main result

The study found that sustained engagement with AVA was associated with substantial time savings, with Difference-in-Differences estimates associating "sustained engagement with 2.4–3.9 hours saved weekly." Qualitatively, participants used AVA as a specialized "evidence engine"; reasoned abstention clarified scope boundaries, and trust was calibrated through institutional provenance and page-anchored citations.

Reports effect sizes.

Research paradigm

Mixed-methods pragmatism (quantitative + qualitative)

Author conclusions

The authors conclude that "We derive generalisable lessons for building trustworthy knowledge systems that can inform deployments in other high-stakes domains, including the need for an end-to-end trust pipeline (from corpus curation through abstention to verification), strategies for managing corpus quality–coverage trade-offs, and interface patterns that prioritise verification over disclosure. We also articulate a vision for ecosystem-aware specialised AI systems that prioritise collaborative interoperability with general-purpose models."

Risk of bias

Selection bias: Voluntary participation and self-selection into treatment/control groups despite randomization; Attrition bias: 243 treatment and 121 control participants completed endline survey from 2,764 initial registrants (8.8% treatment, 21.9% control completion rates); Language bias: 83.8% of interactions in English despite 60+ language support, limiting non-English sample conclusions; Institutional affiliation bias: 95.4% external participants may skew findings toward non-embedded users; Social desirability bias in qualitative interviews (n=20) and in-app pop-up surveys; Hawthorne effect: Being observed may alter user behavior in-the-wild deployment; Selection bias: 95.4% external participants; voluntary participation via recruitment through multilateral development bank may introduce professional/institutional bias; Attrition bias: Only 364 of 2,764 registrants (13.2%) completed endline survey (243 treatment, 121 control); Self-selection: Participants who sustained engagement may differ systematically from those who disengaged; Language bias: 83.8% of interactions in English despite 60+ language support, limiting multilingual findings; Interview sampling: 20 semi-structured interviews via purposive sampling may not represent full user population; Attrition bias: Low endline survey completion rate (9.5% treatment, 21.9% control) relative to registrants; Selection bias: Voluntary participation and self-selection into treatment arm; Open-label design: Participants and researchers aware of treatment assignment; Language bias: 83.8% of interactions in English despite 60+ language support; Self-reported outcomes: Productivity and efficiency measures rely on participant perception rather than objective measurement

Open questions raised

  • Limited understanding of how Humble AI mechanisms function in generative systems (versus predictive systems that classify or rank data)
  • Lack of longitudinal, in-the-wild evidence on how such systems integrate into everyday professional practice
  • Few multi-month, mixed-methods deployments in evidence-dependent professional domains connecting log-level behaviors with self-reported outcomes
  • Limited methods for evaluating multilingual coverage and workflow integration in domain-bounded tools used by globally distributed professionals
  • Need for ecosystem-aware AI systems that prioritize collaborative interoperability with general-purpose models
  • Lack of longitudinal, in-the-wild evidence on how humility mechanisms function in systems that produce new text (generative AI) versus merely score/rank existing data (predictive AI)
Data: No public datasets mentioned as available in the provided text. The paper notes that operational logs were pseudonymized, but no data repository URL or access instructions are provided.; Operational logs were "pseudonymized with hashed identifiers." No explicit statement that data will be made publicly available. The paper mentions analysis of "usage logs linked to baseline and endline surveys" but does not state availability.Code: No code repositories mentioned in the provided text; Not mentioned in the extracted textExtracted from: pdfAgreement 43%

Explore related topics

Related papers