Learning from AVA: Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research
Nimisha Karnatak, Mohamad Chatila, Daniel Alejandro Pinzon Hernandez, Reza Yazdanfar, Michelle Dugas, Renos Vakis · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1145/3772318.3791062
Methodology & findings
Study design
Mixed-methods, randomized-invitation field evaluation combining: (1) longitudinal platform logs (timestamps, session identifiers, query text, response metadata); (2) baseline and endline surveys (2,259 registrants matched n=1,029); (3) in-app pop-up surveys during deployment; and (4) 20 semi-structured qualitative interviews.
Sample
N = 2764, 17 groups
Primary method
Intention-to-treat (ITT) analysis comparing treatment vs. control groups using endline survey outcomes; Difference-in-Differences (DiD) estimation leveraging baseline and endline responses; Linear regression with robust standard errors; Natural language processing (NLP) for query classification using deterministic rule-based taxonomy system; Forward-filling imputation for uncategorized queries (one-hour inactivity window); Session-based user segmentation (single-session vs. multi-session cohorts)
Main result
The study found that sustained engagement with AVA was associated with substantial time savings, with Difference-in-Differences estimates associating "sustained engagement with 2.4–3.9 hours saved weekly." Qualitatively, participants used AVA as a specialized "evidence engine"; reasoned abstention clarified scope boundaries, and trust was calibrated through institutional provenance and page-anchored citations.
Reports effect sizes.
Research paradigm
Mixed-methods pragmatism (quantitative + qualitative)
Author conclusions
The authors conclude that "We derive generalisable lessons for building trustworthy knowledge systems that can inform deployments in other high-stakes domains, including the need for an end-to-end trust pipeline (from corpus curation through abstention to verification), strategies for managing corpus quality–coverage trade-offs, and interface patterns that prioritise verification over disclosure. We also articulate a vision for ecosystem-aware specialised AI systems that prioritise collaborative interoperability with general-purpose models."
Risk of bias
Selection bias: Voluntary participation and self-selection into treatment/control groups despite randomization; Attrition bias: 243 treatment and 121 control participants completed endline survey from 2,764 initial registrants (8.8% treatment, 21.9% control completion rates); Language bias: 83.8% of interactions in English despite 60+ language support, limiting non-English sample conclusions; Institutional affiliation bias: 95.4% external participants may skew findings toward non-embedded users; Social desirability bias in qualitative interviews (n=20) and in-app pop-up surveys; Hawthorne effect: Being observed may alter user behavior in-the-wild deployment; Selection bias: 95.4% external participants; voluntary participation via recruitment through multilateral development bank may introduce professional/institutional bias; Attrition bias: Only 364 of 2,764 registrants (13.2%) completed endline survey (243 treatment, 121 control); Self-selection: Participants who sustained engagement may differ systematically from those who disengaged; Language bias: 83.8% of interactions in English despite 60+ language support, limiting multilingual findings; Interview sampling: 20 semi-structured interviews via purposive sampling may not represent full user population; Attrition bias: Low endline survey completion rate (9.5% treatment, 21.9% control) relative to registrants; Selection bias: Voluntary participation and self-selection into treatment arm; Open-label design: Participants and researchers aware of treatment assignment; Language bias: 83.8% of interactions in English despite 60+ language support; Self-reported outcomes: Productivity and efficiency measures rely on participant perception rather than objective measurement
Open questions raised
- Limited understanding of how Humble AI mechanisms function in generative systems (versus predictive systems that classify or rank data)
- Lack of longitudinal, in-the-wild evidence on how such systems integrate into everyday professional practice
- Few multi-month, mixed-methods deployments in evidence-dependent professional domains connecting log-level behaviors with self-reported outcomes
- Limited methods for evaluating multilingual coverage and workflow integration in domain-bounded tools used by globally distributed professionals
- Need for ecosystem-aware AI systems that prioritize collaborative interoperability with general-purpose models
- Lack of longitudinal, in-the-wild evidence on how humility mechanisms function in systems that produce new text (generative AI) versus merely score/rank existing data (predictive AI)
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations
- Human-AI collaboration patterns in AI-assisted academic writingAndy Nguyen · 2024 · 301 citations