A Framework For Designing Ai-Supported Literature Reviews Under Varying Agency Roles: A Design Science Research Approach
Masoumeh Tavakoligargari, Matthias Bertram, Harald F. O. von Korflesch · Journal of the Association for Information Systems · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Design Science Research following Hevner et al.
Primary method
Design Science Research (DSR)
Main result
The study found that "many core principles, particularly transparency, systematic process rigor, monitoring, goal alignment, and adaptive collaboration, are readily implementable with current AI technologies." Additionally, "the prototype yielded several overarching insights into the practical feasibility and limitations of the proposed framework" including that "the design principles are broadly implementable with current-generation lightweight open source LLMs" and "low-resource deployment proved realistic; all tasks including multi-model comparison and reasoning-trace logging, ran smoothly on standard hardware without fine-tuning."
Research paradigm
Design Science Research (DSR)
Author conclusions
The authors conclude: "Our findings highlight that many core principles, particularly transparency, systematic process rigor, monitoring, goal alignment, and adaptive collaboration, are readily implementable with current AI technologies. At the same time, the scenarios reveal important boundaries: autonomous or principal-level AI behaviour still requires sustained human oversight for interpretive and theory-building tasks, and data-access limitations remain significant constraints." Furthermore, "The study offers several contributions. First, it provides a dual-sourced set of design requirements that integrates procedural needs of literature review practice with governance insights from PAT."
Risk of bias
Selection bias in choice of design scenarios; Prototypical evaluation limited to feasibility demonstration rather than empirical user validation; Limited to three specific open-source LLMs which may not represent broader model diversity; Prototype implementation does not include full-scale real-world deployment evaluation; Limited evaluation scope: prototypical rather than full-fledged implementation limits ability to assess real-world governance effects; Absence of empirical user validation: no user studies conducted with real researchers; Potential selection bias in LLM selection: only three open-source models tested (Mistral-7B-Instruct, Mixtral-8×7B-Instruct, Phi-3-Medium); Scenario-based evaluation rather than longitudinal testing may not capture long-term effects; Limited domain testing: framework evaluated primarily in Information Systems context; Limited to lightweight open-source LLMs; proprietary models not evaluated; Prototype evaluation through scenario-based reasoning rather than empirical user studies; Potential bias in PAT application to AI systems not originally designed for principal-agent contexts; Evaluation limited to IS research domain; generalizability to other fields unclear
Limitations
- The authors acknowledge several limitations: "First, as discussed in Section 2, PAT provides a useful lens for structuring human-AI delegation
- However, it does not explicitly theorize different degrees of autonomy within the agent role." Additionally, "the prototypical implementation serves as a feasibility demonstration of the proposed design principles rather than a fully developed software artifact
- The evaluation follows Hevner's descriptive evaluation strategy (Hevner et al., 2004), focusing primarily on internal coherence and practical plausibility
- While appropriate for early-stage design knowledge, this form of evaluation remains limited in its ability to assess usability and long-term governance effects in real-world settings." The study also identifies that "Access to bibliographical databases remains highly fragmented
- Key IS resources such as the AIS eLibrary offer no free API access, limiting the system's ability to conduct comprehensive searches and hindering open scientific knowledge dissemination."
Open questions raised
- Further theoretical refinement needed to fully capture the spectrum of AI agency and different degrees of autonomy within the agent role
- Need for full-fledged software artifact with comprehensive implementation of proposed framework
- Empirical validation with real users through user studies and longitudinal analyses required
- Domain-specific implementations needed
- Examination of whether PAT sufficiently captures differentiated and hybrid AI agency configurations
- Empirical investigation of real-world researcher-AI collaboration to inform theoretical extensions modeling graded autonomy and evolving delegation structures
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations