Research LLM
Guillaume Guérard, Sonia Djebali, Maxime Hanus, Mark-Killian Zinenberg · Tehnički glasnik · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.31803/tg-20250827131608
Methodology & findings
Study design
Qualitative case study with illustrative examples combined with an exploratory pilot study involving Master's-level students applying the three-stage framework to thesis proposal development.
Primary method
Design Science Research with human-in-the-loop methodology
Main result
The study found that "structured prompting improves traceability and broadens the set of considered alternatives, while verification steps curb overconfident errors." The pilot study with Master's students revealed that the framework provides valuable scaffolding for novice researchers, particularly in enhancing methodological literacy through structured comparisons of research designs, though it also highlighted challenges such as prompt engineering skill barriers and the 'credibility illusion' where fluent AI-generated text obscures methodological flaws.
Research paradigm
Pragmatist/Design Science
Author conclusions
The authors conclude that "the efficacy of AI in research hinges upon the non-negotiable role of human intellect. The framework is intentionally designed with critical 'Human Checkpoints' where the researcher's domain expertise, critical judgment, and ethical oversight are not just beneficial, but indispensable." They state: "AI-generated content, while often structurally coherent, lacks genuine comprehension and remains susceptible to factual inaccuracies and embedded biases. The AI can provide the scaffolding, but it is the human scholar who must act as the architect—validating claims, challenging assumptions, and infusing the work with the novel interpretation that constitutes a true contribution to knowledge." They further conclude that "the responsible integration of AI into research workflows presents a profound opportunity to deepen analysis, optimize efficiency, and expand the frontiers of inquiry" and that "the challenge for the academic community will be to co-evolve with them, refining collaborative frameworks that balance the automation of tasks with the amplification of intellect."
Risk of bias
Confirmation bias in AI outputs - AI may overemphasize dominant theoretical perspectives; Anchoring bias in novice researchers - students tend to over-rely on initial AI-generated outputs; Hallucination risk - LLMs can generate fabricated citations and data; Training data bias - LLMs trained on biased datasets may amplify societal biases; Selection bias in pilot study - small cohort of Master's students, not representative sample; Selection bias in pilot study: only Master's students in research-focused track were recruited; sample representativeness not specified; Observer bias: observational approach with researchers conducting the study likely familiar with their own framework; Lack of control group: no comparison with traditional research methods or students not using the framework; Confirmation bias: framework designed by authors may bias interpretation of pilot results; LLM training data bias: paper acknowledges 'LLMs are fundamentally shaped by the vast, static datasets on which they were trained, and these datasets inevitably reflect existing societal and historical biases'; Anchoring bias observed in pilot: 'students frequently demonstrated a reluctance to challenge or significantly refine the AI's first suggestions'; Attrition/incomplete reporting: pilot study findings are largely qualitative observations without systematic data collection methods described; Bias amplification from LLM training data reflecting societal biases; Anchoring bias in users over-relying on initial AI-generated outputs; Credibility illusion from fluent but potentially flawed AI text; Confirmation bias in researchers favoring supportive information; Measurement bias and sampling bias in case study example; Hallucination risks producing fabricated citations and references
Limitations
- The authors acknowledge that "there are currently no standardized Key Performance Indicators (KPIs) to objectively compare the quality of its outputs against that of traditional, human-only research
- While metrics of efficiency like time saved are easily measured, they fail to capture the core dimensions of scholarly value: the novelty of insights, the depth of critical analysis, or the serendipity of discovery." Additionally, "the framework's intrinsic dependence on researcher expertise" means that "the quality of the final output is inextricably linked to the user's ability to craft precise prompts and critically vet AI-generated content." Further limitations include the risk of "hallucination" where LLMs generate "outputs that are grammatically correct and stylistically plausible but factually inaccurate, logically flawed, or entirely nonsensical," and the challenge that "students often struggled to write effective prompts," suggesting prompt engineering itself is a skill barrier requiring training.
Open questions raised
- Lack of standardized KPIs for evaluating research quality produced with AI assistance
- Need for discipline-specific adaptations of the framework through cross-disciplinary case studies
- Absence of structured training materials for students on prompt engineering and critical AI evaluation
- Lack of robust metrics to assess research proposal quality (clarity, feasibility, novelty)
- Need for integrated software tools to guide researchers through the framework and facilitate transparent reporting
- Lack of standardized KPIs for objectively comparing AI-assisted research quality against traditional human-only research methods
Explore related topics
Related papers
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations
- ChatGPT: Bullshit spewer or the end of traditional assessments in higher education?Jürgen Rudolph · 2023 · 1,674 citations
- A comprehensive AI policy education framework for university teaching and learningCecilia Ka Yuk Chan · 2023 · 1,160 citations
- ChatGPT for Education and Research: Opportunities, Threats, and StrategiesMd. Mostafizer Rahman · 2023 · 904 citations
- ChatGPT and a new academic reality: Artificial Intelligence‐written research papers and the ethics of the large language models in scholarly publishingBrady Lund · 2023 · 769 citations
- ChatGPT in higher education: Considerations for academic integrity and student learningMiriam Sullivan · 2023 · 740 citations