12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Differences in User Perception of Artificial Intelligence-Driven Chatbots and Traditional Tools in Qualitative Data Analysis

Boštjan Šumak, Maja Pušnik, Ines Kožuh, Andrej Šorgo, Saša Brdnik · Applied Sciences · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
E
Evidence
3
Citations
1.45
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3390/app15020631

Methodology & findings

Study design

Empirical laboratory experiment with within-subjects design.

Sample

N = 85, 5 groups

Primary method

Kruskal-Wallis H test (non-parametric ANOVA) for group comparisons; Dunn's post-hoc test for pairwise comparisons; Cronbach's alpha for internal consistency/reliability; Kendall's tau-b correlation for non-parametric correlation between SEQ and NASA-RTLX; Point-Biserial correlation for error rates; Pearson Chi-Square for association analysis; Mann-Whitney U test for between-group emotion comparisons; effect sizes calculated using eta-squared (η²). Software: Not explicitly stated.

Main result

The study found that "ChatGPT obtained the highest SUS score (SUS = 79.03)" compared to Taguette (74.95) and Gemini (75.08), with "all usability scores are acceptable and can be categorized as good." More significantly, "the difference in the UEQ Scale scores between tools was analyzed with a Kruskall-Wallis H test. The results showed a significant difference in overall UEQ score, χ 2 (2) = 102.888, p < 0.001, with a very large effect size (η 2 = 0.4186)," indicating that "AI-based tools, particularly ChatGPT, consistently foster more positive user experiences, reduce negative emotional impacts, and lower cognitive workload compared to traditional tools like Taguette."

Reports effect sizes and confidence intervals.

Research paradigm

positivist/empiricist

Author conclusions

"AI-based tools, particularly ChatGPT, consistently foster more positive user experiences, reduce negative emotional impacts, and lower cognitive workload compared to traditional tools like Taguette." The authors emphasize that "Tools like ChatGPT, Gemini, and Taguette each present unique strengths and weaknesses in these areas, highlighting the importance of tailoring tools to users' needs," and note that "These insights provide valuable guidance for the design and implementation of effective QDA tools," though they caution that "It is important to note that this study involved a limited group of respondents, primarily Master's students in informatics and data technologies, which may affect the generalizability of these findings."

Risk of bias

Selection bias: Participants were recruited from a single university program (informatics and data technologies); Self-report bias: All measures were self-reported via questionnaires; Familiarity bias: Participants had high prior familiarity with ChatGPT but low familiarity with Gemini, potentially influencing responses; Learning effects: Different datasets used for each tool may have influenced comparability; Order effects: Taguette was always tested first, potentially affecting perception of subsequent tools; Selection bias: Participants were Master's students with prior UX evaluation experience, not representative of broader user populations; Self-report bias: All measurements based on self-reported scales, not objective behavioral data; Familiarity bias: Participants had higher familiarity with ChatGPT (M=4.62) than Gemini (M=2.65), potentially influencing trust and perception scores; Task assignment bias: Different datasets assigned to each tool sequentially, though authors attempted to mitigate with fresh datasets; Order effects: Sequential tool evaluation may create learning or fatigue effects; Response validation: Only partial validation of user responses (25% error rate in ChatGPT Task 2 responses); Selection bias: Participants were Master's students in informatics with above-average technical familiarity, not representative of broader user populations; Familiarity bias: Participants had prior experience with ChatGPT (mean 4.62/5 familiarity) but minimal Gemini exposure (2.65/5), introducing confounding by tool familiarity; Self-report bias: All measurements based on self-reported perception scales rather than objective behavioral measures; Task design bias: Different datasets assigned to each tool (content.txt, ux.txt, structure.txt) limits direct comparison; Learning effect: Second task showed different patterns than first task, suggesting cumulative experience effects; Incomplete data validation: Authors note 'User responses were only partly validated, raising concerns about the reliability and accuracy of the feedback received'; Attrition/incomplete responses: For task 2 sentiment analysis, response rates differed (Taguette 80 valid, ChatGPT 64, Gemini 58) and 25% of ChatGPT responses contained categorization errors

Limitations

  • "The participant sample consisted primarily of Master's students in informatics and data technologies, which limits the generalizability of the findings to other user groups, such as secondary school students, PhD candidates, or professional staff." Additionally, "The evaluations relied on self-reported scales, reflecting participants' perceptions rather than objective behavior or cognitive process measures
  • It is important to note that reported trust levels may not directly correspond to trust exhibited in decision-making, and perceptions of workload or difficulty may diverge from actual task performance or cognitive load." Furthermore, "This study did not establish a baseline comparison of the task difficulty without the use of tools, which could provide a critical context for evaluating the effectiveness of the tools."

Open questions raised

  • Comparative studies between tasks performed by chatbots and 'traditional digital tools' are limited and require additional research
  • Very limited number of studies specifically analyzing usability of chatbots for qualitative data analysis
  • Need for broader and more representative user groups in future research
  • Lack of research on establishing baseline comparison of task difficulty without tools
  • Need for objective data such as task completion times to complement self-reported measures
  • Research needed on how training and prior tool familiarity shape user experiences
Data: Qualitative survey responses from 100 users (students and faculty) about Faculty of Electrical Engineering, Computer Science and Informatics website evaluation - available as three text files: content.txt, ux.txt, and structure.txt (each containing 100 user responses); Qualitative data from user evaluations of Faculty of Electrical Engineering, Computer Science and Informatics website (content.txt, ux.txt, structure.txt containing 100 user responses each); Website evaluated: https://feri.um.si/en/; Qualitative data from preliminary user study of Faculty of Electrical Engineering, Computer Science and Informatics website (University of Maribor): Three datasets (content.txt, ux.txt, structure.txt) each containing 100 user responses from online survey of website https://feri.um.si/en/Extracted from: pdfAgreement 50%

Explore related topics

Related papers