12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

GenAI-supported portfolio assessment for complex thinking: a GPT-based innovation in business education

May Portuguez Castro, Isolda Margarita Castillo-Martínez · Frontiers in Education · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

6/10
Relevance
0/4
Quality (LMQS)
C
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3389/feduc.2026.1729156

Methodology & findings

Study design

Qualitative exploratory-documentary approach complemented by correlational statistical analysis.

Main result

The study found that "the GPT-eComplex Assistant closely mirrored human evaluators' judgments, reinforcing consistency, traceability, and transparency, although still requiring teacher calibration" and revealed that "the digital portfolio served as an authentic learning artifact to capture systemic and critical thinking, while showing limitations in the scientific dimension." Overall mean scores for complex thinking across 120 portfolios indicated "an overall mean score of 2.40, corresponding to an intermediate–advanced level, with a clear tendency toward higher performance in the innovative (2.66) and systemic (2.56) dimensions, compared to lower values in the critical (2.31) and scientific (2.06) dimensions."

Research paradigm

Mixed methods (qualitative exploratory-documentary with correlational statistical validation)

Author conclusions

The authors conclude that "the responsible integration of GenAI demands active teacher mediation, ethical awareness, and institutional transparency, ensuring that technology complements rather than replaces pedagogical judgment." They further state that "this study contributes a pedagogical strategy that combines formative portfolio assessment, AI-supported co-evaluation, and the AI-PROMPT Framework, offering a replicable model for embedding GenAI into authentic, reflective, and ethically grounded assessment practices in higher education." They emphasize that "the results represent an innovative contribution by positioning GPT as a complementary tool in authentic assessment, reinforcing the central role of human judgment and opening new perspectives for AI-supported evaluation in the Ibero-American and beyond."

Risk of bias

Selection bias: Pilot subsample (n=12) used for agreement analysis may not be representative of full sample (n=120); Hawthorne effect: Students knew portfolios would be assessed, potentially influencing quality; Single institution: Research conducted at one graduate business school in Peru may limit generalizability to other contexts; Single country context: Ibero-American focus may not transfer to other educational systems; Teacher calibration dependency: Results require human teacher involvement for validity, reducing standardization; Model-specific effects: Findings tied to ChatGPT/GPT-4 may not generalize to other LLMs; Selection bias: Pilot subsample (n=12) selected from corpus of 120 portfolios may not be representative; Attrition/sampling: Only 10% sample used for correlation analysis, non-random selection process not fully described; Evaluator bias: Two instructors conducting initial human evaluations; inter-rater reliability not explicitly reported for human evaluators; AI bias: GPT-eComplex Assistant configuration may embed biases from training data; prompt engineering decisions influence outputs; Context bias: Study conducted at single institution with specific curricular focus (Positive Business Impact Model); findings may not generalize; Single cohort: Portfolios from one graduate business school in Peru; regional and institutional specificity limits generalizability; Selection bias: pilot subsample of 12 portfolios selected from 120 may not be representative; Single institution context (one graduate business school in Peru) limits generalizability; Potential rater bias in human evaluator judgments used for comparison; AI model (ChatGPT/GPT-4) may encode biases from training data; No blinding of evaluators to treatment condition reported

Limitations

  • The authors state that "the use of GenAI for assessment in higher education remains incipient and under-researched, underscoring the need for empirical evidence and ethical guidelines." Additionally, they note that "the digital portfolio served as an authentic learning artifact to capture systemic and critical thinking, while showing limitations in the scientific dimension." The study acknowledges that "complete agreement was not achieved, indicating that the agent should be regarded as a supportive analytical tool rather than a replacement for human judgment." Furthermore, "the correlational analysis was conducted on a pilot subsample (n=12) was used for the correlational analysis comparing human and AI-assisted assessments, while descriptive analyses of complex thinking levels were conducted on the full sample (n=120)," suggesting limited generalizability of the agreement findings.

Open questions raised

  • Empirical evidence on GenAI applications in authentic and formative assessment in Ibero-American higher education remains scarce
  • Limited research on reliability, consistency, and pedagogical validity of AI-assisted qualitative evaluation
  • Few studies have systematically examined the use of GenAI in assessing complex learning processes
  • Need for more empirical evidence and practical guidance on pedagogical application of GenAI in assessment
  • Opportunities to strengthen scientific thinking skills through portfolio design and data-based analysis integration
  • Need for further research on optimizing calibration, feedback, and ethical integration of AI in complex thinking assessment
Extracted from: pdfAgreement 68%

Explore related topics

Related papers