12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Evaluating ChatGPT’s Cognitive Performance in Chemical Engineering Education

Salman Shahid, Shaun Walmsley · Information · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

6/10
Relevance
0/4
Quality (LMQS)
E
Evidence
1
Citations
8.40
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3390/info17020162

Methodology & findings

Study design

Systematic evaluation using a diverse dataset of undergraduate-level chemical engineering problems mapped to Bloom's Taxonomy.

Main result

The study found that "Results show significant differences in ChatGPT performance across Bloom levels, revealing three distinct tiers of capability. Strong performance was observed at lower cognitive levels (Remember–Apply), while substantial degradation occurred at Analyze, Evaluate, and specially Create." This demonstrates that ChatGPT's competency varies significantly depending on the cognitive complexity level required by Chemical Engineering problems.

Reports effect sizes.

Research paradigm

Empiricist/Positivist

Author conclusions

The authors conclude that "The findings provide a nuanced, empirically grounded understanding of current LLM capability limits, with practical recommendations for educators integrating LLMs into engineering curricula." This suggests that the study offers educators concrete guidance on when and how to appropriately integrate ChatGPT into Chemical Engineering education based on cognitive domain requirements.

Risk of bias

Dataset representativeness: Limited to undergraduate-level chemical engineering problems; generalizability to other engineering disciplines unclear; Model version specificity: Results tied to a specific ChatGPT version; performance may vary with model updates; Evaluator bias: Single-model evaluation without comparison to other LLMs; Selection bias: Dataset composition and problem selection criteria not specified in abstract; Potential selection bias in dataset construction; Evaluator bias in categorizing error types and assessing accuracy; Model-specific bias (only ChatGPT evaluated, not comparative across LLMs); Temporal bias (LLM capabilities may change with version updates); Model selection bias: only ChatGPT evaluated, not other LLMs; Dataset composition: only undergraduate-level Chemical Engineering problems; Evaluator bias: not specified whether evaluation was blinded or inter-rater reliability assessed

Limitations

  • The authors note that "Despite widespread informal use, few empirical studies have evaluated LLM performance using a systematically designed dataset mapped directly to Bloom's Taxonomy," indicating that their work addresses a gap but also implying limitations in the broader empirical foundation for LLM evaluation in engineering education.

Open questions raised

  • The authors identify the need for empirical evaluation of LLMs in engineering education, noting that "Despite widespread informal use, few empirical studies have evaluated LLM performance using a systematically designed dataset mapped directly to Bloom's Taxonomy." Future research should examine integration strategies and domain-specific performance across other engineering disciplines.
  • The abstract indicates a gap in empirical studies evaluating LLM performance using systematically designed datasets mapped to Bloom's Taxonomy in chemical engineering education. The authors implicitly call for further research on best practices for integrating LLMs into engineering curricula.
  • The paper identifies the gap that "Despite widespread informal use, few empirical studies have evaluated LLM performance using a systematically designed dataset mapped directly to Bloom's Taxonomy." Future research directions would include evaluation of other LLMs, graduate-level problems, and domain-specific applications beyond Chemical Engineering.
Data: not_statedCode: not_statedExtracted from: pdfAgreement 55%

Explore related topics

Related papers