PEER: Empowering Writing with Large Language Models
Kathrin Seßler, Tao Xiang, Lukas Bogenrieder, Enkelejda Kasneci · Lecture notes in computer science · 2023
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/978-3-031-42682-7_73
Methodology & findings
Study design
Case study with prototype development and preliminary user feedback collection.
Sample
N = 4000, 2 groups
Primary method
Elo rating system for comparative evaluation of prompts; weighted lottery system for prompt selection based on ratings. No traditional statistical hypothesis testing reported.
Main result
The study found that "over 4000 essays have already been uploaded for evaluation, with argumentation being the most frequently requested category" and that "based on the Elo scores, prompting the model to approach the task as a friendly teacher and providing it with additional information about the specific essay type leads to the best results." Teachers acknowledged "PEER's usefulness for both students and teachers, highlighting its user-friendliness, respectful tone, and timely feedback that facilitates individualized learning."
Reports effect sizes.
Research paradigm
pragmatist/design-oriented
Author conclusions
The authors conclude that "Based on the amount of feedback we have received from teachers so far, it is clear that PEER is meeting a need in schools for both learners and teachers" and that their "objective is to bring PEER to schools and establish it as a valuable assistant in the process of learning how to write an essay." They emphasize that users should "carefully evaluate the feedback provided and selectively incorporate only the pertinent critiques" and that "critical thinking is not limited to PEER but should be a fundamental aspect of interacting with any generative AI model."
Risk of bias
Selection bias: Platform users may not be representative of all students; Feedback bias: Teachers providing qualitative feedback were self-selected volunteers; Measurement bias: User preferences for prompt quality may be influenced by interface presentation; Language bias: Tool designed specifically for German language and educational context; Selection bias: Users self-selected to use the platform; no random assignment; Lack of control group: No comparison with traditional teacher feedback or alternative systems; User feedback bias: Elo rating system based on user preferences may not reflect actual quality; Teacher feedback bias: Qualitative feedback from teachers may be influenced by positive initial impressions; Sampling bias: Platform usage skewed toward middle and upper level students; No systematic evaluation methodology: Preliminary results lack rigorous measurement; Selection bias: Users self-selected to upload essays to the platform; Lack of randomization: No control group for comparison; Potential response bias: Teachers who provided feedback may have been more favorably disposed toward the tool; Limited demographic diversity: Platform primarily used by students from middle and upper school levels
Limitations
- The authors acknowledge that "Given the inherent stochastic nature of the underlying model, it is important to acknowledge that a fully error-free outcome cannot be guaranteed." Additionally, they note that "the feedback provided by PEER can be too general at times, such as suggesting to 'use more adjectives'" and that the tool "currently does not account for both versions in its output text" regarding gender-inclusive language
- The authors also state "our project is still in its early stages and requires further development, model fine-tuning and user evaluations."
Open questions raised
- Need for comprehensive user study at schools
- Model fine-tuning using real-world essay data and teacher feedback
- Expansion to other languages and educational systems beyond German
- Improvement of feedback specificity and cultural sensitivity (e.g., gender-inclusive language in German)
- Addition of missing essay types to platform
- Grammar error reduction in generated feedback
Explore related topics
Related papers
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations
- ChatGPT: Bullshit spewer or the end of traditional assessments in higher education?Jürgen Rudolph · 2023 · 1,674 citations
- ChatGPT for Education and Research: Opportunities, Threats, and StrategiesMd. Mostafizer Rahman · 2023 · 904 citations
- ChatGPT and a new academic reality: Artificial Intelligence‐written research papers and the ethics of the large language models in scholarly publishingBrady Lund · 2023 · 769 citations
- ChatGPT in higher education: Considerations for academic integrity and student learningMiriam Sullivan · 2023 · 740 citations
- Practical and ethical challenges of large language models in education: A systematic scoping reviewLixiang Yan · 2023 · 699 citations