Thinking Through Signs: PEEL as a Semiotic Scaffolding for Epistemically Accountable AI-Enabled Research
Clarisse de Souza, Gabriel Barbosa, Simone Diniz Junqueira Barbosa, Barbara BETTS, Renato Cerqueira, Juliana Jansen Ferreira · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Proof-of-concept case study using three scholarly texts (Alvarado 2023, Ferrario 2024, Boisseau 2026).
Sample
N = 3, 4 groups
Primary method
Voyant Tools provides deterministic natural language processing including tokenization, term frequency analysis (Raw Frequency, Relative Frequency, Relative Peakedness, Relative Skewness, Distribution), type-to-token ratio measurement, and visualization tools. Comparison of AI outputs measured against source material using document-level metrics (size compliance, term ratio preservation). No inferential statistics reported.
Main result
The study found that "The three summarization modes differ substantially in the epistemic resources they offer. The free summarizations (AI-1 and AI-2) produce fluent, broadly accurate accounts, but neither indicates which sentences are verbatim, which arguments have been compressed, or which examples were dropped." Additionally, "Claude's outputs ranged from 24.9% to 29.7%, remaining in the right order of magnitude and erring upward rather than downward," while "AI-1 and AI-2 produced condensations of approximately 12% of the original—less than half of what was requested." The research also revealed that "in Boisseau (2026), whose argument turns on the relationship between trust and epistemic responsibility, the source ratio of responsib* to trust* is 23:132, approximately 17%. Claude's condensation tracks this closely at 16%. AI-1 produces a ratio of 60%; AI-2, 31%."
Reports effect sizes.
Research paradigm
Interpretivist/constructivist with semiotics and abductive reasoning
Author conclusions
"This commentary invites the IS community to consider how non-AI technologies, combined with AI, can address critical challenges of interpretive, text-based research. We have shown how signs externalized by Claude and Voyant seed abductive reasoning, helping researchers generate creative hypotheses about text sources even before reading them. The process accelerates interpretive research while maintaining the trace of connections and provenance necessary -though not sufficient -to warrant epistemic authority. The researcher drives the process, all the time." The authors conclude that "Three design implications follow: deterministic instruments must accompany AI tools; fluency is not fidelity; and epistemic authority must be designed, not assumed."
Risk of bias
Selection bias in text choice: The three papers selected were from the authors' own bibliography, not randomly chosen, though authors note they were 'not arbitrary choices' as they supported the commentary's argument; Reflexivity bias: Authors disclose that 'as part of this study's design, none of the authors had read these papers in full when the Voyant analyses were conducted,' creating intentional reflexivity but potentially limiting contextual interpretation; Researcher approval bias: The checkpoint protocol in Phase 2 depends on researcher judgment to approve condensations, which the authors acknowledge may 'favor completeness over brevity'; Tool-selection bias: Comparison involves only three AI systems (two anonymized as AI-1 and AI-2, plus Claude), limiting generalizability; Instruction interpretation: The simple instruction to 'produce a summary 25% of original size with 5% tolerance' may have been interpreted differently by different systems; Researcher prior knowledge bias: Authors acknowledged they had not read the source papers in full before analysis, potentially biasing interpretation of PEEL outputs; Selection bias: Only three texts were analyzed, all from the authors' own bibliography; Anonymization inconsistency: AI-1 and AI-2 were kept anonymous while Claude was named, potentially biasing comparative evaluation; Investigator bias: Authors designed the skill used to guide Claude, potentially creating confirmation bias in favor of Claude's performance; Selection bias: Three specific texts chosen from authors' own bibliography; Observer bias: Authors' familiarity with PEEL system and Claude may influence interpretation; Timing bias: Vouant analyses conducted before full reading of source material; Measurement bias: Researchers explicitly approved text condensations at checkpoint, potentially influencing compression rates; Anonymization bias: AI-1 and AI-2 kept anonymous while Claude identified, potentially affecting interpretation
Limitations
- The authors state that "Larger numbers of texts can be analyzed and compared, although further investigation is required regarding the scalability of this model." They also note: "The data cannot disambiguate between them [four competing hypotheses about AI-1 and AI-2 deviations]
- PEEL makes the anomaly visible and the hypotheses formulable
- Without Voyant, there would be nothing to explain." Additionally, they acknowledge "One possibility is that the clusters-to-categories mapping approved by the researcher is not good" and raise an open question: "is greater vocabulary diversity in a summary a door open to the quiet injection of concepts not present in the source? Careful comparative analysis will tell." Finally, they state: "This is only the beginning" and identify remaining work: "Much remains ahead—scalability, software architecture to lower the costs of tasks deterministic computing can resolve, structured artifacts for research transparency, and the long-term consequences of PEEL in scholarly work."
Open questions raised
- Scalability of PEEL to larger numbers of texts and diverse domains
- Software architecture to lower computational costs of deterministic analysis tasks
- Structured artifacts for ensuring research transparency in AI-enabled research workflows
- Long-term consequences and efficacy of PEEL in actual scholarly practice
- Whether greater vocabulary diversity in summaries represents silent injection of source-external concepts
- Systematic patterns in how LLMs process semantic fields, particularly conflation of 'reliability' with 'trustworthiness'
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- A SWOT analysis of ChatGPT: Implications for educational practice and researchMohammadreza Farrokhnia · 2023 · 1,171 citations
- ChatGPT and a new academic reality: Artificial Intelligence‐written research papers and the ethics of the large language models in scholarly publishingBrady Lund · 2023 · 769 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations