What Is a Person? Emerging Interpretations of AI Authorship and Attribution
Heather Lea Moulaison · Proceedings of the Association for Information Science and Technology · 2023
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1002/pra2.788
Methodology & findings
Study design
Mixed-method content analysis combining: (1) a literature review of published resources on ChatGPT authorship, attribution, copyright, and accuracy issues; (2) a weekly Google search-based web scraping study conducted over 6 weeks (March 14 – April 18, 2023) using the search query 'chatgpt apa citation site:.edu' with three pages of results captured each week (88 total webpages); (3) manual coding in MAXQDA of webpage content for presence/accuracy of APA citation guidance; (4) institutional classification of webpages using Carnegie Classification system and AAU membership status..
Main result
The study found that "librarians were quick to provide guidance, but slow to update that guidance, contributing to the potential for misunderstanding the affordances of and best practices for work with LLMs." Over the 6-week study period, the percentage of top-ranked webpages recommending outdated approaches (citing ChatGPT as personal communication or electronic source) remained high, with "over half of the webpages on the first page of Google hits continued to gave out-of-date recommendations" even after official APA guidance was published on April 7, 2023. Additionally, "40% of library sites engaged with ChatGPT to identify best practices for citing it in APA," demonstrating both sophistication and fundamental misunderstanding of LLM capabilities.
Research paradigm
Interpretive/Hermeneutic
Author conclusions
"Government bodies have led the way in formally designating the writing of LLMs as being the product of algorithms that predict text based on a corpus, and as such are not creating text, nor as machines do they have the intellectual capacity to do so. Publishers have followed this lead, indicating that LLMs are incapable of the intellectual effort required for authorship of an article, for example, and in many cases should not be cited due to the problem of inaccuracies that are introduced. Education, however, shows evidence of being split, with some looking to integrate the tools of the future, and with others focusing on the problematic nature of the content produced, including the hallucinations that stand in the way of accuracy." The author notes that information professionals have an obligation to understand the basics of LLM technology before providing public guidance, and that while not all guidance was accurate, the provision of some guidance is preferable to none.
Risk of bias
Selection bias: study limited to .edu domain institutions in the US only; Selection bias: reliance on Google search algorithm rankings, which may favor certain institutional types; Temporal bias: study period coincides with rapidly evolving guidance, limiting generalizability; Measurement bias: manual coding of webpages by single researcher(s) increases risk of subjective interpretation; Survivorship bias: only captured webpages that remained accessible and ranked during the 6-week period; Selection bias: Results depend on Google's ranking algorithm, which may favor certain institutions or webpage types; Sampling bias: Only .edu domain searches in the United States; excludes international or non-institutional sources; Temporal selection bias: Study period (March 14 – April 18, 2023) captures only a narrow window during rapid policy changes; Coder bias: Manual coding by researcher(s) without reported inter-coder reliability testing or blinding; Language bias: Search string and analysis limited to English-language content; Publication bias: Only top three pages of Google results analyzed; lower-ranked pages excluded; Selection bias: Study limited to .edu domain and Google search results only; Search string bias: Single search query may not capture all library guidance; Manual coding bias: Potential for inconsistent coding across 88 webpages; Temporal bias: Search conducted during transitional period when guidance was rapidly changing; Language/platform bias: Focus on English-language webpages in US higher education
Limitations
- The authors state: "Limitations to this project include the problem of relying on Google for ranked results, and using a single search string that focuses on 'APA citation' within academic libraries in the US
- Because Google is the dominant search engine, this problem is nonetheless likely replicated by scholars, especially junior scholars, seeking support for their own work citing ChatGPT
- Other limitations include the problem of manually coding a large set of documents, though the use of MAXQDA is intended to help mitigate this to an extent
- Finally, this project does not break out webpage content by class of institution in the results
- It assumes that practice is concerned with providing accurate and timely information on citing ChatGPT in APA
- Some libraries supporting junior scholars, however, might reasonably have priorities that focus on other aspects of the problem, such as working with instructors to assess the applicability of ChatGPT texts to the assignment."
Open questions raised
- Limited understanding of how LLMs are used in scholarly writing contexts
- Need for more rigorous, peer-reviewed literature on ChatGPT authorship (most available literature was un-peer-reviewed at time of study)
- Unclear why large, elite research institutions were underrepresented in Google rankings for citation guidance
- Missing research on how different institutional types prioritize guidance development (e.g., accuracy vs. other pedagogical concerns)
- Need for understanding of perceptions of authorship and ownership as they pertain to generative AI across different user communities
- Limited peer-reviewed literature on ChatGPT and authorship at the time of study (identified as gap the paper aimed to fill)
Explore related topics
Related papers
- Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statementDavid Moher · 2009 · 83,271 citations
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviewsMatthew J. Page · 2021 · 10,956 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations