12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Exploring the Potential of GPT-2 for Generating Fake Reviews of Research Papers

Alberto Bartoli, Eric Medvet · Frontiers in artificial intelligence and applications · 2020

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
E
Evidence
8
Citations
1.38
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3233/faia200717

Methodology & findings

Study design

Mixed-methods proof-of-concept study combining: (1) Technical implementation of a fake review generation pipeline using GPT-2 fine-tuned on academic peer review corpus (PeerRead dataset + NIPS 2018-2019 reviews); (2) Small-scale user study (n=16 respondents) rating the credibility of fake vs.

Sample

N = 16, 4 groups

Primary method

Descriptive statistics only. Responses summarized as frequencies and percentages. No inferential statistics (no hypothesis tests, confidence intervals, or p-values reported). No statistical software specified.

Main result

The study found that "a simple heuristic based on widely available and easy to use tools may be remarkably effective" at generating fake peer reviews. In the user study, "approximately 75%, 30%, and 25% of the answers for the three papers, respectively, has satisfied this baseline" of being rated at least "Useful," and "for paper 1, the fake reviews were deemed more useful overall than the real one." The authors conclude that "academic frauds based on fake reviews may indeed be feasible and ready to be deployed in the wild."

Reports effect sizes.

Research paradigm

Empiricist/pragmatist (applied AI research with proof-of-concept validation)

Author conclusions

The authors conclude that "Modern tools for natural language processing and natural language generation may enable novel forms of scholarly fraud based on the automatic generation of fake review reports for academic papers" and that "Our analysis thus indicates that academic frauds based on fake reviews could indeed be feasible and ready to be deployed in the wild." They recommend that "journal publishers and conference organizing committees could occasionally inject fake reviews in the reviewing process, in a carefully controlled way, to make sure that these reviews are indeed spotted and discarded."

Risk of bias

Selection bias: Small, self-selected sample of academic contacts (n=16); Social desirability bias: Users informed study was about 'quality of peer reviews' without knowing some were fake; Rater bias: Assessment required subjective human judgment, particularly in step 5-(d) of review generation; Limited generalizability: Only 3 papers from authors' own research group used; not representative of diverse paper types; Familiarity bias: 50% of respondents had low familiarity with paper topics; Small sample size for user study (n=16); Selection bias: respondents were academic contacts of authors; Participant knowledge asymmetry: 50% had low familiarity with paper topics; Experimenter bias: authors selected snippets subjectively in step 5; Lack of blinding: procedure and evaluation both conducted by research group; Unequal distribution of expertise levels among raters (12 experienced, 3 junior, 1 student); Selection bias: Small volunteer sample of academic contacts (n=16), self-selected; Experimenter bias: Authors manually selected snippets from generated samples (step 5) based on subjective criteria including 'looks natural' requiring human judgment; Source bias: Three papers selected from authors' own research group, not representative; Awareness bias: Participants knew they were rating reviews for quality assessment but were unaware some were fake until debriefing; Domain specificity bias: All papers from machine learning/computer science; results may not generalize to other fields; Confounding variables: Participant expertise level (experienced/junior/student) and familiarity with topics (high/medium/low) not controlled

Limitations

  • The authors explicitly state that "Our small user study cannot certainly be conclusive" and acknowledge multiple constraints
  • They note "we are not able to perform this kind of assessment" of detecting fraud in real reviewing processes, and observe that "obtaining a long, coherent text from GPT-2 tends to be difficult," limiting review length
  • The small sample size (n=16 responses from 12 experienced researchers, 3 junior researchers, and 1 student) and self-selected participant pool introduce selection bias
  • The study uses only three papers from the authors' own research group, limiting generalizability.

Open questions raised

  • The authors identify that GPT-3 (175 billion parameters, two orders of magnitude larger than GPT-2) will enable even more powerful fake review generation. They note the difficulty of assessing whether frauds would be detected in real peer review processes and suggest testing fake reviews' impact on actual editors and program committees.
  • Testing in real peer review processes to assess actual detection by editors and authors
  • Evaluation using more powerful language models like GPT-3 (175 billion parameters)
  • More elaborate text generation methods (e.g., concatenating multiple snippets)
  • Systematic detection mechanisms for identifying auto-generated reviews
  • Need for larger-scale user studies with real reviewing processes to assess fraud detectability by editors and conference committees
Data: PeerRead dataset (publicly available), plus NIPS 2018-2019 reviews. User study results and 9 reviews available online at authors' laboratory website (https://machinelearning.inginf.units.it/); PeerRead dataset - publicly available peer reviews from machine learning conferences; User study results and 9 reviews (3 papers × 3 review types): available at laboratory website (https://machinelearning.inginf.units.it/); PeerRead dataset (publicly available): https://github.com/allenai/PeerRead - contains peer reviews from machine learning conferences; User study results and fake reviews: Available on authors' laboratory website (https://machinelearning.inginf.units.it/), referenced in 'Data and tools' section; Fine-tuned GPT-2 model and generated review examples: Availability implied but not explicitly confirmed in paperCode: Publicly available Colaboratory Notebook mentioned for fine-tuning and text generation; GitHub reference (https://github.com/allenai/PeerRead) for PeerRead dataset; Publicly available Colaboratory Notebook [8] used for fine-tuning and text generation (specific URL not provided in paper); Google Colaboratory Notebook for fine-tuning and text generation: publicly available (referenced as [8] but specific URL not provided in paper); GitHub reference implied: https://github.com/allenai/PeerRead (PeerRead dataset repository)Extracted from: pdfAgreement 50%

Explore related topics

Related papers