The Ghost Couple: Correlated LLM Name Priors and Their Haunting of the Web and Academic Publishing
Michał Brzozowski, Neo Christopher Chung · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Multi-method forensic approach combining: (1) systematic API probing of nine Claude versions, ten GPT versions, and Gemini-2.5-flash with controlled prompts (30 single-name, 30 pair, 30 trio prompts at temperature 1.0); (2) web corpus collection via Serper.dev Google Search API across multiple archetypes; (3) academic infrastructure analysis via ResearchGate and Zenodo API queries; (4) manual verification of 436 ResearchGate publication records and 1,655 Zenodo records; (5) temporal analysis using citation_publication_date meta tags and DataCite server-side timestamps; (6) control corpus design using demographically matched non-ghost surnames..
Primary method
Predict-then-confirm methodology: controlled API probing establishes model-specific name priors, which are then used as search signatures to recover AI-generated content in the wild, converting controlled experiments into detection tools. Comparative cross-model analysis (Claude, GPT, Gemini).
Main result
The study demonstrates that "large language models do not merely default to high-probability individual names when generating fictional experts: they produce correlated character ensembles: pairs and trios whose co-occurrence rates far exceed chance and are consistent across independent generations." Additionally, "on Zenodo, a CERN-operated repository that mints real DataCite DOIs, we identify 1,655 ghost-authored records claiming nonexistent journals with fabricated publication dates: server-side DataCite timestamps prove deliberate backdating, and 991 records were registered in a single month." The authors found that Elena Vasquez appears at "67% in claude-sonnet-4-20250514, decaying monotonically" across model versions, and "the pair is fully suppressed in claude-sonnet-4-6."
Research paradigm
Empirical forensic analysis with qualitative and quantitative components
Author conclusions
The authors conclude: "We have shown that LLMs generate correlated character ensembles, not merely high-probability individual names, that are model-family-specific, version-specific, and actively suppressed at release boundaries; the suppression is itself evidence that the priors were strong enough to be noticed. These ghost names propagate from model outputs into AI-generated web content and from there into academic publishing infrastructure. On Zenodo alone, 1,655 ghost-authored records with real DataCite DOIs were registered in a 60-day automated burst, claiming nonexistent journals with backdated publication dates; the infrastructure for large-scale scholarly record contamination is already in place. The academic record is being quietly haunted."
Risk of bias
Selection bias: API probing limited to publicly accessible checkpoints only; Measurement bias: Google Search API subject to recency bias in indexing; Classification bias: Potential misidentification of real researchers whose names coincidentally match ghost names; Sampling bias: Prompt set size (n=30) may miss lower-frequency name priors; Temporal bias: Web page publication dates unreliable for slop sites; only ResearchGate dates trustworthy; API-only probing misses internal model behavior and fine-tuned variants; Prompt set size (30 per condition) may miss lower-frequency name priors; Google Search recency bias in web corpus collection; Manual verification introduces human annotation bias; Possible false positives: real researchers whose names coincidentally match ghost priors; Temporal analysis relies on user-controllable metadata (Zenodo publication_date field); Selection bias: ghost names that are more distinctive may be more discoverable in web searches; Recency bias in Google Search indexing affecting web corpus collection; Unreliability of page-level publication dates on AI-generated content sites; Incomplete coverage of internal/fine-tuned models not accessible via public APIs; Potential false positives from legitimate researchers sharing names with ghost priors; Selection bias in manual review of ResearchGate and Zenodo records; Temporal lag between model release and appearance of ghost content in wild
Limitations
- The authors state that "Our probing study covers only publicly accessible API checkpoints
- internal or fine-tuned models are not covered
- Prompt set size (30 prompts per condition) is sufficient to establish dominant priors but may miss lower-frequency names
- Web corpus collection via Google Search (Serper) is subject to recency bias in the age field
- page-level publication dates from slop sites are unreliable
- ResearchGate paper dates are more trustworthy but require systematic collection at scale, which is ongoing
Open questions raised
- The authors identify the need for: (1) systematic collection of ResearchGate data at scale; (2) investigation of internal or fine-tuned models beyond publicly accessible APIs; (3) larger prompt set sizes to detect lower-frequency name priors; (4) real-time monitoring infrastructure as fake-paper generators continue deployment.
- The authors identify: (1) Internal or fine-tuned models not covered by API probing; (2) Need for systematic ResearchGate collection at scale; (3) Downstream impact on scholarly aggregators (Semantic Scholar, Google Scholar) not yet fully indexed but infrastructure in place; (4) Understanding of why ensemble crystallization occurs (trio > pair > solo structure may reflect narrative fine-tuning data volume/structure); (5) Need to date LLM-generated content using publication date proxies across platforms.
- The paper identifies the need for systematic collection of ResearchGate data at scale and proposes publication dates on fake academic records as a temporal proxy for model deployment windows. It notes that ghost-authored papers cite real work but real papers' 'cited by' lists do not yet include them, suggesting future research into citation network contamination.
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- ChatGPT in higher education: Considerations for academic integrity and student learningMiriam Sullivan · 2023 · 740 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations