🤖 AI Summary
This study reveals a systematic bias in large language models (LLMs) that generates fictional characters with stereotyped name pairings—termed “ghost author pairs” (e.g., Elena Vasquez and Marcus Chen)—which recurrently co-occur across documents and contaminate academic databases. The work demonstrates for the first time that this bias exhibits model-family specificity, version dependency, and traceable timestamp signatures. Through large-scale text mining, co-occurrence analysis, and metadata forensics—including DataCite timestamps—and cross-platform provenance tracing on Zenodo and ResearchGate, the authors identify 1,655 ghost author records, with 991 concentrated within a single month of registration. Critically, these fabricated publications have been assigned valid DOIs and indexed in scholarly databases, providing a reliable proxy indicator for LLM deployment timelines.
📝 Abstract
These names do not exist. Elena Vasquez and Marcus Chen have appeared as volcano experts, astronauts, thriller protagonists, podcast hosts, and academic co-authors across hundreds of independently produced AI-generated documents, never having lived. We show that large language models do not merely default to high-probability individual names when generating fictional experts: they produce correlated character ensembles, pairs and trios whose co-occurrence rates far exceed chance and are consistent across independent generations. These priors are model-family-specific (Claude: Elena Vasquez + Marcus Chen + Amara Okafor; Gemini: Aris Thorne + Lena Petrova; GPT: Elara Voss with no fixed partner), version-specific, and actively suppressed at model release boundaries, leaving dateable behavioral fingerprints in the content they produced. We document a downstream consequence at scale. On Zenodo, a CERN-operated repository that mints real DataCite DOIs, we identify 1,655 ghost-authored records claiming nonexistent journals with fabricated publication dates: server-side DataCite timestamps prove deliberate backdating, and 991 records were registered in a single month; these carry real DOIs registered in DataCite, making them harvestable by any scholarly aggregator that ingests DOI metadata. Ghost names additionally appear on ResearchGate forming synthetic research groups with collaborators drawn from multiple model families; publication dates on these records provide a reliable temporal proxy for model deployment windows.