A new preprint from Samsung and the University of Warsaw lists Elena Vasquez and Marcus Chen as co-authors on hundreds of academic documents, yet these individuals have never lived.
The paper, titled “The Ghost Couple: Correlated LLM Name Priors and Their Haunting of the Web and Academic Publishing,” identifies these figures as names that large language models repeatedly produce when tasked with generating experts in specific fields.
Researchers found that certain models default to the same names in specific contexts. In June, 404 Media reported that ChatGPT, Gemini, and Claude were likely to use the name Elias Thorne for fiction characters, often a lighthouse keeper. Similarly, users noted that asking ChatGPT to generate a software developer frequently resulted in the name Marcus Chen.

The study showed that LLMs do not just default to single names but produce “correlated character ensembles,” meaning specific names are more likely to appear together. Other consistently generated names include Elena Amara Okafor from Claude, Aris Thorne and Lena Petrova from Gemini, and Elara Voss from ChatGPT.
Earlier this month, 404 Media reported on Research Gold, a company claiming to offer human-written medical research that was entirely AI generated. The founder and lead methodologist for that company was named Elena Vasquez. Research Gold removed the name from its website after the story was published.
Michał Brzozowskim, the lead author of the paper, told me that searching for these names on Google revealed other instances of AI-generated personalities. Following the killing of Alex Pretti by U.S. Border Patrol agents in January, a rumor spread on Facebook that he was fired from his nursing job for misconduct allegations. Snopes reported that this false statement was attributed to “executive director Dr. Elena Vasquez,” who does not exist.
Brzozowskim searched academic databases for the names they knew LLMs often generated.
The scale of the contamination
“On Zenodo, a CERN operated repository that mints real DataCite DOIs, we identify 1,655 ghost-authored records claiming nonexistent journals with fabricated publication dates,” Brzozowskim said.
A DOI, or Digital Object Identifier, is a string of characters and numbers used to identify academic papers. The researchers saw that many papers authored by these AI names were backdated, meaning their publication dates differed from the date they were uploaded to Zenodo.
Anyone with a free account can create a DOI on Zenodo, but the existence of AI-generated papers with DOIs has impacts on other parts of the web and academic publishing.
“These [AI generated papers] carry real DOIs harvestable by any scholarly aggregator; the infrastructure for large-scale scholarly record contamination is already in place. Ghost names additionally appear on ResearchGate, forming synthetic research groups with collaborators drawn from multiple model families, and are indexed without verification by Google Scholar and Semantic Scholar […] The academic record is being quietly haunted.”
ResearchGate and Google Scholar are both aggregators of academic publishing that are likely to come up in search results.
The researchers say that different LLMs and versions of those LLMs generate certain names so consistently, they think they can use them to determine the provenance of AI slop. For example, the name Elena Vasquez was particularly common in content generated by Claude Sonnet 4, so papers that list her as an author were likely generated by that LLM. However, Brzozowskim said that this level of accuracy might not hold for long now that AI-generated content with these names is flooding the internet and feeding back into all AI models that are scraping the internet for training data.
The fact that AI academic publishing is struggling to deal with the load of AI-generated content is not new. We have previously reported that scientific journals have published AI-generated text, that AI is impacting the peer review process, and that the open-access repository for preprint academic research Arxiv will now ban authors for a year if they are caught submitting AI-generated work. The potential upside of this research is that these names might allow us to detect this AI-generated content more easily.
What it means
For academics and researchers, this creates a persistent noise floor in search results. Finding a paper by a real author becomes harder when it is buried under records from people who do not exist. The presence of these fake DOIs also complicates citation metrics, as the academic record is polluted with entries that cannot be verified. The study offers a potential detection method, but the rapid spread of these names into training data suggests that future models may generate different fake names, requiring constant updates to verification tools.




![Program misleading high school students into paying to perform academic misconduct in ML Research [D]](https://ai-maestro.online/wp-content/uploads/2026/05/program-misleading-high-school-students-into-paying-to-perfo-768x768.jpg)