The Crisis of AI Slop: Why Research Papers with Fake Authors and Hallucinations are Flooding Academic Conferences
A recent report from reviewers operating in the so-called "slop trenches" has revealed a disturbing trend in academic publishing: approximately 68% of paper submissions across major AI conferences, including NeurIPS and WACV, contain fabricated citations, fake authors, or LLM-generated content. Despite these significant red flags, the current peer-review system is failing to filter out these errors, with some flagged papers even being accepted for prestigious oral presentations. This analysis explores the data behind the "AI slop" epidemic, citing audits that estimate over 146,000 hallucinated citations in 2025 alone. The findings suggest a systemic failure in the review process, where the vast majority of AI-generated hallucinations persist from preprint to final publication, posing a severe threat to the integrity of scientific discourse and the reliability of future research.
Key Takeaways
- High Prevalence of AI Slop: Reviewers found that 68% of paper submissions (15 out of 22) across venues like NeurIPS and WACV contained fabricated citations, fake authors, or LLM-generated technical jargon.
- Systemic Review Failure: Despite being flagged for fraudulent authorship, some papers were still accepted as oral presentations, and 85.3% of hallucinated references in preprints remain in the final published versions.
- Massive Scale of Hallucinations: Audits estimate roughly 146,900 hallucinated citations in 2025 publications, with early-career researchers and small teams identified as the most likely contributors.
- Ineffective Gatekeeping: Current peer-review processes are largely failing to catch AI-generated "slop," leading to a waste of time for reviewers and a potential erosion of scientific integrity.
In-Depth Analysis
The Reality of the "Slop Trenches"
The term "slop trenches" has emerged to describe the grueling task of reviewing academic papers in an era dominated by "AI slop cannons." In a recent audit of 22 paper submissions across NeurIPS (Datasets and Benchmarks and Position Paper tracks), WACV, and TerraBytes (an ECCV workshop), reviewers Caleb and Isaac identified a staggering rate of fabrication. Out of their combined 22 assignments, 15 papers (68%) were found to contain entirely fabricated citations, fabricated author lists for existing papers, or writing that was unmistakably generated by Large Language Models (LLMs).
The breakdown of these findings highlights the severity of the issue across different venues. At WACV, 6 out of 8 reviewed papers (75%) were identified as containing slop. In the NeurIPS Position Paper track, 100% of the assigned papers (2 of 2) were flagged. The reviewers noted that these papers often featured hallucinated technical jargon, nonsensical writing, and irrelevant citations, making the review process a significant waste of time for legitimate researchers.
Quantifying the Hallucination Epidemic
The problem of AI-generated misinformation in science is not isolated to a few bad actors but is a widespread phenomenon. A Nature analysis from April 2025 suggested that tens of thousands of publications likely contain invalid AI-generated references. Furthermore, an extensive audit of major repositories including arXiv, bioRxiv, SSRN, and PubMed Central by Zhao et al. estimated that approximately 146,900 hallucinated citations were produced in 2025 alone.
Interestingly, these hallucinations are not concentrated in a small number of fraudulent papers but are spread thinly across a vast array of publications. The data suggests that early-career researchers and smaller research teams are the most frequent users of these AI-generated elements. This widespread distribution makes the task of cleaning up the scientific record even more daunting, as the errors are integrated into the broader fabric of academic literature rather than being confined to obvious "paper mills."
The Failure of the Peer-Review Process
Perhaps the most concerning aspect of the "slop" epidemic is the inability of the peer-review system to act as an effective filter. The audit by Zhao et al. traced bioRxiv preprints containing hallucinated references through to their final published versions. They discovered that 85.3% of these hallucinations remained in the published papers, indicating that reviewers and editors are either failing to check citations or are unable to distinguish between real and fabricated data.
This failure is exemplified by the experience of reviewers who flagged papers for having fake authors, only to see those same papers accepted as oral presentations—the highest honor for conference submissions. This suggests a breakdown in the quality control mechanisms of major AI conferences. To combat this, some reviewers are turning to new tools, such as the "bib-audit skill" developed by Isaac, to automate the detection of fabricated references. However, the sheer volume of LLM-generated output continues to overwhelm the manual efforts of the scientific community.
Industry Impact
The influx of AI slop has profound implications for the AI industry and the broader scientific community. First, it threatens the integrity of the scientific record. When fabricated citations and fake authors become commonplace, the foundation of trust upon which cumulative research is built begins to crumble. Researchers may find themselves building upon "hallucinated" technical foundations, leading to a cycle of misinformation.
Second, the burden on the peer-review system is becoming unsustainable. Reviewers are volunteers who donate their time to ensure quality; if the majority of their time is spent debunking AI-generated nonsense, the system may face a total collapse or a significant decline in participation from qualified experts. Finally, the reputation of major AI conferences like NeurIPS and WACV is at risk. If these venues cannot effectively filter out fraudulent or low-quality AI-generated content, their value as benchmarks for scientific progress will be severely diminished.
Frequently Asked Questions
Question: What exactly is "AI slop" in the context of research papers?
AI slop refers to low-quality, often nonsensical content generated by LLMs for academic submissions. This includes hallucinated technical jargon, fabricated citations that do not exist, and even entirely fake author lists for existing papers. It is characterized by writing that appears professional at a glance but lacks factual accuracy or logical coherence.
Question: Why are reviewers failing to catch these fabricated citations?
Data suggests that 85.3% of hallucinated citations persist into final publications. This occurs because the volume of submissions is overwhelming, and checking every single citation for authenticity is a time-consuming manual process. Additionally, LLMs can generate citations that look highly plausible, making them difficult to spot without specialized tools or deep familiarity with the specific sub-field.
Question: Who is most likely to include AI-generated hallucinations in their papers?
According to the audit by Zhao et al., hallucinated citations are spread thinly across many papers rather than being concentrated in a few sources. However, the researchers found that early-career researchers and small research teams are the demographics most likely to include these invalid AI-generated references in their work.

