NeurIPS 2026 desk-rejects papers with hallucinated references and runs a randomized LLM-assisted reviewing experiment
NeurIPS 2026 treats hallucinated citations as a Code of Conduct violation. An area chair reported on Aug 18, 2026 that 5 of the 8 submissions in his batch had two or more fabricated references and would likely be desk-rejected. The same cycle brought AI into the process officially: an opt-in Google Gemini "Paper Assistant Tool" gave authors feedback before submission, and an IRB-approved experiment randomly assigned reviewers to no LLM help, open-ended help or structured help from an LLM assistant built into OpenReview.
Key facts
- NeurIPS 2026 Main Track Handbook: 'There have been many cases of hallucinated citations in literature review, which violates the NeurIPS Code of Conduct'
- Area chair Danish Pruthi (X, Aug 18, 2026, ~23k views): '5 out of 8 submissions have 2+ hallucinated references and will likely be desk rejected'. In a follow-up he said NeurIPS is 'quite liberal' in how it defines hallucinations, and a single fabricated citation was not enough for desk rejection. The two-reference threshold comes only from his posts, not from a published NeurIPS rule
- StrictCite (Sept 5) says NeurIPS, ICLR, ICML and ACL all desk-rejected 2026 submissions over hallucinated citations. It cites an audit (arXiv 2607.00738) of ~48,000 accepted papers that found ≥1 hallucinated reference in 26.2% of NeurIPS 2025 papers and ≥2 in 5.1% (StrictCite's summary; the audit was not checked directly)
- AI-assisted reviewing experiment: participating reviewers are randomly assigned per paper to (1) no LLM assistance, (2) open-ended LLM assistance or (3) structured LLM assistance, through an LLM interface inside OpenReview with zero data retention. It only covers papers whose authors opted in at submission. Area chairs, blind to condition, then rate review quality
- Outside the experiment NeurIPS sanctions no reviewer LLM use; violations 'may result in consequences for reviewers and their submitted papers, including desk rejection'
- Pre-submission: authors could run each paper once through Google's Gemini-based Paper Assistant Tool (PAT) during the 7 days before the May 4, 2026 abstract deadline. PAT feedback is not used in reviewing (NeurIPS blog, Apr 21). PAT paper: arXiv 2606.28277 (Google Research, June 26), with earlier pilots at STOC and ICML
- Comparable evidence: an ICML 2026 randomized experiment (arXiv 2609.19420, 24,000+ papers, 17,000 reviewers) found that a permissive vs restrictive LLM policy had 'near-zero effects' on decisions and scores. But 22.5% of reviewers under the restrictive policy admitted using LLMs anyway
What happened
NeurIPS 2026, the largest ML conference, dealt with AI on both sides of peer review. Before submission, authors could opt in to one automated critique per paper from Google's Gemini-based Paper Assistant Tool. During review, NeurIPS ran a randomized, IRB-approved experiment. Volunteer reviewers on opted-in papers got no LLM help, open-ended help or structured help from an assistant built into OpenReview, and blinded area chairs rated the resulting reviews. Any other reviewer LLM use stayed banned.
NeurIPS also enforced its rule against fabricated citations. On Aug 18 area chair Danish Pruthi reported that most papers in his batch had several hallucinated references and faced desk rejection, while one fabricated citation alone was not enough.
Why it matters
Fabricated references are an easy-to-check sign of unchecked LLM writing, and the top ML venues have started enforcing against them. At the same time they are testing, under controlled conditions, whether LLM help improves reviews. The ICML 2026 experiment suggests that policy alone barely changes outcomes and that many reviewers ignore restrictions.
Changelog
- 2026-10-02: created (leads: strictcite.com / Zvi AI #188 trail)
Related posts (2)
- Danish Pruthi original ↗ Danish Pruthi @danish037 · x · 2026-08-18
Cited as a source by: 2026-08-18-neurips-2026-hallucinated-references-ai-reviewing - Danish Pruthi original ↗ Danish Pruthi @danish037 · x · 2026-08-18
Cited as a source by: 2026-08-18-neurips-2026-hallucinated-references-ai-reviewing
Related events
- NeurIPS 2026: 28% of position-track submissions score 100% AI-written, and 178 are desk-rejected ★★
- ICLR 2026 review crisis: 21% of peer reviews flagged fully AI-written, and an OpenReview bug exposes reviewer identities ★★★
- arXiv will ban authors for a year if they post unchecked LLM-generated content ★★★
Sources (8)
- officialNeurIPS 2026 AI-Assisted Reviewing Experiment
- officialNeurIPS 2026 Main Track Handbook
- officialNeurIPS blog: NeurIPS supports authors with Google's Paper Assistant Tool (PAT)
- discussionDanish Pruthi on X: 5 of 8 submissions have 2+ hallucinated references
- discussionDanish Pruthi on X: NeurIPS is quite liberal in how it defines hallucinations
- discussionStrictCite: NeurIPS started desk-rejecting papers over references that don't exist
- paperarXiv 2606.28277: Towards Automating Scientific Review with Google's Paper Assistant Tool
- paperarXiv 2609.19420: Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026
id: 2026-08-18-neurips-2026-hallucinated-references-ai-reviewing · updated 2026-10-02 · open in the interactive timeline