Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2026
  4. Pangram/UMass study: ~31% of filtered web tokens were…

Pangram/UMass study: ~31% of filtered web tokens were AI-generated by Aug 2026, and such 'wild' AI text soon hurts pretraining

★★★after cutoffresearchPangram LabsUMass Amherstconfidence: high

A Sept 30, 2026 arXiv paper (Russell, Iyyer, Spero, Emi et al.) measured that 27.5% of FineWeb-filtered tokens in June 2026 web data were AI-generated (by Pangram's detector), rising to 31.1% in August. Pretraining 800 models, it found that such "wild" AI text helps data-starved models a little but quickly turns harmful, and that Chinchilla-style scaling laws fail to predict this. It proposes a new law and releases the 83B-token WildAI corpus.

Key facts

What happened

The authors labelled recent Common Crawl data after FineWeb quality filtering with Pangram's AI-text detector, then ran a large controlled pretraining study that mixed in AI-written web text (not synthetic data from one model, but text from many models written for human readers). They fitted scaling laws to held-out loss on human and on AI text.

Why it matters

It is a quantitative estimate that almost a third of quality-filtered web text is now machine-written, and evidence that this text has negative value for well-resourced pretraining. Pre-2023 human text and data filtering become more valuable, and labs that cannot detect AI text pay a compute penalty.

Changelog

  • 2026-10-03: created (lead from aiweekly.co, Oct 2)

Related events

  1. Semafor finds AI-written guest op-eds in the NYT, WSJ and Washington Post; Dartmouth provost's case becomes a campus scandal ★★

Sources (5)

id: 2026-09-30-wild-ai-web-text-scaling-laws · updated 2026-10-03 · open in the interactive timeline