Pangram/UMass study: ~31% of filtered web tokens were AI-generated by Aug 2026, and such 'wild' AI text soon hurts pretraining
A Sept 30, 2026 arXiv paper (Russell, Iyyer, Spero, Emi et al.) measured that 27.5% of FineWeb-filtered tokens in June 2026 web data were AI-generated (by Pangram's detector), rising to 31.1% in August. Pretraining 800 models, it found that such "wild" AI text helps data-starved models a little but quickly turns harmful, and that Chinchilla-style scaling laws fail to predict this. It proposes a new law and releases the 83B-token WildAI corpus.
Key facts
- AI share of FineWeb-filtered web tokens (Pangram classifier): 27.5% in June 2026 → 31.1% in August 2026
- 800 language models pretrained with varying ratios of added AI tokens to human tokens
- Data-starved models: AI tokens first lower loss on human text, then the benefit saturates and reverses into harm; models with large human budgets: AI tokens raise loss almost immediately
- Training on unfiltered web text at August 2026's AI share needs about 1.6x the compute of training on its human subset (at 20 tokens per parameter), a gap that grows with the human budget
- New scaling law with separate benefit and harm terms reduces to Chinchilla without AI text; fit on small models, it predicts models up to 3.6x larger with 41% lower error than the best existing law
- Released: WildAI corpus (83B tokens, with AI/topic/format labels), the 800 models and code (CC BY-NC-SA 4.0)
- Caveat: AI-share numbers depend on Pangram's own detector; Pangram co-founders Max Spero and Bradley Emi are authors
What happened
The authors labelled recent Common Crawl data after FineWeb quality filtering with Pangram's AI-text detector, then ran a large controlled pretraining study that mixed in AI-written web text (not synthetic data from one model, but text from many models written for human readers). They fitted scaling laws to held-out loss on human and on AI text.
Why it matters
It is a quantitative estimate that almost a third of quality-filtered web text is now machine-written, and evidence that this text has negative value for well-resourced pretraining. Pre-2023 human text and data filtering become more valuable, and labs that cannot detect AI text pay a compute penalty.
Changelog
- 2026-10-03: created (lead from aiweekly.co, Oct 2)
Related events
Sources (5)
- paperarXiv 2609.40295: How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
- codeGitHub: pangramlabs/WildAI
- codeHugging Face: pangram/WildAI dataset
- codeHugging Face: pangram/WildAI-models
- discussionKen Ashe: Wild AI web text is starting to poison pretraining
id: 2026-09-30-wild-ai-web-text-scaling-laws · updated 2026-10-03 · open in the interactive timeline