Pew Research: AI 'synthetic respondents' miss real survey answers by 12 points on average
Pew Research Center compared answers from AI-simulated survey respondents (mainly Claude Opus 4.6, also GPT-5.1) with its American Trends Panel on nearly 300 questions from early 2026. The AI answers were off by 12 percentage points on average, stereotyped groups, overstated factual knowledge and rarely said "not sure"; Pew concluded that AI polling is not a replacement for surveying real people.
Key facts
- Nearly 300 questions from three ATP waves (January, March, April 2026); primary model Claude Opus 4.6, also GPT-5.1
- Average error 12 percentage points; about 28% of questions off by more than 15 points, some by 40+
- Larger errors for some groups: Republicans 16.1 points, Black adults 15.1 points
- Examples: Trump approval overstated by 12 points; awareness of data centers underestimated by 22 points; 97% of Hispanic adults said to follow the World Cup vs 43% actual
- Knowledge inflation: e.g. 98% of AI respondents vs 52% of humans answered a First Amendment question correctly; humans chose 'not sure' about 4x as often
- Different models produced contradictory opinion profiles
What happened
Pew's Data Labs team (Chapekis, Lau, Bestvater, Shah, Mercer, Smith) generated "synthetic" respondents with demographic personas and asked them the same questions it had put to real panelists, then compared the distributions.
Why it matters
Startups and some pollsters sell LLM-simulated panels as a cheap substitute for surveys. A large, careful test by the best-known US survey organisation found systematic errors that are largest for exactly the groups and topics where polls matter most.
Changelog
- 2026-10-03: created (lead from Oct 2)
Sources (1)
id: 2026-09-30-pew-ai-synthetic-survey-respondents · updated 2026-10-03 · open in the interactive timeline