Anthropic: an agent running an internal model merged 3,000+ changes in two weeks and made Claude.ai ~3x faster
In a Sept 23, 2026 engineering post, Anthropic said a two-week August sprint made Claude.ai and the Claude desktop app about 3x faster, with most of the work done by Claude Tag (Claude in Slack) running "an internal research model roughly comparable to Opus 5.5". Given standing instructions, the agent found bottlenecks, wrote benchmarks, shipped fixes and watched deployments; more than 3,000 changes were merged with no customer-facing incident or rollback.
Key facts
- Fresh page load 3.1 s → 0.55 s (5.6x); desktop cold start 6.3 s → 3.3 s (1.9x); loading conversations 1.6–2.6 s → 0.5–0.7 s; sending messages 2.2x–19x faster depending on platform
- Agent: Claude Tag (beta) in a dedicated Slack channel, running an internal research model 'roughly comparable to Opus 5.5'
- 3,000+ changes merged in two weeks with 'not a single customer-facing incident or rollback'
- Estimated saving: 'tens of thousands of user-hours of waiting every day'
- Authors: Raymond Wang, Sam Attard, Issac G.; 'Once Claude can measure something, it can make it faster. So we kept finding more things to measure.'
What happened
Anthropic pointed a Slack-resident agent at the performance of its own consumer app and let it run as a continuous engineer: measure, change, deploy, monitor, under standing instructions from the team. The blog gives before-and-after numbers for the main user flows.
Why it matters
A first-party account of an AI agent doing a large, production-grade engineering project on a lab's own flagship product, at a volume of changes no small team could match. It is another data point, alongside Z.ai's claim about GLM-5.3 building its own serving stack, that labs now use their models to improve their own products and infrastructure.
Changelog
- 2026-09-30: created (sweep 2026-09-29: Techmeme Sept 24 + HN)
Related events
- Anthropic releases Claude Opus 5.5 — Fable-5.1-level performance at $4/$20, first model of the Claude 5.5 family ★★★★★
- Z.ai says GLM-5.3 largely built the inference stack that serves GLM-5.3-Flash, calling it an early step toward recursive self-improvement ★★★
Sources (2)
- officialAnthropic (claude.dev): How we made Claude.ai faster
- discussionHacker News discussion (230 points)
id: 2026-09-23-anthropic-claude-tag-makes-claude-ai-3x-faster · updated 2026-09-30 · open in the interactive timeline