We Tested Anthropic's Fable 5.1 for a Week
Every · 2026-09-08 · review · 29,999 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Dan Shipper, co-founder and CEO of publication and product lab Every, reviews Anthropic's Claude Fable 5.1 after one week of early testing across coding, knowledge work, and writing workflows. He breaks down where the model excels—notably autonomous coding and delegating multi-hour agentic tasks—and examines benchmark comparisons against Opus 5 and GPT-5.6.
What is shown
- [01:46] Hands Agent Demo: Demonstrates "Hands", an autonomous Mac desktop computer-use agent built end-to-end by Fable 5.1 via UltraCode using ~40 subagents, receiving instructions in Slack and driving browser tasks in ChatGPT.
- [05:07] Internal Agent Benchmark: Every’s internal agent benchmark dashboard comparing token consumption (766 tokens/run for Fable 5.1 vs. 1,939 for Opus 5) and latency (22s vs. 37s).
- [06:36] Knowledge Work - Data Analysis & Dashboard Generation: On EC Bench ("01 dashboard"), Fable 5.1 processes real NPS survey data and builds a clean interactive static HTML dashboard, scoring 88/100 compared to GPT-5.6's 100/100 score [07:26].
- [08:52] Knowledge Work - Presentation Deck: Demonstrates Keynote slides created end-to-end from an essay on "Compound Engineering", highlighting layout execution, diagramming bubbles, and arrow routing compared to GPT-5.6 [10:04].
- [11:00] Meeting Strategy Extraction: A transcript analysis tool summarizing a launch strategy debate and flagging strategic decisions where Shipper needed to act as tiebreaker.
- [13:33] Writing Evaluation: An EC Bench writing test ("03 writeup") converting an interview transcript with Every's Mike Taylor into a structured blog post ("Raise the Ceiling, Not the Floor"), scoring 67/100 on Fable 5.1 versus 78/100 on Opus 5 [14:49].
- [16:47] Prose Critique & Structural Flow: Demonstrates Fable 5.1 analyzing a draft titled "How Codex Happened" to identify where momentum faltered.
- [18:19] Personal Usage Telemetry Dashboard: Displays personal usage shifts after receiving access on August 24, showing prompt frequency and token consumption surges across Codex/ChatGPT vs. Claude Code.
Claims & numbers
- Coding & Speed: The presenter claims Fable 5.1 is roughly twice as fast as the original Claude Fable and uses approximately half the tokens of Claude Opus 5 for comparable tasks.
- Agent Benchmark: On Every's internal agent benchmark, Fable 5.1 averaged 766 tokens per task run versus 1,939 tokens for Opus 5, with an average response latency of 22 seconds compared to 37 seconds for Opus 5.
- Autonomous Coding Cost: Long autonomous UltraCode runs with ~40 subagents can consume 3 to 5 million tokens over a full day.
- Usage Telemetry: After receiving Fable 5.1 access on August 24, Shipper’s Claude model step share rose from 19.6% to 65.4% (+45.8 percentage points), with three long-running parent agent sessions accounting for 98% of all Claude tokens consumed (Ghostseed at 59.1%, personal feed experiment at 26.1%, and Proof benchmark at 12.8%).
Notable quotes
- [02:42] "I have no idea how this works. This was built end-to-end by Fable 5.1 from a couple prompts."
- [05:33] "It's about twice as fast as Opus and it uses about half the tokens."
- [17:30] "It's actually zeroing in on the right part of the problem and then telling me how to fix it."
Assessment
This is an authentic practitioner review and hands-on benchmark evaluation by an early-access user. The presenter provides verifiable screen recordings of internal tools (EC Bench, live agent execution logs, and analytics dashboards) alongside balanced critique of where the model still lags behind competitors like GPT-5.6.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.