Claude Fable 5.1 | First impressions
Arena AI · 2026-09-08 · review · 47,563 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Peter Gostev, AI Capability Lead at Arena, reviews the newly released Claude Fable 5.1 model, evaluating its performance across diverse complex generation benchmarks on Arena's testing platform. He tests and compares Fable 5.1 Max against earlier models like Claude Fable 5, Claude Opus 5, GPT-5.6 Sol, Kimi K3, and others on intricate 3D web environments, interactive browser games, SVG rendering, and data-intensive white-collar research applications.
What is shown
- Anthropic benchmark table and release notes showing Claude Fable 5.1 benchmark improvements and cache-read pricing details [00:28].
- 3D interactive model generation of Westminster in Three.js/HTML, showing Claude-Fable-5.1-Max ($65) alongside GPT-5.6-Sol and Claude-Fable-5 outputs [01:03].
- Procedural Three.js dinosaur sanctuary simulation with animated sauropods, comparing Fable 5.1 Max, Fable 5, GPT-5.6 Sol, and Kimi-K3 [03:30].
- "The Cocoa Conservatory" procedural chocolate factory prompt test across multiple models [06:01].
- Vector SVG generation of the Mona Lisa, contrasting Fable 5.1 Max's detailed portrait against cartoon-style outputs from Fable 5, GPT-5.6 Sol, and Kimi-K3 [09:18].
- Browser game creation: a 3D downhill sandboarding game in Giza, tested on Fable 5.1 Max, Fable 5, GPT-5.6 Sol, Kimi-K3, Qwen3.8-Max, and GLM-5.3 [11:05].
- Browser game creation: "Canal Dash" Venice boat navigation game [15:10] and "Rooftop Rush" runner game [17:16].
- Interactive 3D space elevator climb visualization ("Ascent Line 7") ascending into orbit, comparing Fable 5.1 Max ($47.13) to Fable 5 and GPT-5.6 Sol [19:22].
- Artistic 3D Three.js scene reconstructions: Monet's Japanese footbridge water lilies [21:40] and grain stacks [23:34].
- Massive 3D city generation of Istanbul, comparing Fable 5.1 Max to GLM-5.3, Qwen3.8-Max, Grok-4.6-Xhigh, and DeepSeek-V4-Pro-Max [26:22].
- White-collar research workflows: Swiss Alps interactive hiking terrain dossiers [31:10], AI hiring constellation network visualization [35:56], a 12-month global AI conference itinerary planner [38:51], a 45-person office hub decision brief [40:18], a global AI Compute Atlas tracker [41:14], and an NVIDIA executive statements accountability audit [45:32].
- An interactive exploded 3D assembly and global supply chain explorer for the Boeing 787 Dreamliner [47:12].
- 3D Cappadocia sunrise hot air balloon simulation across all tested models [50:33].
Claims & numbers
- The presenter notes Anthropic's blog states Fable 5.1 will cost an estimated 25% less for typical workloads where usage is billed by tokens due to reductions on cache reads, with savings up to approximately 45% for highly agentic work [00:38].
- On benchmarks shown: Fable 5.1 scores 52.6% on Agentic scientific research (Terminal-Bench-Science 0.1), 55.9% on Agentic coding (Terminal-Bench 4.0), 1853 on Knowledge work (GPQA-AA v2), 77.9% on Computer use (OSWorld 2.0), 41.7% on OSWorld 2.0 without tools, 60.9% on Multidisciplinary reasoning (Humanity's Last Exam), 31.4% on Business workflows (AutomationBench), and 73.4% on Agentic coding (CursorBench 3.0) [00:30].
- The Westminster generation cost $65 on Claude-Fable-5.1-Max versus $3.10 on GPT-5.6-Sol [01:22, 02:24].
- The Mona Lisa SVG cost $22.06 on Claude-Fable-5.1-Max, compared to $0.21 on GPT-5.6-Sol and $0.56 on Kimi-K3 [09:54, 10:28, 10:46].
- The Venice canal game cost $35.65 on Claude-Fable-5.1-Max [15:15], and the Rooftop Rush game cost $40.39 [17:41].
- The space elevator visualization cost $47.13 on Claude-Fable-5.1-Max versus $3.06 on GPT-5.6-Sol [19:50, 21:01].
- The 45-person office hub brief cost $23.53 on Fable 5.1 Max compared to $5.86 on Fable 5 [41:10].
- The AI Compute Atlas research task cost $65.81 on Fable 5.1 Max and $10.49 on Fable 5 [44:12].
- The presenter claims that while Claude Fable 5.1 Max produces exceptionally detailed and realistic outputs, its total execution costs remain significantly higher than alternative models [50:14].
Notable quotes
- "This is the first time we're seeing a new type of model being updated with any kind of fixes that Anthropic saw that maybe they could do to improve the model." [00:06]
- "It did cost me, it's probably the most expensive SVG you will ever see, twenty-two dollars." [10:24]
- "If you get the best Fable generations, they're absolutely insane and really, really excellent." [22:27]
Assessment
This is a hands-on review and comparative analysis conducted by Arena AI evaluating Claude Fable 5.1 Max against earlier Claude checkpoints and rival frontier models. All outputs are demonstrated live inside the browser from authentic model generations and workspace files, honestly highlighting both Fable 5.1's high quality and its steep generation costs.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.