Sonnet 5.5 (Fully Tested): The MOST USEFUL MODEL YET! RIP ASTRA & SOL!
AICodeKing · 2026-09-29 · review · 5,433 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
AICodeKing reviews Anthropic's Claude Sonnet 5.5 (released September 28, 2026), testing it via OpenRouter inside the OpenCode coding-agent harness across the eight interactive tasks of KingBench 3. The video evaluates Sonnet 5.5's code generation, 3D Three.js rendering, algorithmic reasoning, and local model training against Claude Opus 5.5 as a reference standard. Sonnet 5.5 scores 71.5 out of 80 (89.38%), placing just behind Opus 5.5 and GLM 5.3.
What is shown
- [00:08] Anthropic's announcement page for Claude Sonnet 5.5 (released September 28, 2026) and the OpenCode setup interface.
- [00:50] OpenCode prompt terminal configured with
anthropic/claude-sonnet-5.5on OpenRouter, with high reasoning effort and up to 128,000 output tokens. - [02:05] Task 1: Elevator simulation — An 8-floor interactive web simulation with 3 color-coded elevators; passenger delivery verified for 20 passengers, with mild crowding and overlapping at 30 people [02:44].
- [03:03] Task 2: Contact lens case in Three.js — 3D render with interactive unscrewing/flipping blue 'L' and red 'R' caps, internal hollow wells, and orbit controls.
- [03:56] Task 3: Folding table in Three.js — Interactive slider folds and unfolds the tabletop and leg frames, inspected from multiple camera angles.
- [04:40] Sponsored segment for Bambood (
bambood.ai), demonstrating multi-agent workflows (architect and workers), Git worktrees, and embedded browser previews. - [06:00] Task 4: Panda eating a burger SVG — Clean, multi-layered vector illustration of a panda holding a sesame seed burger with bamboo background.
- [06:35] Task 5: Archery game ("Bullseye Rush") — 2D canvas game with power meter, wind, arrow drop, 4 moving targets, and leaderboard; automated verification passed, with an Esc-key draw bug noted [07:05].
- [07:19] Task 6: Counting problem — Mathematical permutation problem on a $3 \times 20$ grid; Sonnet 5.5 generates a C++ solver, resolves macOS compiler issues autonomously, and finds the exact count (20,460).
- [08:12] Task 7: Local fine-tuning — Creates a 119-fact panda dataset, fine-tunes Gemma 2B locally using LoRA, logs validation loss, and serves facts via a local web UI.
- [09:21] Task 8: 3D wristwatch in Three.js ("Chronos") — Dual-timezone watch with smooth sweep second hand, date/day complications, night mode glow, and strap/dial customization.
- [10:19] Final leaderboard & breakdown — Task-by-task rating breakdown and comparison against previous KingBench 3 runs.
Claims & numbers
- The presenter notes Anthropic released Claude Sonnet 5.5 on September 28, 2026.
- The presenter cites Anthropic's published model specs: 1,000,000-token context window, $2 per million input tokens, $10 per million output tokens, and $0.20 per million cached read tokens.
- Testing parameters: High reasoning effort with a 128,000 output token limit per generation.
- Correct solution for Task 6 is stated as exactly 20,460 valid paths, which Sonnet 5.5 computed correctly.
- Task scores out of 10 awarded by the presenter:
- Elevator sim: 8 / 10
- Contact lens case: 9.5 / 10
- Folding table: 8 / 10
- Panda SVG: 9 / 10
- Archery game: 8.5 / 10
- Counting problem: 10 / 10
- Local fine-tuning: 9 / 10
- 3D wristwatch: 9.5 / 10
- Total score: 71.5 / 80 (89.38%, overall grade 8.94 / 10).
- The presenter compares this against KingBench 3 reference scores: Claude Opus 5.5 at 93.75% (75/80), GLM 5.3 at 91.25%, Step 5 Preview at 83.75%, GLM 5.3 Flash at 78.75%, MiMo V2.6 Flash at 72.5%, and MiMo V2.6 Pro at 69.38%.
- Bambood workspace is priced at $20/month for personal use on two devices.
Notable quotes
- [00:30] "Sonnet 5.5 works really well on these interactive coding tasks."
- [03:37] "Those details make this feel like a complete little product demo."
- [10:53] "Sonnet works really well, and Opus remains my preferred model for overall output quality."
Assessment
This is an independent hands-on benchmark review evaluating Claude Sonnet 5.5 through a live agent harness. The code generations, interactive Three.js models, gameplay, and terminal scripts are demonstrated directly on screen, with fair and transparent critiques of small mechanical and visual imperfections.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.