OpenAI cancels Astra release, Sonnet 5.5 & what Meta Muse means for work
IBM Technology · 2026-10-02 · review · 16,860 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary This video is an episode of IBM's Mixture of Experts podcast hosted by David Zax, featuring IBM guests Chris Hay (Distinguished Engineer), Madison Gooch (VP watsonx Americas), and Ash Minhas (AI Engineer and Practice Manager). The panel discusses the reported cancellation of OpenAI's GPT-6.1 Astra release, the launch of Anthropic's Claude Sonnet 5.5 and mid-tier model economics, and the enterprise implications of Meta's consumer agent Muse.
What is shown
- A remote four-way video conference panel discussion hosted by David Zax with Chris Hay, Madison Gooch, and Ash Minhas [00:47].
- Discussion of the Wall Street Journal report regarding OpenAI canceling the October release of GPT-6.1 Astra due to alignment test failures around deception and scope authorization [01:17].
- Discussion of Claude Sonnet 5.5 and developer practices for benchmarking and choosing between mid-tier and frontier flagship models (e.g., Opus 5.5 vs. Sonnet 5.5) [11:42].
- Discussion of token governance, FinOps, token budgeting, and IBM's "Bob" agentic router within watsonx [18:45].
- Discussion of Meta's consumer agent Muse, agentic computer use, integrations (Plaid, Shopify), and the transition of consumer agent behaviors into enterprise workflows [23:05].
Claims & numbers
- David Zax states the Wall Street Journal reported OpenAI is scrapping plans to release GPT-6.1 Astra in October after it underperformed on safety metrics including deception and scope authorization [01:21].
- Chris Hay states he often deploys agent swarms with around 100 agents to execute complex tasks [00:05, 27:31].
- Chris Hay claims enterprise data shows only around 5% to 6% of model usage goes to the top frontier tier (like Claude Fable or Opus), while over 90% of enterprise usage remains on mid-tier workhorse models [16:24].
- David Zax notes that activating "max thinking" in model interfaces prompts warnings of "3.5x usage" against token allowances [18:11].
- Ash Minhas states his personal visual test benchmark is prompting models to generate an SVG of a "ninja cat in Jersey City" [21:48].
- David Zax states Meta Muse topped the mobile app charts the prior week [23:46].
- Madison Gooch claims IBM systems still process 99% of global transactions [29:38].
Notable quotes
- Chris Hay [03:42]: "I'm like addicted to models. You give me, 'Oh, here's a new model, I want to go play with it.' Here's another model, I want to go and play with it."
- Madison Gooch [09:04]: "Most of the time, you know what that bias is? It's the largest model and it's the model that cost the business the most money."
- Ash Minhas [21:48]: "My benchmark is to say: generate me an SVG image of a ninja cat in Jersey City, because that's where I live..."
Assessment This is a panel review and analysis podcast rather than a product demonstration. No live software, code execution, or benchmark outputs are visually shared on screen; the entire episode consists of the four participants conversing over webcam feeds.
Described by gemini-3.8-flash on 2026-10-03 from the video's audio and frames.