Amazon Nova 2 Sonic
Technical report (Amazon Nova 2, Dec 2025, https://www.amazon.science/publications/amazon-nova-2-multimodal-reasoning-and-generation-models): Big Bench Audio 87.0 (Artificial Analysis) vs GPT-Realtime (Aug 2025) 83.0 and Gemini 2.5 Flash Live 71.0; BFCL subset 74.5; ComplexFunction 65.2; Common Voice avg WER 6.5 vs 8.4 (GPT-Realtime) across 7 languages; human-preference win rate vs GPT-Realtime above 50% for 6 of 8 voices (e.g. 68.4% Spanish) but 42.4% Hindi and 26.3% Portuguese; vs Gemini 2.5 Flash Live 47.5-77.9%. Comparisons are against 2025 competitors. Successor to Nova Sonic (amazon.nova-sonic-v1:0, Apr 2025). Bedrock only, In-Region in us-east-1, us-west-2, eu-north-1, ap-northeast-1 (no cross-region inference); Standard tier only. Lifecycle Active, EOL no sooner than 2026-12-02. No newer Nova Sonic found as of 2026-09-29; per July 2026 reports Nova 2 Sonic is among the Nova models Amazon keeps developing after its Nova wind-down. Prices from secondary source (AWS Nova pricing page does not list per-token rates).
- Context window
- 1,000,000 tokens
- Max output
- 64,000 tokens
- Input
- audio, text
- Output
- audio, text
- License
- proprietary
- Pricing
- speech input: $3 · speech output: $12 · text input: $0.33 · text output: $2.75 (USD per 1M tokens (speech in/out, text in/out); secondary source, not confirmed on the AWS pricing page) source
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| AWS Bedrock | amazon.nova-2-sonic-v1:0 | https://bedrock-runtime.{region}.amazonaws.com (InvokeModelWithBidirectionalStream) | docs |
Notable capabilities (3)
- Real-time speech-to-speech: Single model for natural real-time voice conversations over a bidirectional streaming API (no separate ASR/TTS pipeline). source
- 1M-token session context: 1M-token context window and 64K max output listed for long-running voice sessions. source
- Polyglot voices and turn-taking control: Same voice speaks multiple languages natively (Portuguese and Hindi added vs Nova Sonic); developers set low/medium/high pause sensitivity. source
Voice agents and conversational IVR on Bedrock.
Use the Bedrock InvokeModelWithBidirectionalStream API with model id amazon.nova-2-sonic-v1:0 (see AWS samples in the model card).
Sources: https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-amazon-nova-2-sonic.html · https://cdn.amazon.science/c5/3d/84514a224666b5be6de4b43ef4aa/nova-2-0-technical-report2.pdf