Cartesia Ink-2 (streaming STT)
Launched English-only (blog 2026-07-09); the current stable `ink-2` snapshot is dated 2026-09-17 and supports English, French, Hindi, Japanese, Spanish. Some press dates an earlier Ink 2 release to May 2026 (unverified). Query params: model, encoding, sample_rate, cartesia_version=2026-08-14; send `finalize` when user stops. Older model: ink-whisper (1 credit/s streaming). Ink-2 credit price not found on pricing page (plans list included STT hours).
- Input
- audio
- Output
- text
- License
- proprietary
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Cartesia API | ink-2 | wss://api.cartesia.ai/stt/websocket?model=ink-2 | docs |
| Cartesia API (beta) | ink-preview | wss://api.cartesia.ai/stt/websocket | — |
Notable capabilities (3)
- Built-in semantic turn detection: Emits turn.start / turn.update / turn.eager_end / turn.resume / turn.end events so agents need no separate VAD; 89% precision, 93% F1 on endpointing; ~0.1 s time-to-final-transcript. source
- #1 streaming WER on Artificial Analysis at launch: 3.4% WER on AA-AgentTalk, ranked #1 on Artificial Analysis's streaming STT leaderboard (company claim, July 2026). source
- Keyterm prompting (found after launch): Keyterm prompting and configurable turn detection added 2026-08-11. source
Cartesia's streaming speech-to-text built for voice agents; pairs with Sonic-3.6.
Sources: https://docs.cartesia.ai/build-with-cartesia/stt/latest , https://docs.cartesia.ai/api-reference/stt/stt , https://www.cartesia.ai/blog/introducing-ink-2