Post-Cutoff.com
  1. Home
  2. Models
  3. SeedRealtime (Doubao realtime audio-visual model)

SeedRealtime (Doubao realtime audio-visual model)

ByteDancecurrentaudio/speechSeed

Deployed at scale in the Doubao app (Dola internationally). No public API model id, pricing or benchmark numbers published; Volcengine offers a separate Doubao end-to-end realtime dialogue API (/api/v3/realtime/dialogue) whose relation to SeedRealtime is unverified. Some press calls it the first model to watch, listen and speak simultaneously; not claimed by ByteDance, and Gemini Live / GPT-Realtime already accepted video.

Input
audio, video, text
Output
audio, text
License
proprietary

How to call it

ProviderModel idEndpoint / URLDocs
Web app (Doubao / Dola)—dola.com/chat—
BytePlus Playground—ai.byteplus.com/en/playground—

Notable capabilities (2)

Timeline entry

  1. ByteDance deploys SeedRealtime, a native audio-visual full-duplex model, in the Doubao app ★★★

    ByteDance Seed launched SeedRealtime, an end-to-end LLM that listens, watches (live video) and speaks at the same time instead of chaining ASR, vision and TTS, and rolled it out at scale in Doubao (Dola internationally). ByteDance says it halves conversational pacing problems compared with…

Other ByteDance models

Seed Audio 1.0