Post-Cutoff.com
  1. Home
  2. Models
  3. Qwen3.8-Omni-Flash

Qwen3.8-Omni-Flash

Alibaba (Qwen)currentmultimodalQwen3.8

Thinking on by default with adjustable effort. Realtime variant qwen3.8-omni-flash-realtime: $0.93 audio in / $1.87 audio out per 1M tokens (Singapore/Intl pricing page, checked 2026-09-29). For dedicated hosted voice agents Alibaba also offers qwen-audio-3.1-realtime-plus (see qwen-audio-3-1-realtime). Release day not verified (OpenRouter listing 2026-09-21).

Context window
1,000,000 tokens
Max output
131,072 tokens
Input
text, image, audio, video
Output
text
License
proprietary
Pricing
input: $0.15 · output: $0.47 · cache read: $0.016 (per 1M tokens (USD), Singapore/International) source
Verified
2026-09-29

How to call it

ProviderModel idEndpoint / URLDocs
Alibaba Cloud Model Studio (DashScope, Singapore/Intl)qwen3.8-omni-flashhttps://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1docs
Alibaba Cloud Model Studio (realtime voice/video)qwen3.8-omni-flash-realtime—docs
OpenRouterqwen/qwen3.8-omni-flashopenrouter.ai/qwen/qwen3.8-omni-flash—
Web app—chat.qwen.ai—

Notable capabilities (3)

Alibaba's omni model for transcription-plus-reasoning, meeting/video analysis and multimodal agents at Flash prices.

curl "https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1/chat/completions" \
 -H "Authorization: Bearer $DASHSCOPE_API_KEY" -H "Content-Type: application/json" \
 -d '{"model":"qwen3.8-omni-flash","messages":[{"role":"user","content":"Hello"}]}'

For audio/video inputs see the non-real-time guide linked from the model page.

Sources: https://www.alibabacloud.com/help/en/model-studio/qwen3-8-omni-flash · https://www.alibabacloud.com/help/en/model-studio/model-pricing

Other Alibaba (Qwen) models

Qwen-Audio-3.1-ASR (Flash) · Qwen-Audio-3.1-Realtime (Plus) · Qwen-Audio-3.1-TTS-Next · Qwen3.8-LiveTranslate (Flash Realtime) · Qwen3.8-27B · Qwen3.8-Flash · Qwen3.8-Max · Qwen-Image-3.0 (Pro) · Qwen3.7-Plus · Qwen3-ASR (0.6B / 1.7B) + Qwen3-ForcedAligner · Qwen3-TTS (open weights 0.6B / 1.7B; API qwen3-tts-flash)