Qwen3.8-Flash
Low-cost default in Model Studio (maps to 'GPT-5.4-mini / Haiku 4.5' tier per Alibaba). Max output not verified. Release day not verified (OpenRouter listing 2026-08-26).
- Context window
- 1,000,000 tokens
- Input
- text, image, video
- Output
- text
- License
- qwen-community-1.0 (open weights Qwen3.8-Flash-Next)
- Pricing
- input: $0.15 · output: $0.47 (per 1M tokens (USD), Singapore/International region, input up to 1M) source
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Alibaba Cloud Model Studio (DashScope, Singapore/Intl) | qwen3.8-flash | https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1 | docs |
| OpenRouter | qwen/qwen3.8-flash | openrouter.ai/qwen/qwen3.8-flash | — |
| Hugging Face (Qwen3.8-Flash-Next, base of the API model) | — | huggingface.co/Qwen/Qwen3.8-Flash-Next | — |
| Web app | — | chat.qwen.ai | — |
Notable capabilities (3)
- Preview of the Qwen4 architecture: Built on Qwen3.8-Flash-Next, an experimental preview of the architecture that will underpin Qwen4 (Gated DeltaNet + Qwen Sparse Attention, Gated Residual, N-gram Embedding). source
- Block-level sparse attention (QSA): Qwen Sparse Attention selects micro-blocks rather than tokens, cutting long-context latency for agentic workloads. source
- OpenAI + Anthropic protocol compatibility: Works directly with Claude Code and Codex; 1M context, image/video understanding, desktop-app operation. source
Cheap, fast multimodal workhorse for coding assistants, agents and high-concurrency apps; 1M context.
curl "https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1/chat/completions" \
-H "Authorization: Bearer $DASHSCOPE_API_KEY" -H "Content-Type: application/json" \
-d '{"model":"qwen3.8-flash","messages":[{"role":"user","content":"Hello"}]}'
Sources: https://www.alibabacloud.com/help/en/model-studio/qwen3-8-flash · https://www.alibabacloud.com/help/en/model-studio/model-pricing · https://huggingface.co/Qwen/Qwen3.8-Flash-Next
Other Alibaba (Qwen) models
Qwen-Audio-3.1-ASR (Flash) · Qwen-Audio-3.1-Realtime (Plus) · Qwen-Audio-3.1-TTS-Next · Qwen3.8-LiveTranslate (Flash Realtime) · Qwen3.8-Omni-Flash · Qwen3.8-27B · Qwen3.8-Max · Qwen-Image-3.0 (Pro) · Qwen3.7-Plus · Qwen3-ASR (0.6B / 1.7B) + Qwen3-ForcedAligner · Qwen3-TTS (open weights 0.6B / 1.7B; API qwen3-tts-flash)