GLM-5.3-Flash / FlashX
Z.ai says it beats GLM-5.2 at a fraction of the cost; 3x Coding Plan quota vs GLM-5.3 (FlashX not yet on the plan). Thinking cannot be disabled. 'first' claim is the vendor's own.
- Context window
- 1,000,000 tokens
- Max output
- 128,000 tokens
- Input
- text, image, video, pdf
- Output
- text
- License
- mit
- Pricing
- input: $0.15 · output: $0.5 · cache read: $0.03 (per 1M tokens (USD) for glm-5.3-flash; glm-5.3-flashx (~200 tok/s): 0.37 in / 1.25 out / 0.075 cached) source
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Z.ai API | glm-5.3-flash | https://api.z.ai/api/paas/v4 | docs |
| Z.ai API (fast) | glm-5.3-flashx | https://api.z.ai/api/paas/v4 | docs |
| OpenRouter | z-ai/glm-5.3-flash | openrouter.ai/z-ai/glm-5.3-flash | — |
| OpenRouter (FlashX) | z-ai/glm-5.3-flashx | openrouter.ai/z-ai/glm-5.3-flashx | — |
| Hugging Face | — | huggingface.co/zai-org/GLM-5.3-Flash | — |
| Web app | — | chat.z.ai | — |
Notable capabilities (3)
- First native multimodal GLM-5 model: First GLM-5-series model with native vision (image, video, file input); vision used inside the coding loop (UI replication, Blender, browser/computer-use agents). source
- FIRST Sparse + linear attention hybrid: 320B total / 18B active; Z.ai claims it is the first open-source frontier model combining sparse and linear attention (3.01x less attention compute, 4.44x smaller KV cache vs GLM-5.3). source
- Office deliverables with visual self-check: Produces PPTX/PDF/DOCX/XLSX and renders them to catch overflow and layout issues. source
Cheap multimodal GLM for visual coding, agents and office documents; open weights under MIT.
curl https://api.z.ai/api/paas/v4/chat/completions \
-H "Authorization: Bearer $ZAI_API_KEY" -H "Content-Type: application/json" \
-d '{"model":"glm-5.3-flash","messages":[{"role":"user","content":[{"type":"image_url","image_url":{"url":"https://example.com/ui.png"}},{"type":"text","text":"Rebuild this UI in React"}]}]}'
Sources: https://docs.z.ai/guides/vlm/glm-5.3-flash · https://docs.z.ai/guides/overview/pricing · https://huggingface.co/zai-org/GLM-5.3-Flash