Post-Cutoff.com
  1. Home
  2. Models
  3. DeepSeek-V4.1-Flash

DeepSeek-V4.1-Flash

DeepSeekcurrentreasoning-llmDeepSeek V4open weights

Call as deepseek-flash. Legacy ids deepseek-v4-flash and deepseek-v4-flash-vision-exp are routed here and billed at Flash price. Knowledge cutoff not published.

Context window
1,000,000 tokens
Max output
384,000 tokens
Input
text, image
Output
text
License
mit
Pricing
input: $0.3 · output: $1.2 · cache read: $0.006 (per 1M tokens (USD), peak-hour list price; off-peak is half (input 0.15, output 0.6, cache hit 0.003). Peak = 01:00-04:00 and 06:00-10:00 UTC Mon-Fri) source
Verified
2026-09-29

How to call it

ProviderModel idEndpoint / URLDocs
DeepSeek APIdeepseek-flashhttps://api.deepseek.comdocs
DeepSeek API (Anthropic format)deepseek-flashhttps://api.deepseek.com/anthropicdocs
Alibaba Cloud Model Studiodeepseek-v4.1-flash—docs
OpenRouterdeepseek/deepseek-v4.1-flashopenrouter.ai/deepseek/deepseek-v4.1-flash—
Hugging Face—huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash—
Web app—chat.deepseek.com—

Notable capabilities (5)

DeepSeek's cheapest current model: agentic coding, long-context (1M) work and image understanding at very low cost. Schedule batch jobs off-peak for 50% off.

curl https://api.deepseek.com/chat/completions \
 -H "Authorization: Bearer $DEEPSEEK_API_KEY" -H "Content-Type: application/json" \
 -d '{"model":"deepseek-flash","messages":[{"role":"user","content":"Hello"}]}'

Sources: https://api-docs.deepseek.com/quick_start/pricing · https://api-docs.deepseek.com/updates · https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash

Other DeepSeek models

DeepSeek-V4-Pro · DeepSeekMath-V2 · DeepSeek-V3.2