MiniMax-M3
MiniMax flagship LLM. OpenAI-format responses include <think> content that must be preserved across turns. MiniMax-M3.1-Flash-Preview (1M, tunable thinking) exists but only via Token Plan/MiniMax Code. Max output and knowledge cutoff not verified.
- Context window
- 1,000,000 tokens
- Input
- text, image, video
- Output
- text
- License
- minimax-community
- Pricing
- input: $0.3 · output: $1.2 · cache read: $0.06 (per 1M tokens (USD), standard tier, input <=512K (after permanent 50% discount); >512K input: 0.60/2.40/0.12. Priority tier 1.5x) source
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| MiniMax API (Anthropic format) | MiniMax-M3 | https://api.minimax.io/anthropic | docs |
| MiniMax API (OpenAI format) | MiniMax-M3 | https://api.minimax.io/v1 | docs |
| OpenRouter | minimax/minimax-m3 | openrouter.ai/minimax/minimax-m3 | — |
| Hugging Face | — | huggingface.co/MiniMaxAI/MiniMax-M3 | — |
Notable capabilities (3)
- MiniMax Sparse Attention (MSA): New sparse attention for million-token contexts: 9x prefill and 15x decode speed-up vs M2 at 1M context, ~1/20 per-token compute. source
- Native multimodality from step one: Mixed text/image/video training from the start of pre-training (~428B total / ~23B active). source
- Three reasoning modes: thinking parameter selects among three reasoning modes; interleaved thinking with tool use. source
Low-cost 1M-context multimodal coding/agent model; Anthropic-SDK-first API, open weights.
curl https://api.minimax.io/anthropic/v1/messages \
-H "x-api-key: $MINIMAX_API_KEY" -H "anthropic-version: 2023-06-01" -H "Content-Type: application/json" \
-d '{"model":"MiniMax-M3","max_tokens":4096,"messages":[{"role":"user","content":"Hello"}]}'
Sources: https://platform.minimax.io/docs/guides/models-intro · https://platform.minimax.io/docs/guides/pricing-paygo · https://huggingface.co/MiniMaxAI/MiniMax-M3
Other MiniMax models
MiniMax H3 · MiniMax Music 3.0 · MiniMax-M2.7 · MiniMax Speech 2.8 (HD / Turbo)