MiniMax H3
Replaces Hailuo 2.3 / 2.3-Fast / 02 (now legacy: e.g. MiniMax-Hailuo-2.3 0.28 USD per 768P 6s clip). Modes: T2V, I2V, first/last frame, multimodal reference; 4-15 s, 24 fps. Open release is full-attention only.
- Input
- text, image, video, audio
- Output
- video, audio
- License
- minimax-h3-community-license-agreement
- Pricing
- per second 768p: $0.08 · per second 2k: $0.13 (per second of output video (USD). H3-Max (fal.ai post-trained, fast): 480P 0.05/s, 768P 0.08/s. Extra input images 0.04 each after 5 free) source
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| MiniMax API (Video Generation V2) | MiniMax-H3 | https://api.minimax.io | docs |
| MiniMax API (fast variant) | MiniMax-H3-Max | https://api.minimax.io | docs |
| Hugging Face | — | huggingface.co/MiniMaxAI/MiniMax-H3 | — |
Notable capabilities (3)
- Open omni-modal video model with native audio: Understands mixed text/image/video/audio context and generates video with native stereo audio, up to 2K and 15 s. source
- H3-Context-IR prompt pipeline: Hosted system turns free-form multimodal instructions into a structured intermediate representation before generation (API-only, not open-sourced). source
- 768P to 2K regeneration: H3-Regenerate-2K re-renders a 768P result with the original context into 2K (0.05 USD/s). source
MiniMax's current video generator (successor to Hailuo): text/image/video/audio-conditioned clips with sound, up to 2K. Async task API (create task, then query by task_id).
Sources: https://platform.minimax.io/docs/guides/models-intro · https://platform.minimax.io/docs/guides/pricing-paygo · https://huggingface.co/MiniMaxAI/MiniMax-H3 · https://platform.minimax.io/docs/release-notes/models
Other MiniMax models
MiniMax Music 3.0 · MiniMax-M3 · MiniMax-M2.7 · MiniMax Speech 2.8 (HD / Turbo)