LTX-2.5
Full (trainable) and distilled DiT variants; Gemma 4 12B text encoder; up to 4K; frame counts must satisfy n % 8 == 1; ComfyUI native. Licence bans military/weapons use.
- Input
- text, image, video
- Output
- video, audio
- License
- ltx-2-community-license (free under $10M annual revenue)
- Pricing
- fast per second: $0.09 · 4k per second: $0.3 (USD per second of output video on Lightricks' managed API (as reported by VentureBeat)) source
- Verified
- 2026-10-03
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Hugging Face | — | huggingface.co/Lightricks/LTX-2.5 | — |
| GitHub | — | github.com/Lightricks/LTX-2 | — |
| LTX API | — | console.ltx.io/playground/ | docs |
Notable capabilities (3)
- Open audio-video generation with multishot consistency: 22B DiT that generates synchronized video and audio and keeps character, environment, lighting and voice consistent across cuts. source
- Fast generation: 10-second 720p image-to-video clip in 6.8 s on two GB200s (vendor claim); distilled variant uses 8 steps. source
- Robotics / physical-AI checkpoint: Ships a pretrained checkpoint aimed at world-model use in robotics and physical AI. source
Lightricks' open-weights video+audio model; the most downloaded generative model on Hugging Face in early October 2026.
Timeline entry
- Lightricks releases LTX-2.5, a 22B open-weights audio-video 'world model' with multishot generation and up to 4K output ★★★
On Aug 11, 2026 Lightricks released LTX-2.5, a 22B-parameter open-weights model that generates synchronized video and audio from text, images or video, keeps characters and voices consistent across cuts (multishot), and ships a checkpoint for robotics and physical AI. Free for organisations under…