LTX, a company spun out of Lightricks, released the LTX-2.5 video and world model with 22 billion parameters under an open-weights license. The model generates 10-second 720p clips in 6.8 seconds on two NVIDIA GB200 GPUs, which is 7.6–10.4 times faster than the proprietary Gemini Omni Flash and Veo 3.1. The distilled version runs on GPUs with 16 GB VRAM or more, and ComfyUI integration is available from day one.


What happened
The LTX-2.5 release includes a number of architectural changes. Instead of traditional VAE reconstruction, a diffusion video decoder is used — a non-trivial replacement aimed at eliminating artifacts and quality bottlenecks in video diffusion. Text is processed by a custom text encoder based on Gemma 4 12B, moving away from standard CLIP and T5 approaches. The model supports native multi-frame generation, a prompt enhancer, text-to-video, image-to-video, and synchronized audio generation in a single pass. The distilled version requires 16 GB VRAM or more. The community has already released an NVFP4-quantized version by BennyDaBall — 18.7 GB instead of 21.5 GB INT8, running on Blackwell and RTX 50-series. Lightricks also published the official IC-LoRA Pixel Spatial Upscaler x2 for generative resolution enhancement with detail synthesis.
Context
LTX is a company spun out of Lightricks, a well-known developer of mobile tools for working with photos and videos. The launch of LTX-2.5 as an open-weights model with a license free for companies with revenue up to $10 million ARR positions video generation as basic infrastructure rather than a competitive advantage. LTX CEO Zeev Farbman described ComfyUI integration as the main channel for customer acquisition. The model is available on Hugging Face, and the community has quickly adapted it: quantization, specialized LoRA, integration into existing pipelines — all of this is already in progress.
Why this matters for the industry
A speed of 6.8 seconds on 2× GB200 versus 52–70 seconds for Gemini Omni Flash and Veo 3.1 is a shift in the performance baseline in favor of open-source solutions. Replacing VAE with a diffusion decoder could become a winning approach for an entire generation of video models. An open-weights license up to $10 million ARR allows startups to launch video generation MVPs without API costs, reducing time-to-market from months to weeks. The architectural pattern of a single model for video and audio, as well as day-one ComfyUI support, create the conditions for forming an ecosystem similar to Stable Diffusion: LoRA for styles, custom encoders, integrations into existing creative tools. However, dependence on Blackwell architecture for maximum results creates hardware lock-in.
Why this matters for users
The model is available for download on Hugging Face. The distilled version can be run locally on a GPU with 16 GB VRAM or more. On RTX 50-series and Blackwell, the NVFP4 quantization by BennyDaBall works, reducing memory volume to 18.7 GB. The official IC-LoRA Upscaler allows generative resolution enhancement rather than interpolation. ComfyUI integration works from day one, providing ready-made workflows without the need to write code. Supported modes: text-to-video, image-to-video, multi-frame scenes, and synchronized audio.
What is still unknown / limitations
A technical publication on LTX-2.5 is absent. Video quality metrics — FVD, FID, human evaluations — have not been published. The stated speed of 6.8 seconds was measured on two NVIDIA GB200 GPUs, and the benchmark is reproducible only on Blackwell architecture, making the comparison with Gemini Omni Flash and Veo 3.1 methodologically weak — inference conditions may differ. The architecture of the custom text encoder based on Gemma 4 12B is not described: it is unknown whether it was trained jointly with the main model or frozen. The trade-off between speed and quality when replacing VAE with a diffusion decoder is not yet documented. Claims about production readiness are premature without reproducible metrics.
Sources
- Lightricks/LTX-2.5 — Hugging Face
- Lightricks/LTX-2.5-22b-IC-LoRA-Pixel-Spatial-Upscaler — Hugging Face
- LTX-2.5: Open Weights, 6.8-Second Video, ComfyUI Day One — TL Dev Tech
Author
Look at AI, editorial team
