A specialized LoRA, LTX-Best-Face-ID, has been introduced for the LTX-2.3 (22B) video model, allowing for effective preservation of character appearance during video generation. The technology is based on the Reference-to-Video method and a specialized TASS-RoPE mechanism, ensuring high accuracy in transferring facial features and physique.
What Happened
The developer released LTX-Best-Face-ID—a solution for the LTX-2.3 (22B) model designed to address the problem of character consistency. The technology utilizes the injection of reference latents into the RoPE grid of the first frame and a TASS-RoPE mechanism to separate reference data from the generation process. To ensure accuracy, ArcFace identity loss is applied. Users have access to two modes: Face ID Base for working with close-up shots of faces, and Character-Sheet, which allows for preserving clothing and physique through the use of four-panel reference images.
Context
The problem of character consistency is one of the key challenges in the generative video industry. LTX-Best-Face-ID offers a transition from the labor-intensive process of full fine-tuning of large models to a rapid 'reference-to-video' workflow using LoRA, significantly lowering the technical barrier to entry for personalizing video content.
Why It Matters for the Industry
For the AI video industry, this solution offers an efficient way to integrate identity control into existing pipelines. The use of ArcFace loss and the adaptation of Rotary Positional Embedding (RoPE) minimizes artifacts during facial feature transfer, which is critical for professional production and storytelling. In the long term, similar methods may become standard in video model architectures.
Why It Matters for Users
Regular users and content creators can now create videos featuring specific people simply by uploading their photos. This opens up possibilities for rapid video prototyping, creating advertising materials with consistent characters, and developing full series or digital avatars without the need for massive computational resources required for model training.
What Is Not Yet Known / Limitations
Data regarding latency and the exact inference cost when using a 22B-scale model is missing, which is important to consider when planning industrial deployment. There are also security concerns related to the use of biometric data.
Sources
Author
Look at AI, Editorial Team