MagicMakeup has been introduced—a specialized framework based on Diffusion Transformer (DiT) for the most accurate transfer of makeup from a reference image to a face. Built upon the FLUX.1-Kontext-dev foundation, the model allows for detailed editing while maintaining complete identity of the user's facial features.

image
image

What Happened

The MagicMakeup framework has been developed, supporting 1024x1024 resolution and providing precise regional control over eyes, lips, or the entire face. The solution is based on two key modules: TARG (Token-Aligned Region Gating) to prevent makeup from bleeding outside target zones, and CMPG (Cross-Modal Perception Guidance) to decouple the style transfer process from the preservation of anatomical facial features.

Context

Modern face editing methods in image-to-image (i2i) tasks often face problems with inaccurate positioning ("bleeding" effects) and the distortion of person identity during deep style changes. The use of the DiT architecture and token control mechanisms in MagicMakeup is aimed at addressing these fundamental limitations.

Why It Matters for the Industry

For the digital beauty and video production industries, MagicMakeup sets a new standard for specialized i2i solutions. The emergence of a high-quality open-source reference based on FLUX.1 allows companies to create high-precision virtual try-on tools and integrate advanced regional control mechanisms into existing generative media pipelines.

Why It Matters for Users

Content creators, artists, and beauty app developers gain a tool that allows them to virtually "try on" any makeup on a photo while maintaining full recognizability. This opens up possibilities for creating high-quality prototypes for cosmetic virtual try-on and digital makeup applications.

What Is Not Yet Known / Limitations

There are concerns regarding computational complexity and inference costs due to the use of the FLUX.1 architecture, which may limit the model's applicability in real-time scenarios without the creation of lightweight (distilled) versions.

Sources

Author

Look at AI, Editorial Team