ByteDance has introduced SeedVR2-1.4B — a compact version of its SOTA upscaler SeedVR2-7B. Thanks to the application of 6-layer distillation, the new model is 5.7 times smaller than the original in terms of parameter count, but it retains high edge sharpness and demonstrates a significant increase in operating speed.

image
image
image

What happened

The SeedVR2-1.4B model, containing 1.44 billion parameters, has been released. When scaling an image from 512 to 2048 pixels, the model consumes about 4.6 GB of video memory (VRAM). In x8 upscaling mode, the model's performance is 4.7–5.6 times higher than the base SeedVR2-7B version.

Context

The technology is based on deep distillation of diffusion models, during which the architecture is reduced from 36 to 6 layers. This allows transferring the functionality of heavy SOTA solutions, which initially required significant computational resources, to the category of compact tools for edge devices.

Why this is important for the industry

For the industry, this is an important step in the optimization of diffusion architectures. The successful transfer of SOTA capabilities to models with a small number of parameters lowers the barrier to entry for implementing quality upscaling in mobile applications and low-end devices. This opens the way to the mass use of quantized models on consumer hardware.

Why this is important for users

Users with video cards of small VRAM volume (from 6-8 GB) or owners of powerful smartphones can now use quality upscaling locally, without relying on cloud services. The model is also ready for integration into local workflows, such as ComfyUI, providing high image sharpness with minimal delays.

Sources

Author

Look at AI, editorial team