The new GLM-5.2-Vision-NVFP4 multimodal model has been introduced, which combines the GLM-5.2 text backbone with the MoonViT visual encoder from Kimi-K2.6 and is optimized for operation on the latest NVIDIA Blackwell hardware.

What Happened
Developers have released the GLM-5.2-Vision-NVFP4 model, which utilizes the NVFP4 format and a custom PatchMerger projector. This allows for processing high-resolution images (up to 4096 tokens per image) and supports an ultra-long context of up to 1 million tokens.
Context
This release marks a transition from general-purpose multimodal models to specialized solutions that are deeply integrated with specific hardware through new data formats and modality alignment mechanisms.
Why It Matters for the Industry
For the industry, this is a signal of software readiness to work with the NVFP4 format. Optimization for the Blackwell architecture allows for significantly increased inference speed and reduced computational costs when processing high-resolution visual context.
Why It Matters for Users
Users gain access to a tool capable of performing deep analysis of massive amounts of visual-textual data. This opens up possibilities for creating high-performance multimodal agents capable of real-time operation with minimal latency.
Sources
Author
Look at AI, Editorial Team
