🤖 Release of the GLM-5.2-Vision-NVFP4 Multimodal Model
A new model has been introduced that combines the GLM-5.2 text backbone with the MoonViT visual encoder from Kimi-K2.6. The model is optimized for the NVFP4 format to run on the NVIDIA Blackwell architecture and supports a context window of up to 1 million tokens.
🌍 It demonstrates the capabilities of deep integration between visual and textual components through specialized projectors and optimization for new data formats (NVFP4), which is critical for the next generation of VLMs on Blackwell hardware.
👤 It allows working with a massive context (1 million tokens), including detailed analysis of high-resolution images, using modern quantization methods to accelerate inference.
Source 1: https://huggingface.co/baseten/GLM-5.2-Vision-NVFP4
