🤖 Liquid AI Releases LFM2.5-VL-3B — An Open Vision-Language Model for Devices
The model is built on the LFM2.5-2.6B text base and the SigLIP2 400M NaFlex image encoder: 32,768-token context, 16 languages including Russian, ~3 GB of memory, and responses without a chain of thought.
🌍 This is a new reference for on-device VLMs: the 3B class for the first time shows the level of 4B models in screen understanding (ScreenSpot-v2 80.7) and function calling (ToolSandbox 59.5) — useful for local UI agents working with interfaces without the cloud.
👤 You can try it right now: a WebGPU demo in the browser, the Liquid AI Playground, and GGUF, ONNX, and MLX builds for llama.cpp, LM Studio, Ollama, and Jan. On the Galaxy S26 Ultra, the model maintains ~20 tokens/sec.
Source 1: https://www.liquid.ai/blog/lfm2-5-vl-3b Source 2: https://huggingface.co/LiquidAI/LFM2.5-VL-3B
