🤖 DeepSeek Releases V4.1 Flash
On September 10, DeepSeek introduced the smallest model in its new lineup with native image understanding: a 552B-parameter MoE, a Causal Encoder–Decoder with 8B active parameters on input and 16B on output, and up to 1M token context. The KV cache takes up 4 times less HBM and 8 times less SSD than the previous generation.
🌍 In the DeepSeek API, it is available as deepseek-flash and is replacing the flagship: from 04:00 UTC on September 14, all deepseek-v4-pro requests will be routed to it until V4.1-Pro is released. Off-peak prices are half of peak prices, weights are open on Hugging Face, and there is a vLLM recipe for self-hosting.
👤 A single model handles dialogue, analyzes images, and automatically enables vision — no need to manually select a mode. The old deepseek-v4-flash and deepseek-v4-flash-vision-exp are temporarily routed to it.
Source 1: https://www.deepseek.com/en/news/deepseek-v4-1-flash/
Source 2: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
