🛠️ NVIDIA Nemotron 3.5 Lightning — 30B/3B MoE model for AI agents

30 billion parameters, 3 billion active per token. Hybrid Mamba-2 + Transformer architecture, context up to 1 million tokens, runs on a single GPU.

🌍 NeMo Switchyard — routing between models, up to 74% savings.

👤 On Hugging Face. 86% PinchBench, 52.8% SWE-bench. vLLM, Ollama.

Source: https://blogs.nvidia.com/blog/nemotron-lightning-switchyard-rtx-dgx/