🤖 POCKET-35B: Efficient MoE Model for GPU-less Operation
POCKET-35B has been introduced—a 35 billion parameter Mixture-of-Experts (MoE) model developed by VIDRAFT based on Darwin-36B-Opus. It is optimized for CPU-only operation on smartphones and standard PCs using llama.cpp, Ollama, or LM Studio. Thanks to its MoE architecture, the model achieves speeds of 27.0 tok/s on Xeon processors.
🌍 The emergence of efficient MoE models of this scale running on consumer hardware lowers the barrier to entry for local deployment of powerful AI and stimulates the development of On-Device AI.
👤 It is now possible to run an advanced 35B-parameter language model directly on your smartphone or home computer without purchasing an expensive graphics card.
Source 1: https://huggingface.co/FINAL-Bench/POCKET-35B-GGUF
