🤖 🌟 Prism ML Releases Ternary Version of Qwen3.8-27B
Prism ML released Ternary Bonsai 2 27B on September 17, 2026 — a ternary (weights −1/0/+1) rework of Qwen3.8-27B: 5.9 GB instead of ~54 GB in FP16, and 98.2% of the average score across 14 benchmarks. Density — 1.72 bits/weight, context — 262k tokens, license Apache 2.0.
🌍 The 27B model has been compressed to the size of an 8B model: on the same hardware, this means more users per GPU and lower power consumption — 0.714 mWh/token on an RTX 4090, 40% more efficient than an 8B model. Prism ML's focus is private inference: code agents and documents without the cloud.
👤 The model with reasoning and vision capabilities takes up 5.95 GB: an RTX 5090 delivers ~143 tokens/s, M5 Max — ~46.8, and it works on iPhone/iPad via MLX. Nuance: a Prism ML fork of llama.cpp is required — without the built-in Hadamard transformation, the model outputs nonsense. A WebGPU demo is available.
Source 1: https://prismml.com/news/bonsai-2-27b Source 2: https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf
