๐ค September State of Local AI: What Works on Local Hardware
Base Compute released its monthly report on local inference: which models run on what hardware and at what speed. The focus is on Qwen3.8-27B (Apache 2.0, vision-language, 262K context, ~17 GB in Q4), plus GLM-5.3-Flash (MIT) and the Mac Studio M5 Ultra with an RDMA cluster delivering up to 3x on four machines.
๐ Model selection has come down to three constraints โ memory, bandwidth, and compute. Dense 27โ31B models in Q4 are displacing 70B models on 24 GB cards, while open MoE models deliver frontier quality on 128โ192 GB of unified memory โ a case for self-hosting over the cloud. Thresholds: chat 10โ30 tok/s, coding 30+, agents 50+.
๐ค A practical guide to building local AI: from Gemma 4 E4B on 4 GB of RAM to GLM-5.3-Flash (~120 GB in 3-bit), with prices ranging from a laptop to on-prem setups at $150k+. Updated monthly; this is a local inference vendor, so key figures should be verified against the model card.
Source 1: https://www.basecompute.co/stateoflocalai
