Apple unveiled a new Mac Studio with M5 Max and a brand-new M5 Ultra chips and opened preorders on August 25, 2026. The company has for the first time built neural accelerators into the GPU of its flagship Ultra chip and is directly positioning the station as a desktop platform for running large open-weight LLMs locally.

image
image

What happened

The new Mac Studio is available in base configurations with the M5 Max chip and in higher-end configurations with the brand-new M5 Ultra. The flagship M5 Ultra is assembled from four dies using the UltraFusion scheme, features up to a 36-core CPU and an 80-core GPU, which for the first time in the Ultra series has built-in Neural Accelerators. Memory bandwidth increased by 50% compared to M3 Ultra, reaching 1.2 TB/s, and the unified memory capacity is up to 512 GB. The lineup adds PCIe Gen 6 SSDs and a new Core AI framework that works alongside MLX. Preorders open on August 25, 2026, in 30 countries: the configuration with 96 GB of memory costs $5,499, the 256 GB version — $9,499, the 512 GB version is promised for late October, and shipments begin on September 22. The base Mac Studio with M5 Max kept its $2,499 price.

Context

Apple's Ultra chip lineup has been assembled from a pair of dies since the previous generation; the four-die UltraFusion configuration of the M5 Ultra shifts the performance bottleneck from an individual die to the inter-die connection, so the chip's actual computing power will largely be determined by the efficiency of UltraFusion. Historically, running large language models locally was limited not by the number of cores, but by memory: in configurations with discrete video memory, the model simply didn't fit in VRAM, whereas Apple's unified memory of up to 512 GB with 1.2 TB/s bandwidth for the first time makes local inference of open-weight models with hundreds of billions of parameters practically meaningful in a single desktop chassis. The announcement came at a time of rising cloud inference prices and sustained demand for private local LLM runs without per-token payments; an ecosystem around LM Studio and MLX has already formed in this area.

Why this matters for the industry

For the industry, this is a platform shift, not a routine spec update: Apple is for the first time offering a mass-produced desktop on which frontier-class open-weight models can be run locally, and is promoting clustering of up to four Mac Studios via Thunderbolt 5 and RDMA with a shared memory pool as a replacement for a small inference cluster without data center infrastructure. The Core AI framework expands the software stack for local AI alongside MLX. If the claimed metrics are confirmed by independent measurements, companies will have a specific product with a price for calculating TCO against the cloud and, as a result, price pressure on token costs from cloud providers; developers already have reason to adapt their MLX and LM Studio stacks for M5 Ultra.

Why this matters for users

The key point for the reader is accessibility. For $5,499 in a 96 GB configuration, you can already get a machine with 1.2 TB/s memory bandwidth, and the option up to 512 GB means that models with hundreds of billions of parameters for the first time fit entirely in the memory of a desktop computer. This is private inference without per-token payments and without sending data to the cloud, and the base Mac Studio with M5 Max, which stayed at $2,499, makes upgrading the lineup accessible from a low threshold. Physical machines are not yet available: this is the preorder stage, shipments start on September 22, and the 512 GB version will only appear in late October, so now is the right time to plan pilot stands, not production purchases.

What is still unknown / limitations

All quantitative performance claims are Apple's official data in the "up to" format (up to 4.3x on AI tasks compared to M3 Ultra, up to 4x on prompt processing in LM Studio, up to 4.3x on image generation, up to 3x for a cluster of four Mac Studios), i.e., best-case scenarios without methodology or average values. There are no independent measurements yet: there is no third-party data on tokens/s for specific LLMs, power consumption, and price/performance ratio compared to discrete GPUs. The 512 GB configuration is not yet on sale and is only promised for late October, so the claimed scenario of running large models in full is currently unavailable in any shipping version.

Sources

Author

Look at AI, editorial team