Tuesday, August 25, 2026

Mac Studio 2026

Apple (Hacker News, MacRumors):

Mac Studio with M5 Max features an 18-core CPU, an up-to-40-core GPU with Neural Accelerators built into each core, and up to 128GB of unified memory, accelerating complex pro and AI workloads. With the powerful M5 Ultra, Mac Studio scales up to a 36-core CPU, up to an 80-core GPU, and a staggering 512GB of unified memory, enabling users to run enormous LLMs entirely on device. Wi-Fi 7 and Bluetooth 6 come to Mac Studio for the first time, while Thunderbolt 5 rounds out its extensive connectivity, so users can take advantage of blazing-fast external storage, PCIe expansion chassis, and powerful hub solutions for the most intense workloads. Thunderbolt 5 also enables multiple Mac Studio systems to be clustered, bringing up to 3x faster performance for distributed AI inference when compared to a single system.

[…]

Mac Studio with M5 Max starts at $2,499 (U.S.) and $2,299 (U.S.) for education.

[…]

Mac Studio with M5 Ultra starts at $5,499 (U.S.) and $5,099 (U.S.) for education.

It’s not shipping until September 22, with the 512 GB configuration in late October.

John Gruber:

Here’s my attempt to put all of the RAM/SSD configurations into condensed tables, so you can see which storage and memory options are available for each chip, and how much they cost.

[…]

It kind of stinks that there are no RAM options for the Studio between the 96 GB base and the $4,000 256 GB upgrade.

Jason Snell:

The Mac Studio is effectively the replacement for the Mac Pro, and it’s alone in offering the Ultra-class chip. Apple has positioned the Mac Studio as ideal for local AI workflows, and the new models will still be able to cluster via Thunderbolt 5 to create even larger collections of memory and performance. Apple representatives pointed out that a four-Mac Studio AI cluster, like the one I saw on display at WWDC earlier this summer, is so efficient that it can be powered from a single standard wall outlet. (They even showed us a picture of four Mac Studios plugged into a power strip that was plugged into the wall.) In an era of data center excesses, Apple is clearly leaning into the possibility of small, efficient Macs being used to do local AI work rather than relying on huge, power-hungry cloud models.

Federico Viticci:

If these numbers hold up and scale linearly, a local Mixture-of-Experts model such as Qwen 3.5-35B-A3B, which would run at ~17 tokens/sec on average on a base M4 Mac mini with 16 GB of RAM, could realistically generate output at over 60 tokens/second with the base model M6 Mac mini.

Based on what we’ve seen so far, the one downside of the M6 Mac mini is that it does not have Thunderbolt 5 ports; those are exclusive to the M5 Pro model, also announced today.

[…]

Speaking from personal experience, I know that my M3 Ultra Mac Studio using oMLX can run DeepSeek-V4-Flash locally with generation averaging 35 tokens/second. Assuming a linear 4x increase, that would put the same model at over 120 tokens/second on an M5 Ultra Mac Studio. To put things in perspective, that kind of performance would be faster than any AI chatbot website, it’d be faster than many providers who offer a “fast” mode for their models, and it’d only be second to either dedicated NVIDIA PC clusters at home or specialized inference providers such as Cerebras or Groq…which are running in full-blown data centers. Sure, you would need a computer that is likely going to cost more than $20,000 to make it happen, but it’d still be possible on a single machine that is small, quiet, and that – in theory – any consumer can buy off the shelf.

Previously:

Comments RSS · Twitter · Mastodon

Leave a Comment