Apple refreshed its desktop lineup on August 25 with three new machines that change the math for anyone running language models locally. The M6 Mac mini, the M5 Max Mac Studio, and the M5 Ultra Mac Studio join a growing field of 128GB Strix Halo mini PCs from AMD partners. The result is a market where your decision comes down to two numbers: how much memory you need and how fast you can stream it.

Both of those numbers matter, and neither one is the CPU core count. Memory size determines whether a model loads at all. A 27B parameter model quantized to Q4 takes about 16GB on disk and needs slightly more in RAM. A 120B-class mixture-of-experts model at the same quantization level demands 64 to 80GB. The frontier models that launched this month scale into the hundreds of gigabytes at native precision, which is the entire reason the 512GB Mac Studio exists.

Memory bandwidth determines how fast tokens come out once the model is loaded. Generation on unified-memory hardware is bandwidth-bound. Every token requires streaming the active weights through the memory bus once. Double the bandwidth on a dense model and you roughly double the tokens per second.

Why the M6 Mac Mini Misses the Mark

The base M6 Mac mini ships with 16GB of unified memory at 153 GB/s. The 24GB and 32GB configurations push bandwidth to 170 GB/s. That is about two-thirds of a Strix Halo box and roughly a fifth of the older M3 Ultra Mac Studio. The machine can run Qwen3.8 27B at Q4 inside a 32GB configuration with room to spare, but generation speed will be slow. Expect around 6 tokens per second compared to the 28.8 that the M3 Ultra produces. Anything much larger than a 30B-class model will not load at all.

The M5 Pro Mac mini is the version of this box that matters for local inference. Same five-by-five footprint, 307 GB/s of bandwidth, and it configures up to 64GB. That capacity handles the 27B dense models at Q8 and the 35B and 80B mixture-of-experts models at Q4, at speeds that beat the Strix Halo boxes on dense workloads. Amazon lists it from $1,669.99, Apple's base configuration is 24GB, and the 64GB build is a configure-to-order option.

If you want an M6 Mac mini for general desktop use, the base model at $879.99 runs 8B-class models without trouble. Just do not buy it expecting 27B and above to perform well. The memory bus is too narrow.

The Real Fight: M5 Max Mac Studio versus 128GB Strix Halo

This is the tier where most people asking about local LLM hardware should land, and the choice is closer than either camp wants to admit. The 40-core GPU version of the M5 Max Mac Studio delivers 614 GB/s of bandwidth with 64GB of unified memory for $3,499 from Apple. The same chip with 128GB costs $5,099.

The Strix Halo boxes from GMKtec, BOSGAME, and Minisforum offer 128GB of LPDDR5X-8000 memory at roughly 256 GB/s for $3,500 to $3,800. For the same money as the Mac's 64GB configuration, you get twice the memory. At 128GB, you can load the 120B-class mixture-of-experts models at Q4 and the 200B-class ones at aggressive quantizations. You also get Linux, native Docker support, and a machine that behaves like a server.

The Mac's advantage is the memory bus. At 614 GB/s, the M5 Max has close to two and a half times the bandwidth. On dense models, that turns directly into faster generation speed. The Mac Studio is also quieter, and macOS is where the MLX framework lives, which matters more with each passing month.

The decision splits cleanly. If your workload is one large dense model and you want it fast, the M5 Max at 64GB (or 128GB if the budget allows) is the right machine. If your workload is loading the biggest mixture-of-experts model that fits and running an agent on it overnight, the 128GB Strix Halo box wins on value. The 64GB Strix Halo variant at $2,199.99 is worth considering if half the memory still covers your needs.

Who Actually Needs the M5 Ultra

The M5 Ultra Mac Studio starts at $5,499 with 96GB of unified memory and 1.2 TB/s of bandwidth. It is Apple's first quad-die design, and the 256GB and 512GB configurations are where the launch-day attention landed. Deliveries start September 22, and the 512GB option arrives in late October.

Running 256GB of unified memory as a daily machine means you stop asking whether a model will load. The 27B, 35B, 80B, and 120B-class models all load on the first attempt. When a frontier model in the hundreds-of-gigabytes range shows up, the only decision is which quantization level to use.

The M5 Ultra's improvement over the older M3 Ultra is the bus. Fifty percent more bandwidth should produce something close to 50 percent more tokens per second on dense models, and better throughput on the large mixture-of-experts models where the M3 Ultra starts to feel slow. The 512GB configuration keeps the ceiling the M3 Ultra set in 2025, where frontier models run at native or near-native precision, and puts the faster bus under it.

Who should buy it: people running frontier-class models locally for work, and people who cluster machines together. Apple claims up to three times the distributed inference performance over Thunderbolt 5 compared to a single machine, aimed at exactly that buyer. Who should not: anyone whose largest model fits in 128GB, which is a $3,500 decision.

Gorgon Halo: What AMD Has Coming

AMD's Ryzen AI Max+ PRO 495, the successor to the current Strix Halo, showed up at IFA in early September. Minisforum's announcement puts the numbers at 192GB of memory at 8533 MT/s, with up to 160GB usable as graphics memory, and 131 TOPS of AI performance. Chuwi and ACEMAGIC both showed machines with the same chip.

None of them have shipped yet, and the pricing is not encouraging. Minisforum's own teaser points at roughly 7,000 euros for the 192GB workstation, with sales expected to begin in September. Lenovo's ThinkCentre X Ultra with the same chip is announced at 3,100 euros for November. For context, the previous-generation N5 MAX with 128GB sells for $2,399.

At 7,000 euros, 192GB of Gorgon Halo costs as much as the M5 Ultra. The 128GB Strix Halo boxes remain the memory bargain. If a 192GB box appears near Lenovo's price, the M5 Ultra at 256GB faces a real competitor for the first time.

What to Buy at Each Budget

  • Under $1,000: A Ryzen 7 6800H box with 32GB, or the base M6 mini if it must be a Mac. The 16GB M6 stays under 20B models; the 32GB build costs $1,299.
  • Around $2,000: The 64GB EVO-X2 for Linux, or the M5 Pro Mac mini at 48GB ($2,299 from Apple). The 64GB mini is $2,699.
  • $2,500 to $4,000: The 128GB Strix Halo box. Bandwidth or memory, and memory wins: the interesting models keep getting bigger and the speed gap on mixture-of-experts models is small.
  • $9,500 and up: M5 Ultra at 256GB. The $5,499 base is 96GB. The 256GB step adds $4,000, and that is the configuration that justifies the spend.

Is the M3 Ultra Still Worth Buying Used

Yes, if the price has dropped enough. Bandwidth sits at roughly two-thirds of the M5 Ultra, and the 256GB configuration runs everything that has been thrown at it. The used market should shift after September 22 when new units start arriving. One more thing: Ollama on any of these machines generates tokens at about half the speed of llama-bench with the same weights. Any comparison you read using Ollama understates what the hardware can do. The real gap between these machines is probably wider than the benchmark tables show.

More memory than you think you need is the right amount.