Mac Studio M5 Ultra 96GB vs M5 Max 128GB for local LLMs?
I'm about to buy a Mac Studio mainly for running LLMs locally and I'm stuck between two configs: M5 Ultra (30/64) with 96GB : 1.2 TB/s bandwidth, roughly 1.7x faster generation and much faster prefill M5 Max (40-core GPU) with 128GB : 614 GB/s, but 32GB more memory and a bit cheaper The models I care about most right now are Qwen 3.8 27B and Qwen3.8-Flash-Next. The 27B fits easily on both, so the real question is Flash-Next. With the n-gram table offloaded to SSD, it seems to fit on 96GB, but only with the leanest 4-bit builds and very little headroom left for macOS. On 128GB you get more room for better quants, longer context or a second model loaded at the same time. A few questions for those who already made the call: Which one did you go for, and do you regret it? If you're running Flash-Next on a 96GB machine, how's it working in practice? Any issues with memory pressure or long contexts? Is the speed of the Ultra worth giving up the extra memory, or will 96GB feel tight? Thanks!