The M5 Ultra Mac Studio did not arrive at WWDC 2026. What happened to the Mac Studio while everyone waited matters more. Apple quietly cut the machine’s memory ceiling from 512GB to 96GB in two steps, then raised the M3 Ultra’s price from $3,999 to $5,299. On Apple’s own tech specs page today, the M3 Ultra lists 96GB of unified memory with no upgrade option at all.
That reverses the usual advice. This is not a story about a faster chip being worth the wait; it is a story about a memory ceiling that fell by 81% on the machine you can buy right now. Below we cover what the Mac Studio offers today and every credible M5 Ultra leak, including the 768GB figure Apple is reported to have tested. We also measure what actually fits in 96GB, and explain why the next Ultra chip is not due until 2028.
The Key Takeaways
- 96GB is now the only memory configuration on the M3 Ultra Mac Studio. The 512GB option went in March 2026 and the 256GB option went on 5 May 2026
- The M3 Ultra Mac Studio now starts at $5,299 (up from $3,999) and the M4 Max at $2,499 (up from $1,999) after Apple’s 25 June price rise
- No M5 Ultra exists yet. WWDC 2026 passed without it; reporting now points to around October 2026
- Apple has reportedly tested up to 768GB of unified memory for the M5 Ultra, but supply constraints may stop that option shipping
- There will be no M6 Ultra. The next Ultra after the M5 is the M7 Ultra in 2028, so this is the only Ultra upgrade for roughly two years
Where the M5 Ultra Mac Studio Actually Stands
Apple has not announced the M5 Ultra Mac Studio. That is not a hedge, it is verifiable: search Apple’s Mac Studio tech specs, its Mac Studio marketing page and its store page, and M5 Ultra appears zero times on all three. The same pages return the M3 Ultra and M4 Max repeatedly, so the absence is real rather than a page that failed to load.
WWDC 2026 opened on 8 June 2026 and brought macOS 27 Golden Gate, not a new Mac Studio. The current window is around October 2026. AppleInsider reported on 28 June that the M5 Ultra is still due later in 2026, and Bloomberg’s Mark Gurman has pointed at October since April, with the global DRAM shortage as the reason.
Two details from that reporting matter for buyers. Apple is not redesigning the box, so expect the same chassis and ports. And it has reworked the inside, chiefly a better heatsink to hold clocks up under sustained load, which is exactly the workload profile of a long local-inference run.
The Mac Studio’s Memory Ceiling Fell From 512GB to 96GB
This is the most important thing on this page, and it happened in two steps that were reported separately, so it is easy to miss. When Apple launched the M3 Ultra Mac Studio it advertised up to 512GB of unified memory, an unusual amount for a desktop and the reason the machine became popular for local AI. Both configurations above the 96GB base are now gone.
| Date | What changed | Max unified memory |
|---|---|---|
| March 2025 | M3 Ultra Mac Studio launches, advertised at up to 512GB | 512GB |
| 5 March 2026 | 512GB option removed. It had cost $4,000. The 256GB upgrade rises from $1,600 to $2,000 | 256GB |
| 5 May 2026 | 256GB option removed | 96GB |
| 25 June 2026 | Apple raises Mac prices across the line | 96GB |
MacRumors reported the first cut on the day the 512GB option disappeared, along with the $1,600 to $2,000 increase on the 256GB upgrade. Two months later 9to5Mac reported the last remaining upgrade going too, and put it plainly.
This leaves the M3 Ultra Mac Studio with a single memory configuration: 96GB.
Apple’s own specs page confirms it in a way that is hard to argue with. The M4 Max row reads 36GB unified memory, configurable to 64GB. The M3 Ultra row reads 96GB unified memory and then stops, with no configurable-to line at all. Apple’s Mac comparison page lists the Mac Studio memory ladder as 36GB, 64GB, 96GB and describes the machine as from 96GB unified.
Apple has not given a reason. The timing matches the global memory shortage that also drove the 25 June price rise across every Mac. Both 9to5Mac and MacRumors point at AI server demand for DRAM as the cause.
What the Mac Studio Costs Now
Every price here was read from Apple’s live store SKU data, not from coverage. That matters: one widely cited 2026 Mac Studio guide was updated on 22 July and still lists $1,999 and $3,999.
| Configuration | Memory | Storage | Price |
|---|---|---|---|
| M4 Max, 14-core CPU, 32-core GPU | 36GB | 512GB | $2,499 |
| M3 Ultra, 28-core CPU, 60-core GPU | 96GB | 1TB | $5,299 |
The M4 Max configures up to a 16-core CPU and 40-core GPU with 64GB of memory. The M3 Ultra goes up to a 32-core CPU and 80-core GPU. Both are paid upgrades on top of those starting prices, and the memory ladder stops at 96GB no matter which you pick.
Put the price and the memory change side by side and the shape of 2026 is clear. The entry M3 Ultra costs $1,300 more than it did, and the 512GB upgrade that used to sit above it is not available at any price. That is the context for any decision about waiting.
You pay more for less headroom. That is the whole story of this machine in 2026.
Leaked M5 Ultra Specs
The leaks are consistent on shape. The M5 Ultra is expected to fuse two M5 Max dies with Apple’s UltraFusion interconnect, the approach used on every Ultra chip since the M1. MacRumors and Macworld both put it at around 36 CPU cores and up to 80 GPU cores.
| Chip | CPU cores | GPU cores | Memory (max) | Bandwidth | AI hardware | Status |
|---|---|---|---|---|---|---|
| M3 Ultra (Mac Studio) | up to 32 | up to 80 | 96GB | 819 GB/s | 32-core Neural Engine | Shipping |
| M4 Max (Mac Studio) | up to 16 | up to 40 | 64GB | 546 GB/s | 16-core Neural Engine | Shipping |
| M5 Max (MacBook Pro) | up to 16 | up to 40 | 128GB | 614 GB/s | Neural Accelerator in every GPU core | Shipping |
| M5 Ultra (Mac Studio) | ~36 | 80 | 768GB tested, unconfirmed | not disclosed | Neural Accelerator in every GPU core | Unreleased |
Two honest caveats on that bottom row. Nobody has published an M5 Ultra memory bandwidth figure, so we are not printing one; doubling the M5 Max’s measured 614 GB/s is the obvious estimate but it is an estimate, not a leak. And the core counts come from analyst reporting, not from Apple.
The real generational gain is not core count, since the M3 Ultra already reaches 80 GPU cores. It is that the M5 generation puts a Neural Accelerator inside every GPU core, a matrix-multiply unit sitting next to the shader hardware instead of a separate Neural Engine block off to the side. For LLM inference, which is mostly large matrix multiplies, that is where the speedup comes from.
There was never an M4 Ultra. Apple skipped that generation entirely, which is why the current top-end Mac Studio still runs a chip introduced in 2025 alongside a much newer M4 Max.
768GB Is the Number That Matters
On 25 June, MacRumors reported the specification that would change everything about this machine for local AI, and it is worth quoting exactly because the caveat is as important as the number.
Apple has tested support for up to 768GB of unified memory, but supply constraints could prevent it from launching with an option for that much memory.
Read that second clause twice. It is the difference between a landmark machine and a footnote.
That is eight times today’s 96GB ceiling. If it ships, the M5 Ultra Mac Studio becomes the most capable local-inference desktop Apple has ever sold by a wide margin. If it does not, the M5 Ultra is a faster chip attached to the same 96GB wall, and for large-model work that would be a modest upgrade dressed up as a big one.
The same report notes the obvious consequence for price: eight times the memory during a memory shortage could push a loaded Mac Studio past $10,000. Treat 768GB as a possibility Apple has tested in the lab, not a specification you can plan a budget around.
What You Can Actually Run on 96GB Today
Most coverage of this machine quotes model sizes from memory. The table below is measured: each figure is the actual on-disk size of the MLX build of that model, summed from the weight files themselves. Weights are not the whole story, because you also need room for the KV cache and for macOS, so budget roughly 75-80GB usable out of 96GB.
| Model (MLX build) | Measured weights | Fits in 96GB? |
|---|---|---|
| Qwen3.5 9B, 4-bit | 6.0 GB | Yes, trivially |
| DeepSeek R1 Distill 14B, 4-bit | 8.3 GB | Yes, trivially |
| Gemma 4 26B A4B, 4-bit | 15.6 GB | Yes, with room for several more |
| Qwen3.6 27B, 4-bit | 16.1 GB | Yes, with room for several more |
| Qwen3.6 27B, 8-bit | 29.5 GB | Yes, comfortably |
| Llama 3.3 70B, 4-bit | 39.7 GB | Yes, comfortably |
| Qwen3.6 27B, full precision | 55.6 GB | Yes, tight |
| gpt-oss 120B, MXFP4 | 63.4 GB | Yes, but it is the ceiling |
So a 96GB Mac Studio remains a capable local-AI machine. It handles anything up to roughly the 70B dense class at 4-bit, a 27B at full precision, or several mid-size models loaded at once for an agent pipeline. That covers most practical local work, and it is the honest case for buying one today.
What it no longer covers is the top of the open-weights world, which is the specific thing the 512GB configuration existed for. We go deeper on model selection in our guide to the best open-source AI models.
The Frontier Open Models No Longer Fit Any Mac
This is where the Mac Studio’s story changed underneath it. Two years ago the largest open-weight models were in the 70B range and a big-memory Mac was the cheapest way to run one. The frontier has since moved to trillion-parameter mixture-of-experts models, and it moved much faster than Apple’s memory ceiling. These are measured download sizes from each model’s own repository.
| Open-weight model | Parameters | Weights on disk | Runs on a Mac Studio? |
|---|---|---|---|
| Kimi K3 | 2.8T total, 104B active | 1,561 GB | 아니요 |
| GLM-5.2 | 744B total, 40B active | 1,507 GB | 아니요 |
| DeepSeek V4 Pro | 1.6T total | 865 GB | 아니요 |
| Kimi K2.7 Code, 4-bit | not disclosed | 641 GB | 아니요 |
None of them fit. Not on a 96GB Mac Studio, and not on a 768GB one either.
Kimi K3’s weights alone are 1.56 TB. That is more than sixteen times a 96GB Mac Studio, and more than double the 768GB the M5 Ultra has reportedly been tested with. Even the most optimistic version of this machine does not run today’s best open model, and no amount of quantisation closes a gap that size. Our full breakdown of Kimi K3 covers what it takes to serve it.
This is worth being blunt about because the older framing on this page, that Apple Silicon wins by running the biggest models, is no longer true at the very top. What Apple Silicon still wins is the middle: cheap, quiet, private inference on capable mid-size models, with a unified memory pool that a consumer graphics card cannot match. That is a smaller claim and a defensible one.
How UltraFusion Works and Why It Matters for Local AI
Apple’s original M1 Ultra UltraFusion design used a silicon interposer to connect two M1 Max dies across more than 10,000 signals, delivering 2.5TB/s of bandwidth between them. The M5 Ultra is expected to use the same technique on two M5 Max dies.
The payoff for AI work is that software sees one chip with one memory pool, not two chips shuffling weights over a slow link. For a model that spans tens of gigabytes, that is the difference between loading it once and streaming pieces of it on every token.
Unified memory also removes the split between CPU RAM and GPU VRAM. On a workstation with an RTX 5090 you get 32GB of VRAM separate from system RAM, and anything that does not fit has to be offloaded or quantised harder. On a 96GB Mac Studio, CPU and GPU draw from the same pool at the same 819 GB/s. That is a three-times capacity advantage today. It was an eight-times advantage when 256GB was on the menu, which is precisely what Apple gave up.
Ports, Displays, and the Rest of the Box
The chassis is not expected to change, so today’s layout is the right baseline. On the back, every Mac Studio has four Thunderbolt 5 ports at up to 120Gb/s, two USB-A ports at 5Gb/s, HDMI 2.1, 10Gb Ethernet and a headphone jack. The front differs by chip, and it is the one place the two models are not equivalent.
| Front ports | M4 Max | M3 Ultra |
|---|---|---|
| Two ports | USB-C at 10Gb/s | Thunderbolt 5 at 120Gb/s |
| Card slot | SDXC (UHS-II) | SDXC (UHS-II) |
| Total Thunderbolt 5 ports | Four | Six |
Display support is where the Ultra earns its name. The M3 Ultra drives eight displays at 6K and 60Hz, or four at 8K and 60Hz. The M4 Max manages five.
What Apple’s Own Benchmarks Show About M5 AI Performance
Apple’s Machine Learning Research team published measured MLX benchmarks for the M5 in November 2025, and it is the best first-party evidence for what the Neural Accelerators do. On time-to-first-token, the M5 is up to 4x faster than the M4, with per-model speedups ranging from 3.33x to 4.06x across Qwen models at 1.7B, 8B, 14B and 30B mixture-of-experts, plus gpt-oss 20B.
Sustained token generation, which is bandwidth-bound rather than compute-bound, improved by a more modest 19-27%. Apple attributes that to memory bandwidth rising from 120 GB/s on the base M4 to 153 GB/s on the base M5, a 28% increase.
Prompt processing scales with compute. Generation scales with memory bandwidth.
That split is the useful lesson for a Mac Studio buyer. The Neural Accelerators make prompt processing dramatically faster, which you feel most on long documents and large codebases. Generation speed scales with bandwidth instead. That is why the M3 Ultra’s 819 GB/s still holds up against much newer chips. It is also why an M5 Ultra fusing two 614 GB/s dies should be a large step rather than an incremental one.
The Software Stack: MLX, Ollama, and LM Studio
The Apple Silicon software situation improved sharply through 2026, and three tools cover almost everyone.
MLX
MLX is Apple’s own machine-learning framework, written to exploit unified memory and the Neural Accelerators directly. It is the fastest path on Apple Silicon and the right choice if you are building a product on top of local inference rather than just running a model.
Ollama
Ollama is the easiest way to get a local API running: install, pull a model, done. It shipped an MLX backend in preview on 30 March 2026 and has iterated hard since, with its own engineering blog reporting its highest Apple Silicon performance yet in June.
Be careful with the speedup numbers attached to this, including some we have seen repeated as a blanket figure. Ollama’s own strongest claim is up to 90% faster when used with coding agents, measured on the Aider polyglot benchmark, and it refers to a specific multi-token-prediction optimisation. Independent testing shows the gains concentrate heavily in mixture-of-experts models and fall close to zero on some dense ones. Expect a large win on MoE, and test before assuming it applies to your model.
LM Studio
LM Studio is a GUI app for downloading, running and chatting with models, and it supports MLX too. If you have never run a local model, start here, move to Ollama when you want to call it from a script, and reach for MLX directly when you are optimising.
M5 Ultra vs RTX 5090 for Local AI
This comparison rotted in an interesting way: both sides got more expensive for the same reason. The memory shortage that cut Apple’s configurations and raised Mac prices also pushed GDDR7 costs up, and Nvidia reportedly raised board-partner pricing by around $300 in May 2026.
The RTX 5090 launched at a $1,999 list price and has not been reliably available at it for a long time. Current street pricing is contested. Trackers show figures from a median around $2,150 in June up to roughly $4,300 at large retailers, with premium partner cards above $5,000. We are giving you the spread rather than pretending there is one number.
On raw speed the 5090 wins for anything that fits in its 32GB of VRAM, and it pulls up to 575 watts doing it, before you add a CPU, motherboard, power supply and cooling. A Mac Studio is one box at a fraction of the power draw with the MLX stack already installed.
Then you hit the memory wall.
Memory Is Still the Real Difference
A 70B model at 4-bit is a measured 39.7GB. That does not fit on a single RTX 5090, so you run two cards, offload layers to system RAM and lose most of your speed, or quantise harder and lose quality. A 96GB Mac Studio loads it with room to spare, and holds a second model alongside it.
The honest rule: the 5090 for the fastest tokens per second on models that fit in 32GB, the Mac Studio for holding models that do not, quietly, on one desk. Note how much narrower that claim is than it was six months ago. At 256GB the Mac’s capacity advantage was eight to one; at 96GB it is three to one.
M5 Ultra vs M3 Ultra vs M5 Max: Who Should Wait
The buying decision now hinges on memory rather than on chip generation, and that changes the usual answer. Waiting has become more attractive, not less, because the machine on sale today is the weakest Mac Studio for large local models that Apple has offered since the M3 Ultra launched.
| If you are | Do this | Why |
|---|---|---|
| Running models above 70B locally | Wait for the M5 Ultra | 96GB cannot do this job at any price. 768GB might, if it ships |
| Running 27B to 70B models | Buy now, or wait if you can | A 96GB M3 Ultra handles this well today at $5,299 |
| Mostly doing video, 3D or general pro work | Buy the M4 Max at $2,499 | You are not memory-bound and the Ultra premium is $2,800 |
| Wanting portability | M5 Max MacBook Pro | Same per-core AI architecture, up to 128GB, more memory than the Ultra Studio |
| Starting out with local AI | Mac mini | A fraction of the price and enough for mid-size models |
| On an Intel Mac or M1 Ultra | Upgrade, and time it | macOS 27 drops Intel, so this is not optional, only a question of when |
Look at the fourth line in that table again. An M5 Max MacBook Pro configures up to 128GB of unified memory, which is more than the $5,299 M3 Ultra Mac Studio can reach. For memory-bound local inference, Apple’s laptop currently out-specs its desktop workstation. We break that machine down in our M5 MacBook Pro upgrade guide and compare the range in our best MacBook for AI guide. For the cheaper end, see our Mac mini for AI breakdown.
There Is No M6 Ultra, So the Next Wait Runs to 2028
If you decide to wait, it is worth knowing what you would be waiting for after this. Apple is skipping the M6 generation at the high end entirely. There will be no M6 Pro and no M6 Max, and consequently no M6 Ultra, since an Ultra chip is built by fusing two Max dies.
An Ultra needs two Max dies. No M6 Max means no M6 Ultra.
AppleInsider reported in June that the next Ultra after the M5 is the M7 Ultra, expected sometime in 2028. That reshapes the decision in a way most coverage misses. The M5 Ultra is not one step on a yearly ladder; it is the only Ultra-tier upgrade available for roughly two years.
So the choice is narrower than it looks. Buy a 96GB M3 Ultra now, take the M5 Ultra when it lands around October, or hold what you own until 2028. If your work is memory-bound, only the middle option makes sense, and it is worth planning around.
What We Still Don’t Know
A lot, and it is worth being explicit about which parts of this page are solid and which are not. Everything about the current Mac Studio, its prices, its 96GB ceiling, its ports and the measured model sizes, is verified against Apple’s own pages and the model repositories. Everything about the M5 Ultra is reporting.
- Whether 768GB ships at all. Apple has reportedly tested it; supply may kill it
- Memory bandwidth. No credible figure has been published for the M5 Ultra
- Price. Nothing has leaked, and the $5,299 starting point for the M3 Ultra makes older estimates in the $4,000 range look obsolete
- Timing. October 2026 is reporting, not a date, and this machine has already slipped once
- Whether the 96GB cut is permanent or a shortage measure that reverses when DRAM eases
Treat every line above as weather, not as a forecast you can budget against.
We have deliberately not printed an M5 Ultra price estimate or a tokens-per-second projection for hardware that does not exist. Where this page previously carried those, they were extrapolations, and the 96GB cut is a good demonstration of how badly extrapolation can age.
macOS 27 Golden Gate and the End of Intel
The M5 Ultra Mac Studio will ship with macOS 27 Golden Gate, announced at WWDC 2026 on 8 June and due in late 2026. It is the first version of macOS that runs only on Apple Silicon, and the last with full Rosetta 2. We cover it in detail in our macOS 27 Golden Gate guide.
Intel Macs end at macOS 26 Tahoe. If you are on a 2019 Mac Pro, a 2019 16-inch MacBook Pro or a 2020 iMac, that is the end of the line for new Mac AI features. Migrating to Apple Silicon has become a requirement rather than an upgrade.
One clarification worth making, because it is widely garbled. Apple’s advanced Apple Intelligence tier requires Mac models with M3 and later and at least 12GB of unified memory, and that bar gates two specific features, the expressive Siri voice and higher-accuracy on-device dictation. It does not gate Siri’s AI features as a whole, which run on every Apple Silicon Mac. Any Mac Studio clears the 12GB bar eight times over, so none of this is a consideration here.
Local Models on the Mac Studio, Frontier Models via Fello AI
A Mac Studio is excellent at private, unmetered inference on open-weight models. What no local machine can do is run the closed-weights frontier models, because those weights are never published. As the table above shows, it increasingly cannot run the biggest open models either.
That is the gap Fello AI fills. It runs natively on Apple Silicon and puts the leading frontier assistants behind one menu-bar app for $9.99 a month, so you are not holding several separate subscriptions to compare answers. The split works cleanly: open models run locally on your Mac through MLX or Ollama, and frontier models route through Fello AI when a task needs the best available quality. Fello AI holds a 4.7-star rating across 27,000+ reviews.
If you would rather install individual clients, we maintain guides for Claude on Mac, ChatGPT for Mac, Gemini desktop for Mac and Grok on Mac. To see which cloud model currently leads which benchmark, start with our best AI models guide.
Should You Wait for the M5 Ultra Mac Studio?
For memory-bound local AI work, yes, wait, and that is a firmer answer than this page could give three months ago. A $5,299 Mac Studio capped at 96GB is not the machine that made this line famous for local inference, and the 768GB figure, uncertain as it is, is the only route back to that capability.
If your work is not memory-bound, stop waiting. The M4 Max at $2,499 is a strong desktop for video, 3D and general professional work, and nothing in the M5 Ultra rumours changes that. Paying the $2,800 Ultra premium for bandwidth you never saturate has never been a good trade.
If you need portability, buy the M5 Max MacBook Pro, since it reaches 128GB and out-configures the Ultra desktop on memory today. If you are weighing Apple Silicon against the alternatives, our MacBook vs Googlebook vs Chromebook comparison sets out the wider landscape, and our look at Apple’s AI roadmap covers what comes after this hardware cycle.
One last practical note. Check Apple’s configurator yourself on the day you buy. The memory options on this machine changed twice in nine weeks without an announcement, and the direction of travel has been down.
Verify before you spend $5,299.
FAQ
When is the M5 Ultra Mac Studio coming out?
Apple has not confirmed a date, and WWDC 2026 in June passed without an announcement. Reporting from AppleInsider in late June said the M5 Ultra is still due later in 2026, and Bloomberg has pointed to around October 2026. The global DRAM shortage is the reason for the delay.
How much memory does the Mac Studio have now?
96GB, and that is the only option. Apple removed the 512GB configuration in March 2026 and the 256GB configuration on 5 May 2026. Apple’s own tech specs page lists 96GB unified memory for the M3 Ultra with no configurable-to line, and its comparison page describes the machine as from 96GB.
Will the M5 Ultra Mac Studio have 768GB of memory?
It is possible but not confirmed. MacRumors reported in June 2026 that Apple has tested support for up to 768GB of unified memory, while warning that supply constraints could prevent that option from launching. The same report noted it could push a loaded Mac Studio past $10,000.
How much does the Mac Studio cost?
Read live from Apple’s store: the M4 Max starts at $2,499 and the M3 Ultra at $5,299 with 96GB and 1TB. Both rose on 25 June 2026, from $1,999 and $3,999. Some guides still print the older figures, so check Apple directly.
Can a Mac Studio run 70B local LLMs?
Yes, comfortably. The MLX build of Llama 3.3 70B at 4-bit is a measured 39.7GB, which fits in 96GB with room for a second model. Budget around 75-80GB usable after macOS and the KV cache.
Can a Mac Studio run Kimi K3 or DeepSeek V4?
No, and neither would a 768GB version. Kimi K3’s weights are 1.56 TB on disk and DeepSeek V4 Pro’s are 865GB, measured from their own repositories. Today’s frontier open-weight models need server hardware, not a desktop.
Is there going to be an M6 Ultra?
No. Apple is skipping the M6 generation at the high end, so there will be no M6 Pro, M6 Max or M6 Ultra. AppleInsider reports the next Ultra chip is the M7 Ultra in 2028, which makes the M5 Ultra the only Ultra-tier upgrade for roughly two years.
Is a Mac Studio or an RTX 5090 better for local AI?
It depends on model size. The RTX 5090 is faster on anything that fits in its 32GB of VRAM. A 96GB Mac Studio holds models the 5090 cannot, in one quiet box at far lower power. That capacity advantage is now three to one rather than the eight to one it was when 256GB was available.
Does macOS 27 support Intel Macs?
No. macOS 27 Golden Gate, announced at WWDC 2026 on 8 June, is the first Apple Silicon-only version of macOS and the last with full Rosetta 2. Intel Macs stop at macOS 26 Tahoe.




