Apple announced two new desktop Macs and two new chips on Tuesday, August twenty fifth, 2026: a Mac mini built around M6, Apple's first 2 nanometer chip, and a Mac Studio that tops out with M5 Ultra, the first quad die part Apple has built. For most buyers the interesting figure is the price. For anyone who runs models locally it is a different one. M5 Ultra addresses up to 512 gigabytes of unified memory at 1.2 terabytes per second, which puts a class of model that currently needs rented accelerators inside a machine you can buy outright and keep on a desk.
The short answer
Apple announced M6 and M5 Ultra on August 25, 2026, alongside a new Mac mini and Mac Studio. M6 is Apple's first 2 nanometer chip, with a 12 core CPU, a dual 16 core Neural Engine and 170GB/s of bandwidth. M5 Ultra is the first quad die Apple part, joined by UltraFusion at over 4.4TB/s between dies, reaching 36 CPU cores, 80 GPU cores and 512GB of unified memory at 1.2TB/s. Pre orders opened the same day, machines ship September 22, and the 512GB build arrives late October.
Apple shipping faster desktop Macs is not news. Apple shipping a desktop that can hold a model most teams currently rent time on is a different kind of announcement, and it arrived on a Tuesday in August rather than at the autumn event where everyone expected it.
The number that changes a decision
M5 Ultra supports up to 512 gigabytes of unified memory with 1.2 terabytes per second of bandwidth. Apple frames that as 50 percent higher bandwidth than M3 Ultra, which is the right comparison for a chip generation but the wrong one for understanding what the machine is for.
The comparison that matters is against the thing you would otherwise do. Half a terabyte of memory that an accelerator can address coherently, in one pool, without sharding, is not something you assemble cheaply. It normally means multiple cards, a chassis, a power draw that rules out a normal room, and a software layer that knows how your weights are split. Here it is a box on a desk with a 5,499 dollar starting price, and the 512 gigabyte configuration is a build option rather than a different architecture.
Apple's own phrasing is that the machine supports huge LLMs with hundreds of billions of parameters entirely on device. Note what is absent. There is no named model, no tokens per second figure, and no prompt processing benchmark. Capacity is the claim being made. Throughput is the question nobody can answer until hardware is in hands, and it is the question that decides whether this is a serious inference machine or an expensive way to watch a model load.
Four dies, one address space
M5 Ultra is the first quad die part Apple has built. Previous Ultra chips fused two Max dies; this one fuses four, using the same UltraFusion packaging, with inter die bandwidth Apple puts at over 4.4 terabytes per second.
That interconnect figure is worth holding next to the memory figure. Inter die bandwidth is roughly three and a half times the bandwidth to memory, which is a deliberate ratio: it means a die reaching across the package is not obviously worse off than a die reaching into its own memory controller, and the four dies can behave as one chip rather than as four things a scheduler has to be clever about.
That is the actual difference between this and a multi GPU workstation. On a box with four discrete cards you own the placement problem. Weights get sharded, activations cross a bus you are painfully aware of, and a meaningful share of your engineering time goes into not crossing it. Here the sharding does not exist because the memory pool does not have seams. Whether the resulting throughput is competitive is a separate question, but the programming model is simpler in a way that has real cost consequences.
The full configuration ladder is 36 CPU cores at the top, split 12 super cores and 24 performance cores, with up to 80 GPU cores carrying Neural Accelerators and a 32 core Neural Engine. Apple claims up to 1.25 times the single threaded performance and 1.3 times the multithreaded performance of M3 Ultra, and up to 4.5 times the peak GPU compute for AI. Those are generational figures against Apple's own previous part, not against anything else.
M6 and the smaller machine
The Mac mini is the other half of the announcement, and it is the one more people will actually buy.
M6 is Apple's first 2 nanometer chip. The CPU is 12 cores, arranged as 2 super cores, 4 performance cores and 6 efficiency cores, which is two more cores than M4 in both CPU and GPU. The Neural Engine is the structural change: two 16 core engines rather than one, which Apple says is up to twice the peak compute of the previous generation. Memory bandwidth is 170 gigabytes per second, ten percent over M5 and two and a half times M1, with a 32 gigabyte ceiling on unified memory.
Against a Mac mini with M4, Apple quotes up to 40 percent faster CPU, up to twice the graphics and storage performance, and up to four times faster AI performance. Pricing starts at 899 dollars, which is the highest a Mac mini has ever started, in a period where memory pricing has been pushing server costs up across the industry. An M5 Pro configuration sits above it at 1,699 dollars with up to 18 CPU cores, 20 GPU cores, 64 gigabytes of memory and 307 gigabytes per second of bandwidth.
For infrastructure use there are two details in the small print worth pulling out. Ethernet is 2.5 gigabit as standard on both models with a 10 gigabit option, which is a genuine improvement for anyone using these as build agents or small servers. And the three rear ports are Thunderbolt 4 on the M6 model but Thunderbolt 5 on M5 Pro, so the cheaper chip is the one with the older bus.
What we would wait to find out
Three things are unresolved, and all three decide whether the top configuration is interesting to you.
The first is real inference throughput on a long context. Capacity gets a model loaded; bandwidth and the Neural Accelerators decide whether it is usable at a hundred thousand tokens of context. Apple published neither.
The second is availability of the configuration that matters. The 512 gigabyte Mac Studio is listed as late October, a month behind everything else, which usually indicates memory supply rather than a manufacturing choice.
The third is price at that configuration. The 5,499 dollar figure is where M5 Ultra starts, not where a half terabyte machine lands, and Apple has not said what the step costs. Given what memory is doing to bills of materials this year, that gap is unlikely to be small.
None of that changes the shape of the announcement. A single vendor now sells a desk sized machine with half a terabyte of coherent accelerator memory, at a price a small team can approve without a procurement cycle. Whether it is fast enough is the next question, and it is a much better question to have than the one people were asking last year.
Sources and further reading
- Apple introduces M6 and M5 Ultra for a big leap in performance and AI compute, Apple Newsroom, August 25, 2026
- Apple introduces new Mac Studio with M5 Max and M5 Ultra, Apple Newsroom, August 25, 2026
- Apple unveils a more powerful Mac mini featuring the all-new M6 and M5 Pro, Apple Newsroom, August 25, 2026
- Apple Announces New Mac Mini With M6 and M5 Pro Chips and More, MacRumors, August 25, 2026
- Apple unveils M6, M5 Ultra chips and updates its desktop Macs, Six Colors, August 25, 2026
Frequently asked questions
Can a 512GB Mac Studio really run a frontier sized model?
It can hold one, which is the constraint that usually bites first. Apple says M5 Ultra supports huge LLMs with hundreds of billions of parameters entirely on device, and it did not name a model or publish a tokens per second figure, so treat capacity as the claim and throughput as an open question. The arithmetic is straightforward enough to do yourself: a 400 billion parameter model at 4 bit quantisation needs roughly 200 gigabytes for weights before you add context, so 512 gigabytes leaves genuine headroom. The number to watch once machines ship is not whether it loads but what prompt processing looks like on a long context, because that is where a unified memory design and a rack of discrete accelerators diverge most.
How does 1.2TB/s compare to a datacentre accelerator?
It is a fraction of one, and that comparison is the honest way to read the announcement. A current generation datacentre GPU sits in the multiple terabytes per second range on HBM, so M5 Ultra is not competing on raw bandwidth. What it competes on is capacity per box and cost of ownership. Getting 512 gigabytes of accelerator attached memory in one coherent address space normally means several cards, a chassis to hold them, and a power budget that rules out an office. Apple quotes 1.2 terabytes per second as 50 percent higher than M3 Ultra, and the useful question is whether that is enough bandwidth to make the capacity worth having for your workload rather than whether it beats HBM.
What is a quad die chip and why does it matter here?
M5 Ultra is the first Apple silicon part built from four dies rather than two. Apple joins them with UltraFusion, its packaging interconnect, and says inter die bandwidth now exceeds 4.4 terabytes per second. The reason this matters is that the four dies present as one chip with one memory pool, so software does not have to be written to shard across them. That is the difference between a multi die package and a multi GPU box: on the second you manage placement yourself and pay for every crossing. The scaling ceiling is real, and 4.4 terabytes per second between dies against 1.2 to memory tells you Apple sized the interconnect so that die crossings are not the bottleneck.
Is the M6 Mac mini worth it for a build machine or a small server?
At 899 dollars for a 12 core CPU with 16 gigabytes of memory it is a reasonable build box, and the 2.5 gigabit ethernet as standard with a 10 gigabit option matters more for that role than the AI figures do. The caveats are the ones that have always applied to this machine. Memory is not upgradable, so 32 gigabytes is a hard ceiling on the M6 model and you buy it at purchase or not at all. Storage starts at 256 gigabytes. And the three ports on the back are Thunderbolt 4 on M6 while the M5 Pro model gets Thunderbolt 5, which is the kind of split worth checking before you order a rack of them.
When can I actually get one, and what ships late?
Pre orders opened on August 25, 2026, in 30 countries and regions. Both machines start arriving to customers and to stores on Tuesday, September 22. The exception is the configuration most relevant to this article: Mac Studio with 512 gigabytes of unified memory is listed as coming in late October, roughly a month after everything else. If your interest in this announcement is specifically the large memory configuration, plan for a November arrival rather than a September one, and note that Apple has not published pricing for that specific build beyond the 5,499 dollar starting point for M5 Ultra.