Nvidia has told the contract manufacturers who assemble its AI servers that prices on many configurations will rise by more than 15% for systems shipping in early 2027, and the reason is memory rather than silicon. Bloomberg reported the warning on Saturday, August 22, 2026, and CNBC and Tom's Hardware carried it the same day. The increase varies by chip generation and by how much DRAM and high bandwidth memory a configuration carries, which means it is not a flat surcharge you can pencil into a spreadsheet. If you are sizing a 2027 cluster, the assumption that hardware gets cheaper per unit of compute no longer holds.
The short answer
Nvidia has warned its largest customers, through the contract manufacturers that build the servers, that AI system prices rise by more than 15% for hardware shipping in early 2027. The driver is memory cost rather than accelerator cost. Vera Rubin and Grace Blackwell configurations are both affected, with the exact increase depending on chip generation and on how much DRAM and high bandwidth memory a system carries. Microsoft, Google and Oracle are among the operators whose assemblers passed the new pricing on.
For three decades the safe assumption in capacity planning was that waiting made hardware cheaper. That assumption just failed in public.
What Nvidia actually told its customers
The warning did not arrive as a press release. Bloomberg reported on Saturday, August 22, 2026 that Nvidia had communicated the increase privately to customers and to the companies that assemble its servers under contract. Those assemblers then notified the operators they build for, a list that reporting says includes Microsoft, Google and Oracle. CNBC and Tom's Hardware picked the story up the same day. Nvidia did not comment.
The headline figure is more than 15%, and it applies to systems shipping in early 2027. Two details make it harder to plan around than a single percentage suggests.
The first is that the increase varies by chip generation and by memory configuration. Both the Vera Rubin and Grace Blackwell platforms are named, and within each of them a configuration carrying more memory per accelerator absorbs more of the rise. So the surcharge is not uniform across a fleet, and it is not uniform even across two racks of the same generation.
The second is that these are contract prices. There is no list to check. What you get is a number from your supplier, which is exactly the situation where a buyer without volume has the least leverage.
Memory is the bottleneck, not compute
The cause is worth stating plainly, because it changes what a fix would look like. Samsung, SK hynix and Micron produce most of the DRAM made in the world. Output is rising. It is not rising fast enough to meet demand from AI infrastructure, and high bandwidth memory competes for the same advanced packaging capacity that everything else needs.
Each accelerator generation attaches more memory than the one before it, so demand per unit of compute is climbing at the same time as total units. That is a compounding squeeze, and it is not one that additional accelerator supply relieves. Nvidia can build more GPUs than the memory industry can feed.
This is also why the increase lands on system prices rather than on chip prices. The bill of materials for a rack now carries a memory line large enough to move the total by double digits on its own. It also follows a pattern we saw earlier this month, when Samsung raised foundry prices by up to 15% on the back of the same demand. Different company, different part of the chain, same direction of travel.
What it changes for the people signing the orders
If you buy hardware, the practical change is that dollars per accelerator has stopped being a useful planning unit. The increase tracks memory configuration, so a budget line that averages across a fleet will be wrong in both directions: too high for lean configurations, too low for the memory heavy ones you probably wanted.
If you rent rather than buy, the effect arrives later and indirectly. Operators building on 2026 pricing assumptions have to absorb the difference or pass it through. Reporting notes that European AI infrastructure projects budgeted at around 20 billion euros, along with commercial operators such as Nebius, were planned against the previous pricing. Nobody rebuilds a multi year financial model quietly.
And if you are sitting on a multi year roadmap that assumed cost per unit of compute would keep falling, this is the moment to look at it again. The roadmap is not wrong about compute getting cheaper in the long run. It is wrong about the next two years, because the constraint moved to a part of the supply chain that cannot be expanded on the same timescale.
What we would take from this
The interesting thing is not the percentage. It is where the percentage came from.
For most of the AI buildout, the scarce component was the accelerator, and every planning conversation was about allocation: who gets chips, when, and at what priority. Memory was a line item you configured. It is now the item setting the price of the whole system, which means the questions worth asking in a procurement meeting have changed. How much memory does this workload actually need, per accelerator, is suddenly a cost question rather than a performance one.
We would also treat this as a reminder about where fragility lives. A supply chain with three significant DRAM suppliers behaves differently from one with a single dominant accelerator vendor, but not in a reassuring way. Three suppliers all constrained by the same packaging capacity is not diversity, it is one bottleneck wearing three names.
Sources and further reading
- Nvidia Customers Notified About AI-Related Price Hikes Above 15%, Bloomberg, August 22, 2026
- Nvidia customers reportedly warned about AI-related price hikes, CNBC, August 22, 2026
- Nvidia reportedly warns biggest customers of 15% price hikes on AI servers, Toms Hardware
- Nvidia AI server prices are rising more than 15% from early next year, TNW
Frequently asked questions
How much are Nvidia AI server prices going up?
More than 15% on many configurations, with the exact figure varying by chip generation and by memory configuration. That variability is the important part. A rack that carries a lot of high bandwidth memory per accelerator absorbs more of the increase than a lighter configuration, so two systems from the same generation can move by noticeably different amounts. Nobody has published a price list, because these are contract prices communicated privately between Nvidia, the companies that assemble the servers and the operators buying them.
When do the higher prices take effect?
The increases apply to systems shipping in early 2027, not to hardware already on order or being delivered now. That gives roughly a two quarter window between the warning and the first affected deliveries. In practice the useful deadline is earlier than the shipping date, because large orders are placed and priced months ahead, so anything you intend to buy at current pricing needs to be committed well before the calendar turns.
Why is memory the constraint rather than the accelerators?
Because each new accelerator generation needs more memory attached to it, and memory production has not scaled at the same rate. Samsung, SK hynix and Micron supply most of the DRAM made in the world. Their output is rising, but not fast enough to match AI infrastructure demand, and high bandwidth memory competes for the same fabs and packaging capacity as conventional server DRAM. When the shortage is on the memory side, adding accelerator supply does not solve it.
Which systems are affected?
Systems built around the Vera Rubin and Grace Blackwell platforms. Reporting names Microsoft, Google and Oracle among the operators whose contract assemblers passed the new pricing along, which is a reasonable proxy for saying the increase reaches everyone buying at scale. Smaller buyers and rental operators are exposed the same way, and arguably worse, because they have less room to renegotiate.
What should we do about it in a 2027 budget?
Three things. Price your 2027 capacity with a memory sensitive line rather than a single figure for dollars per accelerator, since the increase tracks memory configuration. Re-examine whether every workload needs the largest memory option, because that is where the premium concentrates. And revisit any multi year plan that assumed falling unit costs, since a plan built on 2025 and 2026 pricing has a gap in it that no amount of scheduling will close.