Insights Business| SaaS| Technology How the Memory Crisis Is Rewriting Software Economics: Cloud Costs, Margins and Build vs Buy
Business
|
SaaS
|
Technology
Sep 16, 2026

How the Memory Crisis Is Rewriting Software Economics: Cloud Costs, Margins and Build vs Buy

AUTHOR

James A. Wondrasek James A. Wondrasek
How the memory crisis is rewriting software economics

The memory crunch stopped being a supply story when the prices stopped behaving like one. Conventional DRAM contract prices rose 90-95% quarter-on-quarter in Q1 2026, then a further 58-63% in Q2, and DDR5 is up roughly 500% year-on-year, according to TrendForce, which projects DRAM and NAND will consume 47% of cloud hardware spend in 2026 and 68% in 2027.

The repricing shows up in specific, named prices. OVHcloud pays six times more for RAM than it did in June 2025 and forecasts twelve times by early 2027, while Hetzner’s CCX13 jumped from €15.99 to €42.99 a month. Those are representative mid-tier providers, and the same force reaches the hyperscalers: AWS raised EC2 Capacity Block prices by 15% in January 2026.

Three forces arrived at once: HBM crowding conventional DRAM off the wafer, hyperscalers pre-paying for most of 2027 output, and cloud hardware spend shifting from compute to memory. The old supply-blip-then-correct model no longer holds. This page is the map; the four articles below are the territory.

In This Series

Why is the memory crisis a structural repricing rather than a normal supply cycle?

A normal cycle self-corrects: demand softens, prices fall, supply catches up. This one does not, because capacity is committed and reallocated before demand is ever met. High-bandwidth memory consumes three to four times the wafer area per bit of conventional DRAM, and hyperscalers have pre-paid for most of 2027 output through long-term agreements. The visible supply is already sold, so prices re-base upward rather than correct, which is the definition of a structural repricing.

DRAM revenue rose roughly 81% quarter-on-quarter in Q1 2026 while bit shipments stayed constrained, the signature of an ASP-driven repricing rather than a demand-volume boom, according to Silicon Analysts. Advanced packaging is geographically concentrated, and Samsung, SK Hynix and Micron have already reallocated wafer capacity toward HBM.

Read Why the memory crisis is structural rather than a normal supply cycle for the full causal case.

How does the memory crunch land on your cloud bill and procurement commitments?

Memory has moved from a subsidiary line item to the dominant input in cloud hardware economics: 47% of cloud hardware spend in 2026, and a projected 68% in 2027. That is why providers are repricing, from that same six-times RAM cost at OVHcloud to NVIDIA passing HBM costs through as a more-than-15% AI server price rise. Your decision is commitment timing: whether to lock pricing now or stay flexible.

TrendForce’s 47% to 68% projection makes memory a planning variable rather than an operations detail. Amazon raised 2026 capex to $220 billion and cited memory for roughly $20 billion of the increase, while Google, Microsoft, Oracle and Meta all re-base. For planning, the distinction is contract versus spot: committed buyers pay contract prices, everyone else faces the volatile spot market.

Read How memory costs are reshaping cloud budgets and procurement commitments to build a board-defensible forecast; the next section turns that cost pressure into an architecture decision.

How does the crunch reprice memory-hungry software and flip build versus buy?

The crunch turns memory from an operating line item into a product-margin variable. Memory-hungry SaaS, per-token pricing and always-on inference all re-base their cost structure as RAM prices climb, and the cloud-first default inverts: at sustained utilisation, fixed-cost dedicated capacity can beat on-demand cloud. The differentiator is the cost floor: incumbents running fully amortised hardware hold a floor that new entrants buying at peak prices cannot match.

Inference has shifted from compute-bound to state-bound. The working set (weights, activations and especially the KV cache) now sets the cost of serving more than raw FLOPs do, so the useful unit becomes dollars per token served, which now tracks dollars per petabyte of working-set memory moved. Quantisation, mixture-of-experts routing and KV-cache compression reduce the pressure, but Jevons paradox means they have not lowered prices. CXL and tiered memory offer a recycling path for operators who can exploit them.

Read How the memory crunch reprices software and flips the build versus buy decision for the TCO and margin detail; the fleet and endpoint implications follow in the final section.

Should you refresh endpoint hardware now or wait out the premium?

Endpoint prices across PCs, phones and smart speakers are a downstream echo of the same DDR5 repricing driving server costs, passed through by OEMs as memory’s share of the bill of materials rises. The refresh-versus-extend decision hinges on fleet age and the cost of carrying ageing kit rather than the calendar. If the premium persists toward 2028, deferring raises your total cost; if it normalises, waiting wins.

Dell, HP, Lenovo, HPE and Cisco have all passed memory costs through, and HP now counts memory at roughly 35% of its PC bill of materials. Even low-memory devices repriced: Amazon’s Echo Dot moved from $49.99 to $79.99. Memory-upgraded or refurbished hardware often beats new-build pricing, and staging refreshes around depreciation schedules defer the peak-pricing hit without freezing capability. OVHcloud’s warning that conditions will last until 2028 sits against the argument that new capacity and ASIC adoption normalise prices sooner, and Gartner, Counterpoint, Goldman Sachs, Morgan Stanley and IDC publish guidance that does not agree.

Read Why endpoint prices are rising and how long the memory premium will last for the evidence and the vintage-exposure analysis.

Resource Hub: The Memory Crisis Deep Dives

Understanding the Repricing

Deciding What to Do

Where to Start

Where this leaves you

The memory crunch is a repricing of software economics, and the four articles above map it to the decision in front of you.

Frequently Asked Questions

Does the memory crisis affect me if I never use AI?

It affects you regardless of whether you use AI, because AI and conventional workloads compete for the same wafer capacity. High-bandwidth memory consumes three to four times the wafer area per bit of standard DRAM, so every HBM wafer is DDR5 that was never built. DDR5 is up roughly 500% year-on-year, and that price reaches every server, laptop and phone.

What is the difference between contract memory prices and spot prices?

The difference is who carries the price risk. Contract prices are negotiated in advance for committed volumes, so buyers pay a fixed, predictable rate; spot prices float with the market and move far more violently. Memory has moved from a subsidiary line item to the dominant input in cloud hardware economics, so the distinction now decides whether your forecast holds or breaks.

Why does HBM cost so much more per bit than conventional DRAM?

High-bandwidth memory is dearer per bit because it is harder to build, not just scarcer. HBM stacks dies vertically and needs advanced packaging, including through-silicon vias and high-precision bonding, which yields less usable output per wafer than conventional DRAM. It also consumes three to four times the wafer area per bit, so the true cost gap is a manufacturing one.

What is the KV cache, and why does it drive my inference bill?

The KV cache is the stored key and value tensors a model keeps for every token already processed, and it is what makes long conversations expensive. It grows with context length and concurrent users, so serving a 100,000-token context can hold far more memory in cache than the model weights themselves. That is why inference is now state-bound rather than compute-bound.

Can quantization or mixture-of-experts routing bring memory costs back down?

They lower memory pressure per model, but they have not lowered prices, because of Jevons paradox. Quantization shrinks weights and KV-cache compression shrinks context, so each token costs less to serve. Cheaper serving then expands demand, absorbing the savings. The useful unit is dollars per petabyte moved per token, and that figure keeps climbing even as efficiency improves.

Does the memory crunch affect GPUs as well as RAM?

GPUs are affected, but through the memory attached to them rather than the silicon itself. AI accelerators carry HBM stacks, and HBM is the same capacity crowding conventional DRAM off the wafer. NVIDIA has passed those costs through as a more than 15% AI server price rise, so GPU economics and memory economics are now the same conversation.

Are refurbished or second-hand servers a safe way to dodge the premium?

They can be a legitimate hedge, with caveats. Second-hand and refurbished hardware often beats new-build pricing because it avoids peak memory costs, and for steady-state workloads the performance gap is small. The risks are shorter warranty cover, uncertain firmware support and slower spare parts, so model the carrying cost of ageing kit before you commit.

Has this happened before, and how did previous memory cycles end?

Memory has cycled before, but not like this. Earlier corrections followed a predictable pattern: demand softened, inventories built, prices fell and marginal capacity exited. This time capacity is committed and reallocated before demand is met, so there is no inventory overhang to clear. The old model rewards waiting for the correction; this reset rewards committing early.

If memory is 35% of a PC’s bill of materials, why hasn’t the price risen by that much?

The memory share of the bill of materials is not the same as the price rise. A 35% memory component rising sharply lifts the total by less than the memory increase alone, because the other 65% (chassis, display, battery, assembly) has not moved. That is why device prices rise by tens of dollars rather than doubling, even when RAM costs surge.

Why can’t cloud providers just absorb the memory cost increases?

They cannot absorb it because memory is now too large a share of their cost base. When memory reaches 47% of cloud hardware spend, a 90-95% quarterly increase cannot be hidden in margin without wiping out the business. That is why OVHcloud, Hetzner and the hyperscalers all repriced rather than absorbed, and why pass-through is now the default response.

What signals should I watch to know when memory prices will start falling?

Watch three signals: HBM capacity announcements, whether hyperscaler pre-purchases slow, and DDR5 contract price prints. Prices normalise when new wafer capacity comes online and pre-commitments ease, which most forecasters place somewhere between 2027 and 2028. OVHcloud warns conditions last until 2028, while others argue ASIC adoption and new capacity pull normalisation forward. Treat any single forecast as a range, not a date.

AUTHOR

James A. Wondrasek James A. Wondrasek

SHARE ARTICLE

Share
Copy Link

Related Articles

Need a reliable team to help achieve your software goals?

Drop us a line! We'd love to discuss your project.

Offices Dots
Offices

BUSINESS HOURS

Monday - Friday
9 AM - 9 PM (Sydney Time)
9 AM - 5 PM (Yogyakarta Time)

Monday - Friday
9 AM - 9 PM (Sydney Time)
9 AM - 5 PM (Yogyakarta Time)

Sydney

SYDNEY

55 Pyrmont Bridge Road
Pyrmont, NSW, 2009
Australia

55 Pyrmont Bridge Road, Pyrmont, NSW, 2009, Australia

+61 2-8123-0997

Yogyakarta

YOGYAKARTA

Unit A & B
Jl. Prof. Herman Yohanes No.1125, Terban, Gondokusuman, Yogyakarta,
Daerah Istimewa Yogyakarta 55223
Indonesia

Unit A & B Jl. Prof. Herman Yohanes No.1125, Yogyakarta, Daerah Istimewa Yogyakarta 55223, Indonesia

+62 274-4539660
Bandung

BANDUNG

JL. Banda No. 30
Bandung 40115
Indonesia

JL. Banda No. 30, Bandung 40115, Indonesia

+62 858-6514-9577

Subscribe to our newsletter