Insights Business| SaaS| Technology How the Memory Crunch Reprices Software and Flips the Build Versus Buy Decision
Business
|
SaaS
|
Technology
Sep 16, 2026

How the Memory Crunch Reprices Software and Flips the Build Versus Buy Decision

AUTHOR

James A. Wondrasek James A. Wondrasek
How the memory crunch reprices software and flips the build versus buy decision

You probably still price the cost of serving software by compute: inference burns FLOPs, and build-versus-buy is a raw-price comparison between renting on-demand cloud and buying hardware at sticker price. That frame will cost you money.

The 2026 memory crunch is a structural repricing. DRAM contract prices rose roughly 90 to 98 per cent in Q1, then another 58 to 63 per cent in Q2, with no supply relief before 2027 or 2028. Memory now sets the cost of serving software, part of the wider repricing of software.

Memory-hungry workloads feel it first, and it flows into cloud bills, per-token pricing and SaaS gross margins. Get this call wrong and you lock in the wrong cost basis for years. Let’s work through it in order: first why inference turned into a memory problem, then what build-versus-buy is actually settled by, and then what the repricing means for your margins.

Why is inference becoming a state-management problem rather than a compute problem?

Inference is bandwidth-bound: every generated token streams the resident weights and KV cache from memory, so memory bandwidth sets the cost. Where the KV cache lives, and how it tiers across VRAM, DRAM and flash, is the primary cost lever.

The memory wall: bandwidth-bound decode

The unit that matters is dollars per petabyte moved, which is nearly model-agnostic and separates the hardware economics from the model choice. On identical H100s, realised dollars per million tokens run from $0.21 to $15.25 depending on utilisation and offered request rate, so state residency sets the spread.

KV cache: from compute to state

MoE sparsifies compute but not memory: Mixtral-8x7B activates 12.9 billion parameters per token yet needs 31.2 GB of KV cache at 128K context, more than its 25.4 GB of active weights. Quantisation from FP8 to FP4 halves bytes per weight, and KV-cache compression like MLA drops the cache 93 per cent. CXL and tiered memory recycle DDR4 and DDR5 from decommissioned servers at a fraction of new-DRAM cost.

ASICs versus GPUs, and flash versus DRAM

GPU-hosted inference keeps the KV cache in HBM. Flash-based serving tiers context to NAND and enterprise SSD, cheaper but slower; flash sustained about 30,000 tokens per second against DRAM’s 17,000 once memory filled. ASICs right-size the memory subsystem, but the saving comes out of the merchant margin, not the memory premium the shortage has inflated, so only operators with the volume justify them.

How do you assess build versus buy for memory-hungry workloads?

With memory established as the cost driver, the build-versus-buy call changes shape. It is one strand of the memory crisis rewriting software economics. Assess on total cost of ownership: compare owned or dedicated capacity against on-demand cloud across cost basis, utilisation and availability. When memory dominates the bill, fixed-cost dedicated capacity undercuts rented cloud for stable workloads, but only at high utilisation, and only if you account for the depreciation conveyor.

Fixed-cost dedicated versus on-demand cloud

Fixed-cost infrastructure is a hedge against upside volatility; it does not guarantee the lowest price at any given moment. A rental reprices at every renewal; an owned asset fixes its cost on the day you buy. On-demand cloud passes the rise into per-hour pricing and sits at the back of the allocation queue when memory is scarce. See the companion piece for board-level memory cost forecasting.

The depreciation conveyor and the incumbent floor

Incumbents amortise capacity bought at pre-crunch prices and run it at marginal cost, a floor new entrants buying at peak prices cannot reach. A new entrant buying GB300-class capacity in 2026 pays near $0.174 per petabyte against an incumbent floor around $0.054, a 3.2 times gap.

Where utilisation decides the answer

Owned memory is an asset at high utilisation and a liability when idle. Break-even utilisation sits around 85 per cent, so steady-state workloads favour fixed-cost capacity while development, test and variable workloads stay on-demand. For most operators the answer is a hybrid, with committed allocations locked in before the next vintage reprices.

How does the memory crunch change the economics of memory-hungry software?

The crunch turns inference into a variable cost of goods sold, compressing SaaS gross margins from the 80 per cent norm toward 60 to 70 per cent — the income-statement edge of the structural repricing of software economics. Per-token and per-petabyte pricing must re-base to the new memory cost, pushing vendors toward consumption or outcome-based pricing that moves variable inference cost back to the customer.

Gross margin: from 80 per cent to the 60 to 70 per cent band

For every $1 million in AI product revenue in 2026, roughly $230,000 walks out the door as inference cost before an engineer or seller gets paid. Snowflake’s trailing product margin sits at 67.2 per cent against a 75 per cent target, and public SaaS companies report margins 10 to 17 points below pre-AI baselines.

Token economics and per-token re-basing

Dollars-per-petabyte and cost-per-million-tokens are the unit-economics frame. Track the inference efficiency ratio, AI-related revenue divided by inference cost. As a rough heuristic, a ratio above five is healthy and anything under three is a structural problem that compounds with usage.

Pricing models that move cost off the balance sheet

Consumption and outcome-based pricing migrate variable inference cost to the customer; Salesforce’s Agentforce, Intercom’s Fin and ServiceNow’s Now Assist all shipped hybrid pricing in 2026. Model routing and prefix caching do the heavy lifting: Coinbase saved 90 per cent of its token spend by routing prompts to open-source models, and providers now offer roughly 90 per cent discounts on cached input tokens, which cuts per-query cost by an order of magnitude.

Build-versus-buy is now a memory-driven cost-structure and margin decision, decided on cost basis, utilisation and the depreciation conveyor. The question is no longer “cloud or own?”; it is “what do my utilisation and vintage say?”. This is the memory crisis and software economics, arriving on your income statement.

Every memory-reduction lever, quantisation, KV-cache compression, CXL recycling, flash tiering, is a margin decision. So work out your memory exposure and re-base pricing before the next vintage reprices your commitments. The alternative is watching the fleet and endpoint implications land on someone else’s books.

Frequently Asked Questions

What is the memory crunch, in plain terms?

The memory crunch is a structural shortage of DRAM, HBM and NAND flash that has pushed contract prices up roughly 90 to 98 per cent in one quarter and a further 58 to 63 per cent in the next. It is driven by AI infrastructure demand outstripping memory supply, and unlike a normal cycle, no meaningful relief is expected before 2027 or 2028.

Is the memory crunch just a normal supply cycle that will correct itself?

No. This is a structural repricing rather than a typical boom-and-bust cycle. Demand from AI data centres is compounding faster than fabrication capacity can expand, and new fabs take years to build and qualify. That is why no supply relief is forecast before 2027 or 2028, and why planning on a quick price correction is a costly assumption.

How long will the memory premium last?

Most forecasts point to no meaningful supply relief before 2027 or 2028, because new memory fabrication capacity takes years to build and qualify. Treat the premium as a multi-year cost basis, not a short spike. Any build-versus-buy or pricing decision should be modelled across that horizon, not against a hoped-for return to pre-crunch prices.

What is the depreciation conveyor, and why does it give incumbents an advantage?

The depreciation conveyor is the advantage incumbents hold by amortising memory and compute bought at pre-crunch prices while running it at marginal cost. New entrants buying at peak prices carry a higher cost basis they cannot match. The build-versus-buy flip is therefore won on when you bought capacity, not on how well you operate it.

Should a startup build or buy if its workload is unpredictable?

Buy. Owned memory is an asset at high utilisation and a liability when idle, so unpredictable or spiky workloads usually favour on-demand cloud despite rising per-hour rates. Build only for stable, high-utilisation inference where fixed-cost capacity can be kept busy; otherwise idle capacity erodes the very cost advantage that justified owning it.

How are ASIC-powered inference accelerators different from general-purpose GPUs?

ASICs are purpose-built for inference and right-size the memory subsystem, which improves memory-per-token economics. But they attack the merchant margin, not the underlying memory premium, so they only pay off at volume. Operators without predictable, high-volume demand rarely justify the engineering and capital commitment that purpose-built silicon requires.

What is the difference between GPU-hosted inference and flash-based KV cache serving?

GPU-hosted inference keeps the KV cache in expensive HBM or DRAM for maximum speed, while flash-based serving tiers context to NAND flash and enterprise SSD. Flash is slower per access but far cheaper per terabyte, so it holds context DRAM cannot and sustains throughput once memory fills. The trade-off is latency against cost per token.

Does the memory crunch affect smaller software companies, or only large AI labs?

It reaches smaller companies through the cloud bill. Even if you never buy a GPU, your provider passes rising memory costs into per-hour and per-token pricing. Any memory-hungry workload, from retrieval-augmented generation to long-context features, sees its unit economics shift. That is why the crunch is a margin problem for ordinary SaaS businesses, not just hyperscalers.

How do I work out my own exposure to memory costs?

Start by measuring where inference dominates your cost of goods sold. Track cost per million tokens and dollars-per-petabyte moved, then forecast that line against rising memory prices. If variable inference cost is material, model how it compresses gross margin and re-base pricing before the next vintage reprices your commitments.

Is it worth buying hardware now, or should I wait for prices to fall?

Waiting carries its own cost: current allocations are scarce, and the next vintage will be priced higher if the crunch persists as forecast. Lock in committed allocations now if your utilisation is stable and high. If your load is genuinely uncertain, stay on-demand rather than buy capacity you cannot keep busy.

How does model quantisation help with memory costs?

Quantisation shrinks each weight to lower precision, such as FP8 or FP4, so the working set fits in less memory and streams faster. Combined with KV-cache compression methods like MLA and quantised cache, it reduces the memory each token must carry. That pushes the memory wall back and claws back gross margin without easing the underlying shortage.

What is CXL, and can it really cut my memory bill?

CXL is a memory interconnect that lets servers attach additional DRAM or recycling pools over a standard link. It helps most by recovering DDR4 and DDR5 from decommissioned servers and reusing it as tiered capacity, extending effective memory at a fraction of new-DRAM cost. It softens the premium for operators who can exploit it, and does nothing for those who cannot.

AUTHOR

James A. Wondrasek James A. Wondrasek

SHARE ARTICLE

Share
Copy Link

Related Articles

Need a reliable team to help achieve your software goals?

Drop us a line! We'd love to discuss your project.

Offices Dots
Offices

BUSINESS HOURS

Monday - Friday
9 AM - 9 PM (Sydney Time)
9 AM - 5 PM (Yogyakarta Time)

Monday - Friday
9 AM - 9 PM (Sydney Time)
9 AM - 5 PM (Yogyakarta Time)

Sydney

SYDNEY

55 Pyrmont Bridge Road
Pyrmont, NSW, 2009
Australia

55 Pyrmont Bridge Road, Pyrmont, NSW, 2009, Australia

+61 2-8123-0997

Yogyakarta

YOGYAKARTA

Unit A & B
Jl. Prof. Herman Yohanes No.1125, Terban, Gondokusuman, Yogyakarta,
Daerah Istimewa Yogyakarta 55223
Indonesia

Unit A & B Jl. Prof. Herman Yohanes No.1125, Yogyakarta, Daerah Istimewa Yogyakarta 55223, Indonesia

+62 274-4539660
Bandung

BANDUNG

JL. Banda No. 30
Bandung 40115
Indonesia

JL. Banda No. 30, Bandung 40115, Indonesia

+62 858-6514-9577

Subscribe to our newsletter