Insights Business| SaaS| Technology What Local AI Hardware Can Run and Why Hybrid Inference Exists for Teams
Business
|
SaaS
|
Technology
•
Oct 5, 2026

What Local AI Hardware Can Run and Why Hybrid Inference Exists for Teams

AUTHOR

James A. Wondrasek James A. Wondrasek
What Local AI Hardware Can Run and Why Hybrid Inference Exists

A box with 128GB of unified memory should run a 70B model comfortably, right? That’s what the spec sheet implies. But local LLM decoding is bound by memory bandwidth: generating each token means re-reading the model’s weights, so what fits and how fast it runs are both set by bandwidth.

By the end you’ll know what a DGX Spark can actually run, why local inference degrades under concurrency, and how a routing layer keeps the routine majority on-device while escalating only the hard prompts. If you’re still unsure local models are worth it, why small local models finally qualify makes the case, and the broader local-first picture shows where hardware sits in the wider shift. Let’s start with the box itself.

What Can a DGX Spark (~$5,000 Box) Realistically Run, and Where Does It Fall Short?

A DGX Spark runs 7B to 14B models fluidly and quantised 32B models at usable speeds, but dense 70B decoding is bound by memory bandwidth. Its 128GB of unified memory at 273 GB/s means a 70B model fits, yet streams at single-digit tokens per second.

The bandwidth ceiling: why 70B fits but crawls

A 70B at 4-bit takes about 40GB, which fits in 128GB with room to spare. The constraint is speed: capacity is not the issue. Decoding re-reads the whole model for every token, so a 273 GB/s pipe caps a dense 70B at around 7 tokens per second. Small models tell the opposite story: 7B to 14B at Q4 run at 30 to 70 tokens per second, fast enough to feel interactive.

DGX Spark vs RTX 5090: capacity or speed

The RTX 5090 has only 32GB of VRAM but 1.79 TB/s of bandwidth, so it runs a 32B Q4 model at 90+ tokens per second yet can’t hold a 70B without CPU offload, which drops throughput below 10 tokens per second. The Spark offers capacity at slow speed; the 5090 offers speed within its 32GB limit.

Apple Silicon vs NVIDIA: unified memory vs CUDA

M-series unified memory, up to 192GB at 546 to 819 GB/s, lifts the 24 to 32GB VRAM ceiling for single-user 70B decoding, and MLX delivers 20 to 87% higher throughput than llama.cpp. The trade-off is tooling: CUDA’s vLLM and TensorRT-LLM ecosystem is deeper, and MLX serving is weak under concurrency. If you’re doing solo work, pick the Mac; if you’re serving a team, pick the Spark. Both are part of the local-first shift. Solo work is the easy case; serving is where a single box starts to break.

Why Does Local Inference Struggle to Serve Many Concurrent Users?

Concurrent requests share one box’s memory bandwidth and KV cache, so throughput splits and queues form. A box that feels instant for you alone starts queueing at two or three people, and your team’s latency spikes non-linearly.

Add a second concurrent stream and each stream’s throughput drops sharply, because both share the same memory interface, while concurrent prefills spike time-to-first-token (TTFT). Every concurrent user holds a context-length slice of the KV cache, competing with the weights for memory. At 32K context across three or four users, that headroom evaporates fast. Memory is the binding constraint.

Ollama vs vLLM: convenience vs throughput

Ollama wraps llama.cpp and is well suited to a single workstation, holding around 10 to 50 queries per second. vLLM adds PagedAttention and continuous batching, cutting KV-cache waste (memory reserved but unused) to under 4% and sustaining 100+ queries per second on the same GPU. That’s a 16 to 20x difference on identical hardware, so the engine is a key decision for teams.

PAIR as the clustering workaround

NVIDIA’s Personal AI Router (PAIR) schedules independent requests across paired local machines, without pooling VRAM or sharding a model. In one five-subagent benchmark, pooling an idle RTX laptop, a DGX Spark and an RTX 5090 cut the workload from 18 minutes to 8 minutes 48 seconds. It adds machines while the per-box limits stay put.

The remedies for all of this are a throughput-oriented engine such as vLLM, and clustering independent requests across machines. Once local can’t keep up, the question becomes where that workload should run.

How Does Hybrid Cloud-Edge LLM Inference Actually Work?

Hybrid cloud-edge inference routes each request between a local model and a cloud frontier model rather than sending everything to one endpoint. Complexity-based routing sends simple queries to the local model and hard ones to the cloud, while confidence-based cascading lets the local model escalate when it is unsure. Most traffic stays on-device, where latency is lowest and data stays local, and cloud spend is reserved for the hard prompts.

Real query distributions are bimodal: lots of trivial work and a little hard work. Cloud round-trips add 200 to 500ms before the first token, where a local model can feel instant. That is why your business wants hybrid: it wins on latency and privacy for the long tail without paying cloud prices for it.

Complexity-based routing: deciding what stays local

A lightweight classifier estimates difficulty and dispatches simple tasks to the local model, complex ones to the cloud. Done well, it can cut cost by 85% while keeping 95% of frontier performance. The router has to stay cheap, or it eats the latency budget it was meant to save.

Confidence-based cascading: early-exit to the cloud

The local model attempts every query first, then escalates when its token-level confidence falls below a threshold. In the Local-First pattern that looks like 70 to 80% deterministic local extraction, 20 to 30% cloud fallback and about 5% human review, cutting API cost by roughly 75%.

The small language model tier

Distilled 1.5B to 7B models fill the local tier for classification, extraction, summarisation and simple Q&A. Apple ships this shape at billion-device scale, resolving about five in six requests on-device and escalating the rest. That routing is itself a governance decision about which data may leave the machine.

A ~$5,000 box earns its keep as the cheap, private long tail of a system whose expensive moments are routed upward. Seven to 14B models absorb the routine work without any of it crossing the network. The moment concurrency or a hard prompt appears, the answer is to route rather than buy a second box and hope.

The useful question is where each query should run. The next decision is a governance decision, and it sits inside the local-first architecture story.

Frequently Asked Questions

How do I work out how much VRAM or unified memory a local model needs?

Start with the weight size at your chosen quantisation, then add room for the KV cache. A rough guide is about 0.5 to 0.6GB per billion parameters at 4-bit, so a 7B model needs roughly 4GB of weights while a 70B needs about 40GB, which is why the DGX Spark’s 128GB unified memory holds a 70B Q4 comfortably. Leave 20 to 30% spare for long prompts, because context length eats that headroom fast.

Does Q4 quantisation hurt output quality enough to matter?

Usually not for the workloads local hardware is good at. Four-bit quantisation trades a small amount of accuracy for roughly a fourfold reduction in memory use, which is what lets a 32B or 70B model fit at all. For classification, extraction, summarisation and simple Q&A the quality drop is rarely noticeable. Push to Q3 or below and reasoning starts to degrade, so treat 4-bit as the practical floor.

Can a decent CPU run an LLM locally without a GPU?

It can, but slowly, and it is not the setup you want for anything interactive. CPU inference is bound by system memory bandwidth just like a GPU, and typical desktop memory delivers a fraction of the 273 GB/s in a DGX Spark or the 1.79 TB/s in an RTX 5090. A small 3B to 7B model at Q4 will reply in a few tokens per second, fine for background jobs but painful in conversation.

What’s the difference between an NPU and a GPU for local AI inference?

An NPU is a low-power accelerator tuned for small, efficient models, while a GPU is built for high-bandwidth parallel work. NPUs win on power draw and cost, which makes them a natural fit for laptops, AI PCs and always-on edge devices running a 1B to 3B model. They lose on memory bandwidth and capacity, so they cannot hold large models or serve the concurrency a discrete GPU can.

Is a Mac Studio a better buy than a DGX Spark for running large models locally?

It depends on the workload, not the spec sheet. An M-series Mac with up to 192GB unified memory and 546 to 819 GB/s of bandwidth gives comfortable single-user 70B decode, and MLX runs 20 to 87% faster than llama.cpp on that hardware. The DGX Spark is slower per token but sits in the deeper CUDA tooling ecosystem. Choose the Mac for solo work, the Spark for serving.

Why does the same model run at different speeds on different machines?

Because batch-1 decode is bound by memory bandwidth, not raw compute. Generating each token means re-reading the model weights from memory, so throughput tracks how fast memory can feed the chip. That is why an RTX 5090 with only 32GB of VRAM runs a 32B model at 90+ tokens per second while a box with more capacity but less bandwidth crawls. Capacity decides what fits, bandwidth decides how fast.

What happens if the cloud goes down while I’m relying on hybrid inference?

Well-designed routing degrades to local rather than stopping. Because the local tier already handles the trivial majority (classification, extraction, simple Q&A), the system keeps working on those tasks when the cloud call fails. The hard minority either queues, retries, or falls back to the best local answer with a confidence flag. The real risk is silent degradation, so log escalations and surface them rather than hiding them.

Is hybrid cloud-edge inference safe enough for sensitive health or financial data?

Yes, if the routing rules are explicit about what may leave the machine. The whole point of a local tier is that sensitive records never cross the network; only genuinely hard, de-identified prompts are escalated. For regulated work you would pair that with data-residency controls and, where required, a private cloud endpoint or an air-gapped local-only mode. The governance decision matters more than the hardware.

When does local AI hardware break even against cloud API costs?

It depends on volume and utilisation, and it is rarely a clean comparison. A ~$5,000 box only pays for itself if it stays busy, and an idle GPU is the most expensive asset in the rack. Heavy steady traffic, the kind that can put $12,000 to $19,000 a month through APIs, favours local for the cheap long tail, while spiky, occasional work usually stays cheaper in the cloud. Model the utilisation, not the sticker price.

What’s the smallest local model that’s genuinely useful in production?

Distilled 1.5B to 7B small language models are already production-grade for the right jobs. They handle classification, extraction, summarisation and simple Q&A at 30 to 70 tokens per second on modest hardware, which is exactly the deterministic long tail a router keeps on-device. They will not do long reasoning or open-ended generation, and that is fine: those are the queries you escalate. Use the smallest model that passes your accuracy bar.

How long will a local AI box stay useful before it becomes obsolete?

Longer than a spec-sheet buyer expects, because the model tier moves faster than the hardware. A box that comfortably runs 7B to 14B models today will still run the next generation of small models, since the local tier is defined by the cheap, deterministic workloads rather than frontier scale. What ages it is the ceiling you already hit: capacity and bandwidth stay fixed. Plan to route and it stays useful.

AUTHOR

James A. Wondrasek James A. Wondrasek

SHARE ARTICLE

Share
Copy Link

Related Articles

Need a reliable team to help achieve your software goals?

Drop us a line! We'd love to discuss your project.

Offices Dots
Offices

BUSINESS HOURS

Monday - Friday
9 AM - 9 PM (Sydney Time)
9 AM - 5 PM (Yogyakarta Time)

Monday - Friday
9 AM - 9 PM (Sydney Time)
9 AM - 5 PM (Yogyakarta Time)

Sydney

SYDNEY

55 Pyrmont Bridge Road
Pyrmont, NSW, 2009
Australia

55 Pyrmont Bridge Road, Pyrmont, NSW, 2009, Australia

+61 2-8123-0997

Yogyakarta

YOGYAKARTA

Unit A & B
Jl. Prof. Herman Yohanes No.1125, Terban, Gondokusuman, Yogyakarta,
Daerah Istimewa Yogyakarta 55223
Indonesia

Unit A & B Jl. Prof. Herman Yohanes No.1125, Yogyakarta, Daerah Istimewa Yogyakarta 55223, Indonesia

+62 274-4539660
Bandung

BANDUNG

JL. Banda No. 30
Bandung 40115
Indonesia

JL. Banda No. 30, Bandung 40115, Indonesia

+62 858-6514-9577

Subscribe to our newsletter