Video RAM Killed the Radio Star
Tier 00 — I See You

Video RAM Killed the Radio Star

Sooooo.... what's VRAM?

Nobody sets out to learn about VRAM. You back into it. You are mid-game and the frame rate falls off a cliff, or a model you were promised would run on your card refuses to load, or someone in finance asks why the graphics card line on the bill doubled and you realise you cannot explain it either.

A gamer, an engineer, a platform team and an accountant walk into the same problem. None of them use the same words for it. All of them are talking about a few gigabytes of memory soldered to a circuit board next to a chip, and almost none of them can say what it does.

That memory is the most important number in modern computing that nobody can define at a dinner party. It decides which models exist for you. It decides how fast the answer comes back. It sets the price of the card, the size of the cluster, and, since the memory makers started feeding the AI datacentres first, the price of your console and your laptop too. It is a very small thing that has quietly become everyone's problem.

So this is the whole thing in one climb. It starts with a desk and an artist and ends with a Kubernetes cluster that budgets memory the way a bank budgets cash. Five tiers, and they are a spectrum, not a difficulty setting:

01 The Origin Novice For anyone who has heard the word and wants it explained with a desk and an artist. →
02 The Vector Explorer For the person who wants to know what is actually on the card and why it is glued down. →
03 The Nexus Intermediate For the hobbyist about to spend money, who wants to know why a mini PC can beat a graphics card at one thing and lose at everything else. →
04 The Apex Advanced For the engineer who has to make a model fit, run fast and not fall over, and who needs the arithmetic. →
05 The Zenith Expert For the architect who has to make all of that work for an organisation: many cards, many teams, a pipeline that runs while everyone is asleep. →
06 The Reference Look it up Every term the guide uses, and every subject it points at without covering. →

Read to the tier you need and stop. Come back when the next one starts to itch. Nobody is grading you.

Where a section leans on something you might not know yet, there is a short box explaining it in a few sentences and telling you it deserves a study of its own. Those boxes are the map of the guides that follow this one.

Figures are current to mid-September 2026.

Tier 01 · Novice

The Origin

A computer has two workers. Each needs a desk for whatever they are working on right now. That desk is memory.

A computer has two workers.
Opens programs, moves files, talks to the network, keeps the lights on. Bit of everything. Paints every picture you see on screen, sixty or more times a second. Lately it also does the maths behind AI models, which turns out to be the same kind of work: millions of identical operations at once.

Standard RAM is the manager's desk. Built so the manager can grab any single item instantly.

VRAM is the artist's desk. Built so the artist can pour a whole bucket of material across it at once, because painting a picture means touching millions of dots at the same time.

One desk cannot be best at both jobs. So each worker gets a desk shaped for their work, and the artist's desk is bolted directly to the artist's chair so nothing has to travel across the room.

Bolted to the chair
Across the room
Same figure, two speeds. The dots on the second rail move at the pace the picture stutters.
Two things matter about the artist's desk: how big it is, and how fast material can be poured across it.

If the artist's desk fills up, one of two things happens. Sometimes the overflow goes on the manager's desk across the room, and every time the artist needs it someone has to walk over and fetch it. That walk is slow and the picture stutters. Other times the artist just stops and says there is no room. Which one you get depends on what kind of work it is.

You can buy more sticks and plug them into the manager's desk. The artist's desk is glued to the artist.

To get a bigger one, you buy a different artist.

Everything below is that story with numbers.
Tier 02 · Explorer

The Vector

VRAM stands for Video Random Access Memory. It is memory that lives physically on the graphics card, wired straight to the GPU, and used only by the GPU.

System RAM lives on the motherboard, wired to the CPU, and is used by everything else.

The name is left over from the 1980s and 1990s, when "VRAM" meant a specific chip design: dual-ported DRAM that let the CPU write one frame into memory while the display hardware read the previous frame out at the same time. That design is long dead. Today "VRAM" just means "whatever memory is bolted to the GPU", regardless of the technology underneath. You will still find consumer explainers repeating the dual-port story as if it describes the GDDR7 on a current card. It does not. Treat it as history.

For graphics
▸The framebuffer: the image being sent to the monitor, plus the one being drawn behind it ▸Textures: the images wrapped onto 3D objects ▸Geometry: the vertices and triangles that make up the scene ▸Shader programs and their intermediate data ▸Depth buffers, shadow maps, post-processing buffers
For compute and AI
▸Model weights: the billions of numbers that make up a neural network ▸Activations: intermediate results as data flows through the network ▸The KV cache: the memory an LLM keeps of the conversation so far ▸Gradients and optimiser state during training
The framebuffer is literally the grid of pixels about to be shown. A texture is an image stretched over a 3D shape so a flat triangle looks like brick. A shader is a tiny program the GPU runs per pixel or per vertex to decide colour and position. If graphics rendering interests you, the rendering pipeline is a subject worth a guide of its own.
The gap is not subtle.

A CPU runs a handful of threads that jump around unpredictably. It wants any single byte back as fast as possible. DDR system memory is optimised for low latency, measured in nanoseconds.

A GPU runs tens of thousands of threads doing the same operation on adjacent data. It does not much care whether one byte arrives in 100 or 400 nanoseconds. It cares whether it can move hundreds of gigabytes per second. Graphics memory is optimised for bandwidth and placed as physically close to the GPU die as manufacturing allows.

The RTX 5090 gives its 32 GB of GDDR7 about 1.8 TB/s. DDR5 system memory on a typical desktop manages 75 to 100 GB/s. That is a difference of around 20 times.
Latency is how long one request takes to come back. Bandwidth is how much data moves per second. A motorway has high bandwidth and high latency; a footpath to the corner shop is the reverse. Memory systems are tuned for one or the other, rarely both. This distinction runs through the whole guide.
Either throughput collapses by an order of magnitude, or the process dies outright.

When a workload needs more memory than the GPU can keep resident, what happens depends on the memory model. Graphics drivers page textures and buffers into system RAM and fetch them back over PCIe as needed, and CUDA Unified Memory allocations can migrate between device and host the same way. Ordinary CUDA device allocations do not spill; they fail with an out-of-memory error. When paging does happen it runs over PCIe, which peaks around 32 GB/s on 4.0 x16 and double that on 5.0, a small fraction of VRAM bandwidth. In games this shows up as stutter, texture pop-in and sudden frame-rate drops. In AI workloads it shows up either as throughput collapsing by an order of magnitude or as the process dying outright.

The second rail is drawn at the speed ratio it actually runs at. That is the stutter.
PCIe is the slot and bus a graphics card plugs into; on a conventional PC it is the main road between a discrete card and the rest of the machine. CUDA is NVIDIA's programming platform for running general computation on the GPU. Almost every AI framework sits on top of it. Both are worth understanding on their own.

VRAM chips are soldered to the card's circuit board around the GPU. There is no slot. The number on the box is the number you will ever have on that card.

There is no ninth square to add.
Unified memory trades peak bandwidth for capacity.

A discrete GPU is a separate card with its own VRAM.

An integrated GPU is built into the CPU package and has no VRAM of its own. It borrows a slice of system RAM. This is called shared graphics memory, or a unified memory architecture. Every laptop without a separate graphics chip works this way, as does every Apple Silicon Mac and most consoles. Unified memory trades peak bandwidth for capacity: a Mac Studio can hold 512 GB of unified memory, which no discrete card at any price can match.

Discrete · two pools
CPU GPU System RAM 75 to 100 GB/s VRAM up to 1.8 TB/s
PCIe ≈32 GB/s
Each worker owns its desk. Anything that has to cross between them goes over the slow road.
Unified · one pool
CPU GPU Unified memory to 512 GB · 1.2 TB/s on M5 Ultra
no crossing
One desk, shared. Nothing has to be copied across, and capacity no discrete card offers, at a fraction of its bandwidth.
Windows also reports a "shared GPU memory" figure for discrete cards. That is not a second kind of memory. It is the system RAM the driver is allowed to spill into.
Windows
Task Manager, Performance tab, GPU. "Dedicated GPU memory" is VRAM.
Linux with NVIDIA
nvidia-smi
The memory column shows used and total.
macOS
About This Mac. Apple Silicon reports unified memory, not VRAM.
Sources for this tier
Tier 03 · Intermediate

The Nexus

This tier is for the person who wants to buy something and understand what they are paying for. It is also where the interesting trade-offs live.

GDDR sits on the board. HBM sits on the package.
GDDR Graphics Double Data Rate
8 chips × 32-bit around the die
The consumer and workstation standard. Individual chips sit on the board around the GPU, each connected over a 32-bit interface. GDDR6 is mainstream, GDDR6X is a faster Micron variant on high-end RTX 30 and 40 series cards, and GDDR7 arrived with the RTX 50 series and RTX PRO Blackwell. Per-pin rates have climbed from around 14 Gbps on early GDDR6 to 28 to 32 Gbps on GDDR7.
Current RTX 50 cards use 2 GB GDDR7 modules; the 3 GB modules that would give a 5080-class card 24 GB exist and ship on the RTX PRO 6000 and laptop 5090, but as of September 2026 they cost roughly three times the 2 GB part and the desktop "Super" refresh that depends on them has slipped to a reported CES 2027 window.
HBM High Bandwidth Memory
dies stacked, 1,024-bit per stack
The datacentre standard. Memory dies are stacked vertically and connected with through-silicon vias, then placed on the same package as the GPU die rather than on the board. Each stack has a 1,024-bit interface (2,048 on HBM4), so bandwidth per stack is enormous and the whole assembly is compact. It costs far more per gigabyte and only makes sense at datacentre prices.
Generations: HBM2, HBM2e, HBM3, HBM3e, and now HBM4 on NVIDIA's Rubin, which entered production in early 2026 with cloud availability rolling out in the second half of the year.
A vertical wire drilled through a stack of silicon dies so they can talk without going out to a circuit board. It is the trick that makes HBM possible and one of the reasons it is expensive to make. Advanced packaging (CoWoS, chiplets, interposers) is its own rabbit hole and is currently the bottleneck of the entire AI hardware industry.
Bus width is the variable people forget.

Bandwidth is roughly memory clock multiplied by bits per clock multiplied by bus width. A 256-bit bus with GDDR6 at 16 Gbps effective gets about 512 GB/s. A 512-bit bus with GDDR7 at 28 Gbps gets about 1.8 TB/s. HBM3e on an H200 reaches 4.8 TB/s across six stacks. Bus width is why two cards with the same memory generation can have wildly different bandwidth, and why cheap cards with a narrow bus underperform their headline capacity.

256-bit with GDDR6 at 16 Gbps effective → about 512 GB/s. 512-bit with GDDR7 at 28 Gbps → about 1.8 TB/s. Six HBM3e stacks on an H200 → 4.8 TB/s. The headline capacity figure predicts none of this.
Bars drawn to bus width · the term the headline figure hides
256-bit GDDR6, 16 Gbps 512 GB/s
512-bit GDDR7, 28 Gbps 1.8 TB/s
1,024-bit HBM3e, one stack × 6 on H200
6,144-bit H200, six stacks 4.8 TB/s
Drawing the bars to width rather than bandwidth separates the width term from the clock term. Doubling the bus from 256 to 512 bits and moving GDDR6 to GDDR7 together give roughly 3.5×. One HBM3e stack is already four times the width of a consumer bus, and the H200 carries six, past the top of this chart, so the last bar is full-bleed rather than to scale.
These are the two numbers that matter, and they answer different questions.
Capacity decides whether a workload fits at all. If a model needs 40 GB and you have 32, it does not run on that card without intervention. Bandwidth decides how fast it runs once it fits. For most inference and plenty of graphics work, bandwidth is the actual limit, not compute.
Buying on TFLOPS alone. For inference, a card with lower TFLOPS and higher bandwidth frequently wins. The card with more compute loses to the card with less. That surprises people every time.
Part Memory Type Bandwidth
{{ r.part }} {{ r.mem }} {{ r.type }} {{ r.bw }}
The B200 shows up as both 180 GB and 192 GB depending on which NVIDIA document you read. Both are real; the HGX spec sheet says 180. Rubin's 22 TB/s is NVIDIA's target; supply-chain reporting says the first HBM4 shipments land nearer 20 TB/s.
Capacity against bandwidth · both axes logarithmic
32 64 128 256 512 0.1 0.25 0.5 1 2 5 10 20 RTX 4090 RTX 5090 RTX PRO 6000 A100 H100 SXM H200 B200 B300 Rubin Ryzen AI Max+ RTX Spark M5 Max M5 Ultra DDR5 desktop memory capacity, GB → TB/s ↑
Same fourteen parts as the table. Fit is the horizontal axis, speed the vertical one, and the unified-memory platforms sit in the bottom right: more capacity than any discrete card at a fraction of the bandwidth. The HBM parts climb almost straight up. Both axes are logarithmic, so each gridline is a doubling or better; the vertical spread from DDR5 to Rubin is roughly 250×.
In a game, a flipped bit is one wrong pixel for one frame. In a three-week training run, a flipped bit silently corrupts a checkpoint.

That is the case for error correction.

The consumer versus professional split is less clean than it used to be. GDDR7 on RTX Blackwell cards has single-bit error correction built in and always on, with no toggle and no performance cost, so a GeForce RTX 50 card does correct single-bit flips. What it does not have is the end-to-end RAS stack of a datacentre part: double-bit detection with error containment, dynamic page offlining, row remapping, and driver-level reporting you can alert on. Professional and datacentre GPUs expose all of that, plus validation and support. ECC is part of why they cost several times more, but it is not the main reason.

ECC is error-correcting code, extra bits stored alongside data so a single flipped bit can be detected and repaired. RAS stands for reliability, availability and serviceability, the datacentre umbrella term for everything that keeps a machine correct and running when hardware misbehaves. Both matter more the longer a job runs.
A mini PC can load a model that the most expensive consumer graphics card cannot. It will then run it at a fifth of the speed.

That is unified memory in one sentence, and it has become the practical answer to the consumer VRAM ceiling. Three platforms matter right now.

Announced 25 August 2026, goes to 512 GB at 1.2 TB/s. That is a 50% bandwidth jump over the M3 Ultra and enough to run quantised 70B dense models at long context, or quantised 400B-class MoE models, entirely in memory. Strix Halo puts 128 GB of LPDDR5X on a 256-bit bus in a mini PC, of which up to 96 GB on Windows or around 108 GB on Linux is assignable to the GPU. It fits quantised 70B and 120B-class models but runs dense 70B at Q4 at roughly 5 tokens/s. MoE models, which stream far fewer bytes per token, are where it makes sense. Announced with Microsoft on 31 May 2026 and due in machines this autumn, a Grace Blackwell superchip for Windows on Arm: 20-core Arm CPU, a Blackwell GPU with 6,144 CUDA cores, up to 128 GB of LPDDR5X at ~300 GB/s. Same trade as Strix Halo with a better software story.
Capacity that no discrete card can offer, at bandwidth a quarter to two-thirds of a discrete card. They fit big models. They do not race them. A mixture-of-experts (MoE) model is built from many sub-networks and only wakes a few of them per token, so it reads far fewer bytes per token than a dense model of the same total size. Quantised means the model's numbers have been stored at lower precision to make it smaller. Both get proper treatment in the Apex tier; both deserve a guide each.
This is worth its own section because it changes the economics of everything below.

Samsung, SK hynix and Micron have moved wafer capacity to HBM for AI accelerators, where margins are far higher. HBM consumes three to four times the wafer area per gigabyte of standard DRAM, and hyperscaler orders are effectively open-ended. The result is DRAM and GDDR tightness that TrendForce expects to run through 2027, with SK hynix warning 2027 could be the worst year of it and some vendors pointing at 2028 before normalisation. NAND has been caught up in the same cycle but is expected to ease sooner, in the second half of 2027.

Concrete effects as of September 2026: memory is now the largest single component cost on a high-end consumer card. The RTX 5090 has traded at roughly double its launch price through the summer. Tom's Hardware reported that NVIDIA moved memory procurement onto board partners in late 2025, leaving them to buy GDDR7 at spot prices. NVIDIA has not confirmed it. Apple pulled the 512 GB Mac Studio option for months before restoring it with the M5 Ultra. Console makers have raised prices citing memory cost. HBM is sold out for 2026.

What the reporting says, on one axis
HBM sold out
DRAM and GDDR tightness, TrendForce
SK hynix: 2027 the worst year
RTX 50 Super, reported window
NAND expected to ease
Some vendors: normalisation
2026 H2 2027 2027 H2 2028 2028 H2
now, September 2026 forecast, no firm date
Every bar is a vendor or analyst statement as reported, not a measurement. The overlap is the useful part: nothing in the reporting has GDDR loosening before the RTX 50 Super window, and the earliest any of it points to normalisation is 2028. These are the dates most likely to move.
The implication is simple: VRAM is not just the scarce resource inside the cluster, it is the scarce resource in the supply chain. Lead times, capacity reservations and right-sizing matter more than they did two years ago.
VRAM does not pool across cards by default.

Two 24 GB cards are not one 48 GB card. Software has to split the work deliberately. NVLink (datacentre cards, and some older consumer cards) makes that split efficient at hundreds of GB/s; over plain PCIe it is much slower. RTX 40 and 50 series consumer cards have no NVLink, so multi-card builds with them are PCIe-bound.

· ≠ no NVLink on RTX 40/50 · PCIe-bound
NVIDIA's private high-speed link between GPUs, many times faster than PCIe, that lets several cards behave almost like one. Its absence on consumer cards is one of the clearest lines between hobbyist and datacentre hardware.
Sources for this tier
Tier 04 · Advanced

The Apex

This tier is for the engineer who has to make a model fit, run fast, and not fall over. It is also where the arithmetic starts to bite.

A model with X billion parameters needs roughly 2X GB of VRAM for its weights at 16-bit precision.

The practical rule from Hugging Face's documentation: a model with X billion parameters needs roughly 2X GB of VRAM for its weights at 16-bit precision. A 7B model needs about 14 GB. A 70B model needs about 140 GB, which is why a single 80 GB H100 cannot hold one at full precision and two are the floor.

70B at FP16, 128K context, one sequence

That is weights alone. Add:

▸The KV cache, which grows with context length and batch size
▸Activations for the current batch
▸Framework overhead: CUDA context, cuBLAS workspaces, allocator fragmentation. Budget 1 to 2 GB minimum
▸For training: gradients (same size as weights) and optimiser state. Adam keeps two extra copies, so training at full precision needs roughly four times the weight footprint before you even count activations
Inference against training · 70B at FP16 · weights only, before activations
Inference: weights resident, nothing else 140 GB
weights 140 GB
Training with Adam: four copies of every parameter ≈560 GB
weights gradients Adam m Adam v
Gradients are the same size as the weights, and Adam keeps two more copies of its own, so training a model needs roughly four times the footprint inference does, before a single activation is counted. The 70B that needs two H100s to serve needs eight to train this way, which is why optimiser sharding and mixed precision exist.
A model's parameters are the numbers it learned. Most of them are weights; the rest are biases, scale factors and other learned terms. In sizing conversations people say "weights" for the whole lot, which is a simplification you will now recognise. Activations are the intermediate values produced as an input passes through the layers. Training also needs gradients (how each weight should change) and optimiser state (the optimiser's own bookkeeping). If you have never looked at how a neural network is actually stored in memory, that is the study that makes every number in this tier obvious.
Precision is bits per parameter.
Format Bytes per parameter 70B model weights
{{ q.fmt }} {{ q.bytes }} {{ q.total }}

Quantisation converts weights from 16-bit to 8-bit or 4-bit. It cuts VRAM by 50 to 75% and, because inference is bandwidth-bound, usually raises tokens per second too, since fewer bytes get read per token. The cost is some accuracy loss, which varies by model and method. Common formats: GGUF (llama.cpp, Ollama), AWQ and GPTQ (vLLM and other GPU servers), and FP8 natively accelerated on Hopper, Blackwell and Rubin, with native FP4 (NVFP4) starting on Blackwell and continuing on Rubin.

A 70B model at FP16 is about 140 GB of weights alone, so even a 141 GB H200 is too tight for practical serving once the CUDA context, allocator and KV cache are counted; it needs two cards or a quantised copy. At Q4 the same model fits on a pair of RTX 5090s or a single 96 GB workstation card with room for context. These are number formats. The letters say how the number is stored (floating point or integer) and the digits say how many bits it takes. Fewer bits means smaller and faster but less precise. Quantisation is the craft of dropping bits without dropping quality, and it is a whole discipline with its own guide coming.
It is the most misunderstood number in LLM sizing.

A 128K-token context on Llama 3.1 70B with an FP16 KV cache costs about 40 GB of memory before the model itself. Other 70B designs and lower KV precisions come in far smaller, but the shape of the problem is the same. That is why long conversations eat serving headroom, and why chat systems truncate, compress or refuse once they hit a context limit. It is also the most misunderstood number in LLM sizing.

During generation, an LLM keeps the key and value tensors for every token in the context so it does not recompute them. That store is the KV cache. Size per token, in bytes:

2 × layers × KV heads × head dimension × bytes per element
2 × 80 × 8 × 128 × 2 = 327,680 bytes, roughly 320 KB per token. A 128K-token context therefore costs around 40 GB per sequence, on top of the weights. Multiply by concurrent sequences and the KV cache, not the weights, becomes the dominant VRAM consumer on a serving node.
KV cache by context length · one sequence · 320 KB per token
8K 2.5 GB
16K 5 GB
32K 10 GB
64K 20 GB
128K 40 GB
KV cache by concurrency · 32K context each
1 seq 10 GB
2 seqs 20 GB
4 seqs 40 GB
8 seqs 80 GB
Both axes are linear and every step is a doubling, which is the whole point: the KV cache scales with context and with concurrency independently, while the 140 GB of weights stays fixed. Eight users at 32K cost the same 80 GB as one user at 256K. Neither bar includes the weights.
If you see 160 KB per token quoted for this model, that is the INT8 figure, not FP16. Models using grouped-query attention or multi-head latent attention have far smaller KV footprints per token. That is a deliberate design choice to reduce VRAM pressure, and Llama 3.1's 8 KV heads against 64 query heads is already GQA. Full multi-head attention on a model this size would need around 2.5 MB per token.
A token is a chunk of text, roughly three-quarters of a word. Attention is the mechanism that lets each new token look back at every previous one; keys and values are the two things it looks at. The KV cache is where those are kept so they are not recomputed every step. This one concept explains most of the memory behaviour of every LLM you will ever run, and it is the next guide in this series.
At batch size 1 the compute units mostly sit idle waiting on memory.

Autoregressive decode produces one token at a time, and each step reads the entire weight tensor from VRAM. At batch size 1 the compute units mostly sit idle waiting on memory. The ceiling is:

tokens per second ≈ memory bandwidth ÷ bytes of weights read per token
32B at FP16 · ≈64 GB of weights · single sequence
Same model, same architecture, faster memory, proportionally faster output. Run the same division for a 70B FP16 model and you get 24, 34 and 57, which is the figure you will see quoted in vendor guides, but note that 140 GB of weights does not fit on an H100 or comfortably on an H200, so those numbers are a pure bandwidth thought experiment.

Batching amortises the weight read across many sequences, which is how serving frameworks hit thousands of aggregate tokens per second on one card.

Prefill (processing the prompt) is compute-bound rather than memory-bound. Hold that thought; it matters when scheduling.

Prefill is the model reading your prompt, all tokens at once, which is heavy on compute. Decode is the model writing its answer, one token at a time, which is heavy on memory bandwidth. Almost every serving optimisation exists because these two phases want different hardware.
Serving frameworks expose these as flags.
Colour is which device holds which block. Tensor parallelism cuts across every layer, so every layer needs a synchronisation. Pipeline parallelism cuts between layers, so only the boundaries talk.
Every layer requires a synchronisation, so it is communication-heavy and strongly benefits from a high-bandwidth, low-latency interconnect such as NVLink or NVSwitch. It runs over PCIe or RDMA fabrics too, just slower. Usually kept within a node. Puts different layers on different GPUs. Less communication, but GPUs idle waiting for their turn unless micro-batched. Works across nodes. Replicates the whole model on each GPU and splits the batch. Needs the model to fit on one card. Used for training and for scaling serving throughput. Places different experts of a mixture-of-experts model on different GPUs.
Tensor parallelism across four 24 GB cards behaves like roughly 96 GB minus per-card overhead, but only if the interconnect keeps up.
Three mechanisms exist to run several workloads on one physical card.
The driver context-switches between processes. No memory isolation; every process sees the whole card and can starve the others. Cheap, simple, fine for dev clusters. Several CUDA processes share the GPU's execution resources so their kernels run concurrently instead of being time-sliced. On Volta and newer, each client owns its own GPU address space and submits work directly to the GPU; the MPS server holds the shared scheduling resources. Better utilisation; memory limits can be enforced per client, but fault isolation is weaker than MIG. Available on A100, A30, H100 variants, H200, B200, GB200 and the RTX PRO 6000 Blackwell series. The card is partitioned in hardware into up to seven instances, each with its own slice of VRAM, compute and cache, and with the performance and fault isolation that time-slicing and MPS do not offer.
H100 80GB · a valid mixed layout
On an H100 80GB the valid profiles are 1g.10gb, 1g.10gb+me, 1g.20gb, 2g.20gb, 3g.40gb, 4g.40gb and 7g.80gb. The card has seven compute slices and each profile consumes the number in its name, so a mixed layout has to add up to seven: one 3g.40gb, one 2g.20gb and two 1g.10gb is a valid and common production config. Seven 1g.10gb instances is the other common one.

MIG is the only one of the three that gives a hardware-isolated, fixed memory partition with its own bandwidth. MPS narrowed the gap in 2026: CUDA 13.4's MPS v3 adds cgroup-integrated device-memory limits, so an allocation that exceeds the hard limit fails with an out-of-memory error rather than starving a neighbour. That is a real budget, but it is enforced in software over a shared card, and a fault in one client can still take down the others.

The three mechanisms, side by side
Mechanism Memory limit Fault isolation Hardware partition Configured
Time-slicing ○none ○none ○no Device plugin
MPS ◐per client, v3 ◐weaker ○no Device plugin
MIG ●fixed slice ●full ●yes GPU Operator
●enforced in hardware ◐enforced in software ○not available
Read the middle column first: CUDA 13.4's MPS v3 gives a real per-client memory budget, but it is enforced over a shared card, so a neighbour's fault still reaches you. Only MIG moves both of those into hardware. Time-slicing offers neither and is a dev-cluster tool.
Linux control groups, the kernel feature containers use to cap how much CPU, memory and now GPU memory a process tree can take. If you run anything in Docker or Kubernetes you are already using them without seeing them. Worth an afternoon.
Two teams both wrote correct code. One team's test suite kept failing on a card that was fine. The other team's process was holding 78 GB of memory it was not using.

Here is how you see that.

nvidia-smi          # point-in-time
nvidia-smi dmon     # streaming
nvtop               # interactive, across cards
▸DCGM (Data Center GPU Manager) and its exporter for Prometheus: memory used, memory bandwidth utilisation, ECC errors, temperature, power. This is the basis for all cluster-level GPU observability
▸Framework level: PyTorch's memory summary reports allocated versus reserved, which exposes allocator fragmentation
"Used" memory as the driver reports it includes what the framework's caching allocator has reserved but is not actively using. A process can appear to use 78 GB while the model needs 50. The allocator keeps freed blocks to avoid slow reallocation. Two processes both doing this on one card will fight, and neither will win. Prometheus is the standard open-source metrics database in cloud infrastructure. An exporter is a small program that turns something's internal numbers into a format Prometheus can scrape. DCGM's exporter is how GPU memory numbers end up on a dashboard. Observability is a field of its own and pays for itself quickly.
Quantise, shard, or change hardware. Weights fit, but KV cache or activations exceed the remainder as concurrency or context grows. Cap max context, cap concurrent sequences, or reserve headroom. Enough total free memory exists but no contiguous block large enough. Restart or tune the allocator. On unified-memory or oversubscription setups the process keeps running at a tenth of the speed instead of failing. On datacentre cards, a rising correctable error rate can indicate deteriorating memory and is worth monitoring. An isolated corrected bit is not a failure prediction; a trend is a warning.
Sources for this tier
Tier 05 · Expert

The Zenith

This tier is for the architect. Everything above was about one card and one workload. Here VRAM becomes a resource an organisation has to allocate, isolate, meter, autoscale and gate, across many cards, many teams and a pipeline that runs without anyone watching.

The short version: it is the system that takes a fleet of machines and turns them into one pool where you describe what you want to run and it decides where. The unit of work is a pod (one or more containers), machines are nodes, the component that places pods is the scheduler, and a DaemonSet is a pod that runs on every node. Kubernetes is a large subject and a prerequisite for the rest of this tier.
Kubernetes does not know what a GPU is.

Historically, GPU scheduling has relied on a device plugin that tells the kubelet "this node has four of something", and most clusters still work that way. DRA is replacing that model. The core went GA in Kubernetes 1.34 back in 2025; what 1.37 added this year is the piece that makes migration practical, a DRA driver answering the old nvidia.com/gpu request directly. Details further down this section.

How the plugin works

The vendor provides a device plugin (a DaemonSet) that discovers GPUs on each node and advertises them to the kubelet as an extended resource, conventionally nvidia.com/gpu (AMD uses amd.com/gpu). A pod requests an integer count of that resource in its container limits. The scheduler places the pod on a node with enough unallocated units and the device plugin mounts the device files into the container.

Three properties of this model shape everything downstream.

There is no nvidia.com/vram: 20Gi. You cannot ask for "a card with at least 24 GB free". You ask for one GPU and get whatever that node happens to have. Requests must equal limits for extended resources. A node with four cards schedules at most four GPU pods regardless of how little VRAM each actually uses. A pod that requests no GPU gets no GPU, even on a GPU node. Non-GPU workloads will happily land on expensive GPU nodes and eat the CPU and memory, which is why GPU nodes get tainted.
A label is a tag on a node ("this one has 80 GB cards"). A taint is a repellent on a node ("nothing lands here unless it explicitly tolerates me"). Together they are how you steer expensive hardware to the workloads that deserve it.

Making memory size visible to the scheduler

Since the resource is opaque, VRAM capacity has to be exposed as node labels so pods can select on it. NVIDIA's GPU Feature Discovery (part of the GPU Operator) labels each node with product name, memory size, compute capability, driver version and MIG configuration. A pod that needs an 80 GB card uses a node selector or affinity on the memory label. Without this, a pod requesting one GPU can land on a 16 GB T4 when it needed an 80 GB H100, and OOM on start.

Separate node pools per GPU class, each tainted, with workloads tolerating the taint and selecting on the label. Mixed pools with scheduler roulette are a common cause of "works on my cluster, fails in production".

Giving a pod less than a whole card

Three ways, mapping to the three sharing mechanisms from the Apex tier.

A card is advertised as, say, four units of nvidia.com/gpu. Four pods each get one unit and share the card with no memory isolation. Fine for notebooks and low-traffic dev services. Never for anything with a VRAM guarantee. Configured similarly and adds concurrent kernel execution plus per-client memory limits. Better throughput than time-slicing for many small inference workloads, weaker fault isolation. Configured on the node via the GPU Operator's MIG manager and each instance is advertised as its own resource: either as generic nvidia.com/gpu (single strategy) or with the profile in the name, such as nvidia.com/mig-1g.10gb (mixed strategy). A pod requesting one mig-2g.20gb gets exactly 20 GB of hardware-isolated VRAM. This is the only way to give a Kubernetes pod a hardware-isolated, fixed VRAM partition below a full card. MPS v3 can enforce a hard software budget on a shared card; MIG is the only actual partition.

Changing MIG geometry is disruptive. The GPU Operator's MIG Manager handles the reconfiguration and restarts the Operator-managed GPU components, but user workloads have to be evacuated first, and on some driver and platform combinations a reboot is required. There is no fixed duration you can promise. I treat MIG geometry as node-pool infrastructure configuration, dedicate pools to fixed layouts, and do not let workloads change them.

The plugin's replacement: Dynamic Resource Allocation

DRA is the replacement for the device plugin model, and as of September 2026 it is real enough to plan around.

The timeline: core DRA APIs went GA in Kubernetes 1.34 (August 2025) in the resource.k8s.io/v1 group. Kubernetes 1.37, released 26 August 2026, graduated three more pieces to GA: DRA Extended Resource support, device taints and tolerations, and ResourceClaim status. Extended Resource support is the one that matters operationally. A DRA driver can now satisfy a plain nvidia.com/gpu: 1 request in a pod spec with no ResourceClaim on the workload and no device plugin running beside the driver. Tenants keep their manifests; the allocation underneath moves to DRA. That removes the migration blocker that kept most clusters on the device plugin.

The model: workloads describe what they need through ResourceClaims and ResourceClaimTemplates, drivers publish structured device attributes as ResourceSlices, and the scheduler filters and selects on them. Practically, this is "give me any GPU with at least 40 GB and compute capability 9.0" as a first-class scheduling constraint. Device taints let a driver or admin mark degraded hardware so nothing new lands on it.

NVIDIA's side: GPU Operator 26.7.0 manages the DRA driver natively through two new custom resources, NVIDIADriver and GPUCluster, instead of ClusterPolicy. A cluster runs one or the other, not both, and there is no in-place migration between them. Prerequisites are Kubernetes 1.34.2 or later, driver 580 or later, and a CDI-capable container runtime. Workloads can claim full GPUs and preconfigured MIG devices through ResourceClaims and select by attribute. The upstream DRA driver v0.5.0 adds multi-user MPS, NUMA attributes and system-mediated GPU sharing across namespaces via DRA's consumable capacity feature, though several of those remain alpha and off by default.

DRA on its own does not create fractional GPUs or guarantee safe sharing; it changes how devices are described and selected, not what the hardware can isolate. Managed Kubernetes support is uneven, and Karpenter is where it bites. Karpenter will not provision a new node for a pod whose only GPU request is a ResourceClaim; it can place that pod on a GPU node that already exists, and that is all. So on EKS, DRA works with Karpenter only when the node pool is static, or with managed and self-managed node groups. Dynamic Karpenter provisioning still needs the device plugin. EKS Auto Mode does not do DRA at all. The control plane supporting DRA and the autoscaler supporting DRA are two different questions; ask both, and expect to run the device plugin path and the DRA path side by side on separate node pools for a while.

Related and worth watching: the PodGroup gang-scheduling API reached beta in 1.37, though like other new beta APIs it is not on by default; it sits behind the GenericWorkload feature gate and the scheduling.k8s.io/v1beta1 API has to be enabled. Multi-GPU training jobs that need all their pods or none have been a persistent gap; this is the upstream fix.

Some jobs need all their pieces to start together or not at all; a distributed training run with seven of eight workers is just seven GPUs burning money. Gang scheduling is the scheduler understanding that. It has been solved by add-ons for years and is only now arriving in core Kubernetes.

Stopping one team from eating the cluster

▸ResourceQuota can cap requests.nvidia.com/gpu per namespace. This is the primary lever for stopping one team consuming the cluster.
▸LimitRange can set defaults, though for GPUs the useful default is zero.
▸Admission policy (Kyverno, OPA Gatekeeper, or ValidatingAdmissionPolicy) should reject GPU pods that lack a node selector on memory class, lack a toleration, or request more than a defined ceiling. It should also reject GPU pods without CPU and memory limits, because a GPU pod with unbounded system memory takes the whole node down when it spills.
▸Priority classes and preemption let production inference pre-empt batch training and CI jobs when capacity is tight.

Growing and shrinking the GPU fleet

Cluster Autoscaler and Karpenter both understand extended resources and will provision a GPU node when a pod is pending on nvidia.com/gpu. Two things bite.

Node start time on GPU instances is long: image pull for the driver container, driver load, device plugin registration, plus a serving pod that may need to pull a 40 GB model. Budget several minutes, set readiness accordingly, or pre-warm.

Scale-down behaviour is autoscaler-specific. Cluster Autoscaler has a separate GPU utilisation threshold for scale-down, based on requested GPU resources on accelerator nodes. Karpenter uses scheduling simulation and consolidation. Neither observes actual VRAM utilisation as a pressure signal, which is why DCGM metrics belong in application autoscaling and capacity policy, not in node scale-down. Annotate long-running GPU pods as non-evictable or use a pod disruption budget.

Karpenter's node pools can express GPU instance families and let the scheduler pick the cheapest instance that satisfies the memory label. Materially cheaper than fixed node groups when workload sizes vary.

The one thing to install first

On any non-trivial cluster, install the GPU Operator rather than hand-managing drivers. A default install deploys the driver as a container, the container toolkit, the device plugin, DCGM exporter and MIG manager on every GPU node, and keeps them version-aligned. Driver, CUDA runtime and container toolkit version skew is a common cause of GPU pods failing to start.
A model that fits on the 96 GB workstation card in staging and OOMs on the 80 GB card in production is not an ops incident. It is a test that was never written. Continuous integration and delivery is the automated pipeline that builds, tests and ships code every time someone changes it. A runner is the machine or container that executes a pipeline job. Adding GPUs to that pipeline is what this section is about; the pipeline itself is a discipline worth learning properly.

Why the pipeline needs a GPU at all

The moment a repository contains a model, a CUDA kernel, an inference service, or a data pipeline that runs on GPU, the pipeline needs GPU stages: unit tests that exercise CUDA paths, integration tests that load a model, regression tests on output quality, latency and throughput benchmarks, and image builds that bake in weights or compile kernels. On CPU these either do not work or take so long the pipeline is useless.

Three ways to give it one

Simple, always warm, expensive at idle. One or more VMs with a GPU and a runner agent installed. The GPU is time-shared between jobs by the runner's concurrency setting with no isolation. Acceptable for a small team and nothing larger. GitLab Runner's Kubernetes executor, GitHub's Actions Runner Controller, and equivalents spawn each job as a pod. The job definition sets nvidia.com/gpu in the pod's resource limits, the node selector for memory class, and the toleration. Each job gets a clean GPU, the cluster autoscaler brings GPU nodes up on demand and removes them after, and GPU quota is enforced at the namespace level like any other workload. This is the standard pattern and the one I default to. Each job provisions a GPU VM (often spot) via the runner's autoscaling driver, runs, terminates. Cheapest at low utilisation, slowest to start.

What the job needs to actually run

▸A container image with the CUDA runtime matching the node's driver, or a base image from the framework vendor
▸The container toolkit on the node so the device files get mounted (the GPU Operator handles this)
▸The runner's pod template or job spec passing the GPU request through. In GitLab this is set on the runner config or overridden per job with resource variables; in Actions Runner Controller it lives in the runner scale set's pod template
▸Model weights cached somewhere fast. Pulling 40 GB from object storage on every job is the dominant cost. Use a persistent volume, a node-local cache, or bake weights into a layer the node image cache keeps warm

Treating memory like a test

A mature pipeline treats VRAM the way it treats test coverage or bundle size: measured on every run, compared to a threshold, failing the build on regression.

Load the model, run a representative batch at production context length, record peak allocated memory via the framework's tracker, assert it is below the target card's capacity minus a headroom margin (10 to 15%). A change that pushes peak from 70 GB to 82 GB on an 80 GB target fails here, not in production. Run a fixed load, record tokens per second or frames per second, compare to the last main-branch baseline, fail on regression beyond tolerance. Because throughput is bandwidth-bound, a change that increases bytes read per token shows up here even when peak memory does not move. If the deployable is a quantised model, the pipeline produces it, measures size and accuracy on a held-out set, and refuses to promote if either regresses. Where production runs on more than one GPU class, run the memory test on each. That is the test from the opening line of this section.

Record these as pipeline artefacts and plot them. VRAM creep is gradual and only visible in trend.

Keeping jobs from poisoning each other

▸Never let two CI jobs share a GPU without MIG. A test that leaks memory in one job OOMs the next job on the same card and produces a flaky failure nobody can reproduce.
▸Have the job assert the GPU is empty at start (query the driver, fail if used memory exceeds a small threshold). This catches a previous job that did not clean up.
▸Kill the job on timeout. A hung CUDA process holds VRAM until it dies; a runner that leaves it behind silently poisons the node.
▸Pin the CUDA base image digest. Framework updates change memory behaviour.
Every cloud sells GPU capacity as a named instance shape, and every shape hard-codes the memory.

The job of infrastructure-as-code is to make "how much VRAM" a decision someone writes down, not an accident of which shape got picked.

IaC means describing servers, networks and node pools in files (Terraform, Pulumi, OpenTofu) that a tool applies, so infrastructure is versioned and reviewed like software. A quota is the ceiling a cloud lets your account provision. Spot capacity is spare hardware sold cheap that the cloud can take back at short notice. All three are foundational for cloud work.

Naming the memory class, not the shape

Define a variable or module per GPU class (for example "inference-80gb", "dev-24gb", "training-141gb-x8"), map each to the shapes that satisfy it per region, and have node pools and VMs reference the class, not the shape. Changing vendor or generation then becomes a mapping change.

Quota is a prerequisite, capacity is a rumour

GPU quota is separate from general compute quota on every major cloud and defaults to zero or close to it. Increases take days and are per region and per shape family. Treat quota as a prerequisite tracked alongside the IaC, and have plan-time checks (or a pre-apply script) fail fast when a plan exceeds it. Capacity is also not guaranteed: a region can have quota available and no physical cards, and in the current memory market that happens more than it used to. For production, use capacity reservations or committed capacity; for CI and batch, use spot with a fallback shape list.

Where the driver comes from

Node boot is fast and deterministic. Build a machine image with the driver, container toolkit and monitoring agent pre-installed, versioned and tested. Rebuilding the image is the driver upgrade path. Preferred for VM-based runners and standalone inference hosts. For Kubernetes, boot a plain image and let the GPU Operator install everything as containers. Slower first boot, but driver version becomes a cluster configuration value rather than an image rebuild.
A node with a host-installed driver and an operator trying to install a different one fails in confusing ways, and you will lose an afternoon to it.

Rules the pipeline enforces so people do not have to

▸GPU node pools must carry the taint and the memory-class label
▸Every GPU node pool must have a maximum node count
▸Spot GPU pools may only host workloads with a defined toleration for interruption
▸Every GPU resource must carry cost-allocation tags (team, workload, environment)
▸Production inference pools must use reserved capacity

Run these in the pipeline against the plan, before apply.

What you are actually paying for

GPU hours are the most expensive line on the bill and VRAM is what you are actually paying for. Three levers:

▸Right-size the class. A 24 GB dev card costs a fraction of an 80 GB datacentre card. Most notebooks and unit tests do not need the latter.
▸Scale to zero. Every GPU pool that is not production inference should reach zero nodes when idle. That includes CI pools.
▸MIG-partition large cards for small workloads rather than assigning a whole card to a 7B model.

Meter VRAM utilisation per team via DCGM and show it back. Allocated-but-idle VRAM is the usual waste pattern and it is invisible without the metric.

A serving framework is mostly a memory manager with an HTTP endpoint attached. Once you see it that way, every flag makes sense. vLLM, SGLang, TensorRT-LLM, Hugging Face TGI and NVIDIA Triton are programs that load a model once and answer many requests concurrently, doing the batching and memory juggling that a naive script would not. Picking and tuning one is its own subject.

What the framework does with the memory

Allocates the KV cache in fixed-size blocks rather than one contiguous region per sequence, removing fragmentation and letting far more sequences share a card. Admits new requests into the running batch as others finish, keeping the GPU busy and amortising the weight read. Shares KV blocks between requests with a common prompt prefix (system prompts, few-shot examples), cutting KV VRAM and prefill time. Runs the compute-bound prefill and the bandwidth-bound decode on separate GPU pools sized for each. This is the current frontier for large deployments and the payoff for the "hold that thought" in the Apex tier.
Phase one Prefill: reading your prompt
Compute saturated
Bandwidth slack
All prompt tokens go through the model at once, so the compute units are the limit and the weight read is amortised across the whole prompt.
wants compute
Phase two Decode: writing the answer
Compute slack
Bandwidth saturated
One token at a time, and every step reads the entire weight tensor again. The compute units mostly wait on memory.
wants bandwidth
The meters are illustrative, not measured: what matters is that they are mirror images. Run both phases on one pool and whichever resource the current phase does not need sits idle. Disaggregation puts each phase on hardware sized for the resource it actually saturates, and it is why a card with lower TFLOPS and higher bandwidth can win on decode, as the Nexus tier warned.

Setting the budget

Serving frameworks expose controls for their GPU and KV cache memory budgets, each in its own way. In current vLLM releases, gpu_memory_utilization is a fraction of total VRAM, defaults to 0.92, and is a per-instance limit. The framework loads weights, then allocates the remainder up to that fraction as KV cache. The arithmetic:

available KV ≈ (VRAM × utilisation fraction)
             − weights
             − peak activation reserve
             − non-PyTorch overhead (NCCL, backend buffers)
             − CUDA graph reserve
The same subtraction, on an H100 80GB serving a 70B at Q4
H100 80GB, total memory 80 GB
× utilisation fraction 0.92, the budget vLLM may touch 0.92 73.6 GB
− weights, 70B at Q4 35 38.6 GB
− peak activation reserve ≈2 36.6 GB
− non-PyTorch overhead (NCCL, backend buffers) ≈1 35.6 GB
− CUDA graph reserve ≈1.5 34.1 GB
= what the KV cache actually gets 34.1 GB
Each bar is what remains of the 80 GB card after that line. Two terms are arithmetic from figures elsewhere in this guide: 80 × 0.92 gives the 73.6 GB budget, and a 70B model at Q4 is 35 GB by the Apex precision table. The three middle terms are what startup profiling measures on your own box, shown at typical magnitudes: individually small, together 4.5 GB the KV cache does not get. At 320 KB per token the remainder holds roughly 100,000 tokens: three 32K conversations at once, not thirty.

That is what vLLM's startup profiling actually measures: it loads the weights, runs a dummy forward pass to find the activation peak, measures what was allocated outside PyTorch, estimates CUDA graph memory, and hands whatever is left to the KV cache. The startup log prints each term.

Divide by KV bytes per token to get the card's total token capacity, then divide by expected tokens per request to get concurrent sequence capacity. That number is the node's practical concurrency and context-capacity limit. It is not a speed; bandwidth sets speed, KV capacity sets how many conversations of what length can be resident at once. It should be an explicit input to capacity planning, not something you discover under load.

Unified-memory systems (DGX Spark, GH200, RTX Spark) are a different animal. There is no separate VRAM to budget; the model, the KV cache, the OS and its page cache all draw from one pool. Older vLLM releases refused to start because they counted reclaimable page cache as used GPU memory. Current releases account for UMA separately, and still get it wrong in new ways: as of September 2026 there are open issues where startup exhausts the pool while the host reports free memory. vLLM's own DGX Spark guidance is to start from the model-specific recipe and tune on measured headroom; its published Nemotron recipe runs at 0.85 with concurrency capped at four. The number is model- and release-specific; that is the lesson. Start from the recipe, leave room for the OS and runtime, pin the release, test on the real box, and do not assume a discrete-GPU default means anything here.

Set maximum model length explicitly. vLLM sizes the KV pool from the memory budget, not per request, so the setting does not reserve space in advance; what it does is bound the requests the server will accept and determine how much concurrency that pool supports at that context length. Do not advertise 128K if the application only needs 32K.

How the pods should look

▸Default to one serving pod per GPU (or per MIG instance, or per tensor-parallel group). Deliberate sharing is possible: vLLM's own docs show two instances on one card at 0.5 each, and MPS v3 can enforce the split. It still leaves a shared fault domain and needs explicit capacity control, so treat it as an exception you design, not a default you drift into.
▸Readiness probe that only passes once weights are loaded and a warm-up request has completed. Load can take minutes; routing traffic early produces timeouts.
▸Startup probe with a generous failure threshold for the same reason.
▸Pod anti-affinity across nodes for availability.
▸Priority class above CI and batch so serving pods pre-empt them under pressure.
▸Rolling update with max surge of at least one, and capacity to host that surge, or the rollout stalls waiting for a GPU.

Scaling on the right signal

CPU-based HPA is usually the wrong signal for GPU inference; the CPU can sit nearly idle while the GPU or the KV cache is saturated. Tokenisation and preprocessing can make CPU matter in some designs, but it is rarely the limit. Scale on a metric that reflects VRAM and GPU pressure:

▸DCGM GPU utilisation and memory utilisation via the Prometheus adapter
▸Framework metrics: vLLM exposes KV cache utilisation percentage, running and waiting request counts, and time to first token. KV utilisation sustained above roughly 80% is the standard scale-out trigger; waiting-queue length is the leading indicator
▸KEDA can drive replica count from any of these, or from queue depth for asynchronous inference

Scale-out lead time is node provisioning plus model load, typically five to fifteen minutes. Either keep headroom replicas or use predictive scaling on traffic patterns.

The Horizontal Pod Autoscaler is Kubernetes' built-in mechanism for adding replicas when a metric rises. KEDA extends it to scale on almost anything, including queue depth and custom GPU metrics. Autoscaling well is harder than it looks and is a subject in itself.

Shipping a model like software

A model version is a deployable artefact with a VRAM footprint, the same way a container image has a size. The promotion pipeline should:

1Build or fetch the model artefact and record its checksum and size
2Run the CI memory and throughput gates from 4.2 against the exact target GPU class
3Publish the artefact to a registry (OCI registry with ORAS, or an object store with immutable versioning)
4Deploy as a canary: one replica on production hardware with a small traffic share, monitored on KV utilisation, latency and quality metrics
5Promote or roll back. Because a new model version can have a different KV footprint per token, canary has to observe VRAM headroom at production context lengths, not just error rate

Weights should be mounted, not pulled, wherever possible: a read-only persistent volume populated once per version, or a node-local cache warmed by a DaemonSet before the rollout, so replica start is seconds not minutes.

What to watch and when to wake someone

Minimum metric set per GPU, scraped from DCGM:

▸Memory used and free
▸Memory bandwidth utilisation (not just compute utilisation; a bandwidth-bound card shows low SM utilisation while fully saturated)
▸ECC single-bit and double-bit error counts
▸Temperature and power (throttling reduces effective bandwidth)
▸XID errors from the driver log, which range from application bugs to hardware failure

Per serving replica, from the framework: KV cache utilisation, running and queued requests, time to first token, tokens per second, request failures by reason (OOM versus timeout versus error).

KV utilisation sustained above threshold, queue length growth, any double-bit ECC error, rising single-bit ECC rate, any XID indicating hardware fault, memory used above 95% on any card.
Those three questions are where most real clusters live.

The device plugin answers "is there a free GPU on this node". It does not answer "should this team get it before that one", "does this job need all eight or none", or "can this notebook have a quarter of a card". Those three questions are where most real clusters live, and the answers come from the scheduling and admission layer around Kubernetes: sometimes in front of the default scheduler, sometimes replacing it.

The default Kubernetes scheduler places one pod at a time onto one node. An admission controller such as Kueue decides whether a workload should be let into the cluster yet, based on quotas and fairness across teams. A batch scheduler such as Volcano or KAI replaces the default placement logic with gang scheduling, fairness and topology awareness. Kueue, Volcano and KAI are the three you will meet. Each is a subject in its own right.

Three tools, three layers

A job's path to a card
01
Kueue holds the job suspended until its queue has quota Admission
↓
02
Volcano / KAI picks the node, as a gang if the job needs one Placement
↓
03
Plugin or DRA mounts the device into the container Handover
↓
04
GPU Operator driver, toolkit, DCGM, MIG geometry Fabric
Each layer hands down to the next, and only the bottom two touch the GPU. That is why the three tools coexist rather than compete: swapping the placement layer for Volcano or KAI leaves admission and handover alone, and none of them replaces the GPU Operator underneath.
Jobs arrive suspended, Kueue checks them against a ClusterQueue with quotas and cohort fairness, and releases them to the ordinary scheduler when there is room. It does not touch the GPU itself. It is the answer to "stop team A's backlog from starving team B" and to "keep excess jobs suspended instead of leaving hundreds of pods pending". Kueue understands GPU memory and compute as quota dimensions when the resources are exposed that way. Gang scheduling, queues with resource caps, topology policies, built in. Its vGPU device plugin exposes shared GPUs with per-container memory and compute limits. Built on kube-batch. Gang scheduling, hierarchical PodGroups, fair share across queues, time-aware fairness, GPU sharing, and since v0.10.0 topology-aware scheduling. It supports DRA for GB200 and GB300 compute domains and integrates with Grove and Dynamo for disaggregated serving. It runs beside the default scheduler and takes only the workloads that ask for it.

All three coexist with the device plugin or with DRA. None of them replaces the GPU Operator.

Slicing a card in software

Apex covered time-slicing, MPS and MIG. There is another common software approach, and it is a library rather than a driver feature.

HAMi (a CNCF project) intercepts CUDA calls inside the container through a preloaded library and enforces a memory ceiling and a compute share that the pod requested in megabytes and percent. A pod asks for one vGPU with nvidia.com/gpu: 1, nvidia.com/gpumem: 8000 (MiB) and nvidia.com/gpucores: 30, sees 8 GB inside the container, and gets an out-of-memory error when it tries to take more. No MIG-capable hardware needed, no profile table, no reset to change the split, and it works on a consumer card. HAMi plugs into Kueue as a schedulable flavour, into Volcano through the vGPU plugin, and into KAI's resource isolator. A published test on an 8x RTX PRO 6000 Blackwell node under Kubernetes 1.35 turned eight cards into eighty schedulable slices and showed the OOM firing at the quota.

The library is enforcing a limit that the hardware does not know about, so anything that bypasses the intercepted CUDA path bypasses the limit: direct driver API calls, a misconfigured runtime, a container that sets the disable flag. HAMi's own documentation lists the ways the limit fails to apply. A neighbour's crash still shares the card. There is no bandwidth isolation.

The decision table an architect actually needs:

▸Notebooks, dev services, small inference, mixed hardware including consumer cards: HAMi or another software slicer, with Kueue on top for quota.
▸Production inference for tenants with different trust levels, or anything with a VRAM guarantee in a contract: MIG.
▸Many small kernels from one trusted tenant: MPS v3 with hard memory limits.
▸Nothing: time-slicing.
Linux lets you load a library ahead of a program's own, so calls it makes to a shared library can be caught and altered. HAMi uses this to sit between the application and the CUDA runtime. It is an old trick and a powerful one, and knowing how it works tells you exactly where it stops working.
In 2026 the answer to "where does that live" is no longer "in VRAM or nowhere".

A 128K-token context on Llama 3.1 70B at FP16 KV is about 40 GB. In 2026 the answer to "where does that live" is no longer "in VRAM or nowhere".

Two requests that begin with the same text (the same system prompt, the same document) share the same KV blocks for that prefix. A serving engine that notices can skip recomputing them. That is a cache hit, and it is the reason a busy chat product is far cheaper per token than the arithmetic in Apex suggests.

Why one card is the wrong unit

vLLM's built-in prefix cache is per replica. Ten replicas behind a load balancer each keep their own, so a request that lands on replica B gets nothing from what replica A computed a minute ago for the same 100K-token document. Agentic workloads make this worse: a coding agent runs a long multi-turn loop, each turn re-sends the whole history, and the working context outgrows any single card's KV budget. vLLM's own analysis of Codex-style traces on SWE-bench is what drove its Mooncake integration.

What replaces it

Sits under vLLM (and others) and moves KV blocks to wherever there is room: CPU RAM, local NVMe, Redis or Valkey, an S3-compatible object store, a Mooncake pool, or between workers over NIXL and RDMA, with storage tiers accelerated by GPUDirect Storage. Same blocks, shared across replicas. It also does non-prefix reuse (reusing cached blocks that appear mid-prompt, with selective recompute to keep quality) and carries the KV transfer for prefill/decode disaggregation. Its multi-node peer-to-peer CPU memory sharing moved from experimental to production in January 2026, the standalone multiprocess architecture followed in April, and it is in use at GKE Inference, CoreWeave and Cohere. It runs in-process for single-node offload, or as a standalone server that several vLLM instances share and that survives an engine crash. From the Kimi team. It is the pool LMCache can write to, the transfer backend SGLang and vLLM use for disaggregated prefill, and the thing under the Kimi K2 deployment on 128 H200s. vLLM featured Mooncake Store as a first-class KV backend in May 2026 and combined it with prefill/decode disaggregation in one deployment.
From LMCache's own technical report, measured on B200 machines and not a universal law: loading a cached KV block over the network beats recomputing it only above a crossover point that depends on link speed. At 32 Gbps between nodes the crossover was around 256K tokens; at 64 or 128 Gbps loading won at every context length tested. Cache tiering on a slow network can be slower than prefill. Measure the crossover on your fabric and your cards before turning it on.

The stack that makes this operational

Prefill/decode disaggregation (Apex, "hold that thought") plus a shared KV pool is now an orchestrated pattern rather than a research setup. NVIDIA's Dynamo is the inference runtime that does the routing, Grove is the Kubernetes API that wraps prefill, decode and router pods into one gang-scheduled unit with startup ordering, and KAI places the gang on hardware that can talk to itself fast. The point for VRAM: prefill pools want compute, decode pools want bandwidth, and the KV cache moves between them over the fabric. Buying identical cards for both is the mistake this stack exists to stop.

RDMA lets one machine write into another machine's memory without the CPU in the way. NIXL is NVIDIA's transfer library that does this between GPUs, CPUs and storage behind one interface. If the KV cache is going to leave the card and come back in time to be useful, this is the road it takes.
Nothing in a request for two GPUs says which pair you get.

Tensor parallelism across two GPUs on the same NVLink pair runs at NVLink speed. The same job on two GPUs that only share a PCIe root complex runs at PCIe speed, and across two that share nothing but the CPU it is worse. Nothing in nvidia.com/gpu: 2 says which pair you get.

The map of what is wired to what: which GPUs share an NVLink domain, which sit behind the same PCIe switch, which are near the same CPU socket and memory (a NUMA node), which network card is closest. On a single HGX board with NVSwitch every card reaches every other at full speed and the question does not arise. On a PCIe box, a mixed node, or anything spanning nodes, it decides whether the interconnect in Apex keeps up.

How the scheduler learns the map

The traditional device plugin knows a little. It can publish NUMA locality and preferred allocations for the node it runs on, and the kubelet's Topology Manager can use that to keep a pod's GPUs and CPUs together inside one node. What it cannot do is make a rack-scale NVLink domain, or a GPU-plus-NIC relationship, a first-class object the cluster scheduler reasons about. Placement of multi-GPU jobs across nodes under the device plugin is affinity heuristics, not scheduler logic.

DRA fixes this at the description layer: the NVIDIA DRA driver publishes GPU attributes including NVLink topology as structured parameters the scheduler can filter on. For multi-node NVLink (GB200 and GB300 racks), the NVIDIA DRA driver manages ComputeDomains, its abstraction for "this set of nodes shares one NVLink domain, keep the whole job inside it", and labels each node with nvidia.com/gpu.clique to say which partition it belongs to.

A scheduler still has to consume that. KAI's topology-aware scheduling reads a cluster-scoped Topology object whose levels run from widest to narrowest and can include the GPU clique and the hostname; it places a gang inside the requested level. NVIDIA's own NVCF stack creates exactly that mapping from the clique labels for its topology-aware path. Kueue has its own topology-aware scheduling built on a labelled node hierarchy. Volcano has a topology policy on its PodGroup. HAMi's scheduler avoids a poorly connected GPU for multi-GPU requests and offers it to single-GPU ones.

KubeCon NA 2025 spent a whole track on this: when GPUs, NICs, CPUs and memory are not aligned within the same NUMA node and PCIe root, distributed jobs lose a large fraction of their throughput. The number varies by workload; the sign does not.

What to actually do

▸Separate node pools by topology class the same way you separate them by memory class: NVSwitch boards, PCIe boards, single-card nodes. A tensor-parallel job never lands on a PCIe pool by accident.
▸For anything multi-GPU, use a scheduler that reads topology (KAI, Volcano, Kueue with labels) rather than the default one with affinity rules.
▸For GB200-class hardware, adopt DRA and ComputeDomains from day one; there is no first-class device-plugin abstraction for a rack-scale NVLink domain.
▸Put the NIC in the same constraint. GPUDirect RDMA wants the GPU and the network card on the same PCIe root, and DRA can express that as a joint claim.
A page to a human at three in the morning is the expensive answer.

The Zenith observability section says to alert on double-bit ECC, rising single-bit ECC, XID faults and thermal throttling. This section is what the platform does when those alerts fire, because a page to a human at three in the morning is the expensive answer.

An XID is an error code the NVIDIA driver writes to the kernel log when the GPU does something wrong; some mean the application crashed, some mean the hardware is failing. A Kubernetes node condition is a flag on a node ("Ready", "MemoryPressure") that the scheduler and other controllers read. GPU health work is largely about turning the first into the second.

Detect, quarantine, drain, fix

The shape every mature platform converges on, and the shape NVIDIA shipped as open source:

1Detect. DCGM raises policy violations (double-bit ECC, XID, NVLink failure, page retirement, thermal and power limits) and an agent subscribed to that channel sees them as they happen; other health checks poll on a configured interval, so their floor is that interval. Syslog catches kernel panics, driver crashes and NVLink errors. The cloud provider's API announces scheduled maintenance and hardware events.
2Quarantine. Cordon the node so nothing new lands.
3Drain. Evict the workloads with the eviction policy the tenants agreed to.
4Remediate. Reset the GPU or reboot the node; if that fails, open the break-fix ticket and replace the node.
5Preflight. Before a new job lands on a returned node, run an active check as an init container so a bad card is never handed to a fresh workload.

NVSentinel is NVIDIA's implementation of exactly that pipeline for Kubernetes: GPU health monitor over DCGM, syslog monitor, cloud-provider monitor, a Kubernetes object monitor that turns any CEL expression into a health event, fault quarantine, node drainer, remediation through node reboot or the cloud API, and preflight. NVIDIA runs it internally with full remediation on by default; the documentation cites deployments across AWS, GCP, Azure and OCI up to 1,100 nodes and around 40,000 GPUs. Status is beta with production use, and it needs the GPU Operator to expose DCGM as a standalone service because it talks to DCGM directly rather than through the exporter. The recommended rollout is the sane one: monitoring only, then cordon, then drain, then remediation, one step at a time.

The managed clouds have their own versions. EKS's node monitoring agent uses the same DCGM push channel and reports a full cycle from fault to a replacement node running workloads in under twelve minutes. AKS runs GPU checks in Node Problem Detector (XID errors, NVLink status, InfiniBand link flapping, ECC) and publishes node conditions; NPD detects and reports only, so remediation is yours to wire up.

The two lines that matter

A rising correctable-ECC rate does not corrupt anything; the hardware fixed every one of those bits. What it can tell you is that the memory may be deteriorating, and a sustained rise justifies closer watching or a proactive quarantine before the first uncorrectable fault lands mid-checkpoint. The trend alert from the observability section is that early warning; the remediation pipeline is what acts on it in time. Pin the control modules to dedicated system or control-plane nodes; NVSentinel provides a system node selector and tolerations for exactly this, and its own demo configuration uses them. A node drainer that evicts itself is a very quiet outage.
A rollout that pulls a 140 GB model from one bucket to fifty nodes at once is a network incident with a model attached.

"Mount, not pull" from 4.4 is the principle. This is the machinery, because a rollout that pulls a 140 GB model from one bucket to fifty nodes at once is a network incident with a model attached.

The time between "start a new replica" and "that replica is answering requests". For an LLM it is dominated by moving the weights into VRAM. Every autoscaling decision in 4.4 is bounded by it, so shaving minutes off cold start is the same thing as being able to scale on real traffic.

The ladder

One ReadWriteMany volume (NFS, CephFS, a cloud file service), one job that populates it per model version, every replica mounts it read-only. This is what NVIDIA's own Dynamo recipes use by default. Add a second volume for the compilation cache (CUDA graphs and the like) so replicas do not recompile on every start. The naive path downloads to local disk and then loads to GPU, two slow copies back to back with the network idle during the second. The Run:ai Model Streamer reads tensors concurrently from object storage and feeds them into GPU memory as they arrive. vLLM wires it in behind --load-format runai_streamer; as of vLLM 0.18 and streamer 0.15.6 it reads directly from S3, GCS and Azure Blob using each cloud's native credential chain, and on AKS a workload identity gets it into Blob with no storage key at all. Once one node has the weights, the others should get them from it, not from the bucket. Dragonfly (CNCF graduated) is the general-purpose P2P layer: it speaks hf:// and modelscope:// natively, is revision-aware, and gets a model onto every node without every node hammering the origin. Serving-framework integrations are on its roadmap, not shipped. For inference fleets running Dynamo, ModelExpress does the same thing at the tensor level: the first worker publishes the weights it holds and later workers pull matching tensors from its GPU memory over NIXL and RDMA, skipping storage entirely. NVIDIA's own guidance is the honest ladder: PVC and a download job for small clusters, ModelExpress when rollout time, autoscale cold start or fleet-wide updates matter more than simplicity. Warm it with a DaemonSet before the rollout so replica start is a local read, not a network fetch. This is the layer that makes scale-out on real traffic possible and is worth doing regardless of which of the above you use for the first copy.

What it changes upstream

A model artefact in the registry (4.4, step 3) now has a distribution plan attached: which volume or cache holds it, which nodes are warm, and how long a cold node takes to become warm. That number goes into the autoscaler's lead time. Without it, the headroom replicas in 4.4 are a guess.

CPU and system memory are cheap and elastic. VRAM is expensive and fixed per device. Some memory models can oversubscribe it into host RAM, but that turns a capacity problem into a severe performance problem, as the Vector tier showed. In 2026 it is supply-constrained upstream as well. Design allocation, quota, scheduling, procurement and cost reporting around it, not around GPU count. The only hardware-isolated VRAM partition is a whole card or a MIG instance. MPS v3 memory limits are a useful software budget; time-slicing and MPS remain for environments where a neighbour's fault is tolerable. Weights, KV per token, headroom and concurrent sequence capacity are numbers that can be computed and asserted in CI. Compute them. Assert them. Production inference, training and batch, CI, and interactive development have different isolation, cost and pre-emption requirements. Give each its own node pool, taint, quota and autoscaling policy. Driver, CUDA runtime, container toolkit and framework versions are a compatibility matrix. Pin all four, upgrade them together, test the upgrade on a canary pool. Allocated VRAM sitting idle is a major source of waste in GPU estates and invisible without per-team memory metrics. CPU-based autoscaling and CPU-based scale-down are both wrong for GPU workloads. Use DCGM and framework metrics for scale-out; protect GPU pods from CPU-based scale-in. With extended-resource support GA in 1.37 and the GPU Operator managing the driver natively, the path exists. Pilot it on a non-production pool this quarter rather than discovering next year that the device plugin is the thing holding your cluster back. HAMi-style interception is a policy fence: it holds for a well-behaved neighbour and fails when the interception path is bypassed. MPS v3 hard memory limits are enforced by the driver, so a runaway allocation fails rather than spreading, but MPS still shares compute, bandwidth and a fault domain. A whole GPU or a MIG instance is the hardware boundary. Put a contract behind hardware; put convenience behind software, and know which kind of software. The time to get weights onto a fresh card bounds how fast the platform can respond to load. Measure it per model and per node class, keep it in the capacity plan next to VRAM per replica, and treat any change that makes it longer as a regression.
Sources for this tier
Roadmap items are vendor statements as reported. Rubin bandwidth, RTX 50 Super timing and memory-shortage recovery dates are the three most likely to move. Check the primary source before you put a date in a plan. Two of the sections above lean on vendor project documentation that is moving monthly (NVSentinel, KAI, Dynamo); version numbers and status labels are the parts most likely to be stale by the time you read this.
Tier 06 — Look It Up

The Reference

Every term the guide uses, and every subject it points at without covering.

Where to go next

Each of these got a box above and each deserves its own climb:

The KV cache and attention Why a long conversation costs more memory than the model itself, and the mechanism that makes it so.
Quantisation and number formats How a 140 GB model becomes a 35 GB one, what you lose, and why it usually gets faster.
The rendering pipeline What the GPU actually does sixty times a second, from triangle to pixel, and where the memory goes.
Advanced packaging: HBM, CoWoS, chiplets Why the memory is stacked on top of the chip now, and why that one step is the bottleneck of the whole AI industry.
Kubernetes for people who run GPUs The parts of Kubernetes you need before any of Zenith makes sense, and none of the parts you do not.
Observability with Prometheus and DCGM Turning a graphics card into numbers on a dashboard, and which of those numbers to wake someone for.
Serving frameworks and the prefill/decode split What vLLM and its rivals do with the memory, and why the two halves of a request want different hardware.
Infrastructure as code for GPU fleets Buying, naming and constraining GPU capacity in files that get reviewed, so nobody picks a shape by accident.
Batch scheduling and fair share (Kueue, Volcano, KAI) Who gets the card when three teams want it, who waits, and how to slice one card in software when the hardware will not.
KV cache tiering and prefill/decode disaggregation Moving the KV cache off the card into host memory, storage and other nodes, and when that is slower than starting again.
Hardware topology and NUMA for accelerators Which two cards you get, not how many, and why the wrong pair runs at PCIe speed.
GPU fleet health and self-healing (NVSentinel, DCGM policy) Detect, quarantine, drain, repair, with no human paged, and how NVIDIA runs it across forty thousand GPUs.
Model artefact distribution at scale Getting 140 GB onto fifty nodes without a network incident, and why cold start is a capacity number.

Quick reference

{{ g.term }} {{ g.meaning }}
Roadmap items are vendor statements as reported. Rubin bandwidth, RTX 50 Super timing and memory-shortage recovery dates are the three most likely to move. Check the primary source before you put a date in a plan.