Video RAM Killed the Radio Star
Sooooo.... what's VRAM?
Nobody sets out to learn about VRAM. You back into it. You are mid-game and the frame rate falls off a cliff, or a model you were promised would run on your card refuses to load, or someone in finance asks why the graphics card line on the bill doubled and you realise you cannot explain it either.
A gamer, an engineer, a platform team and an accountant walk into the same problem. None of them use the same words for it. All of them are talking about a few gigabytes of memory soldered to a circuit board next to a chip, and almost none of them can say what it does.
That memory is the most important number in modern computing that nobody can define at a dinner party. It decides which models exist for you. It decides how fast the answer comes back. It sets the price of the card, the size of the cluster, and, since the memory makers started feeding the AI datacentres first, the price of your console and your laptop too. It is a very small thing that has quietly become everyone's problem.
So this is the whole thing in one climb. It starts with a desk and an artist and ends with a Kubernetes cluster that budgets memory the way a bank budgets cash. Five tiers, and they are a spectrum, not a difficulty setting:
Read to the tier you need and stop. Come back when the next one starts to itch. Nobody is grading you.
Where a section leans on something you might not know yet, there is a short box explaining it in a few sentences and telling you it deserves a study of its own. Those boxes are the map of the guides that follow this one.
Figures are current to mid-September 2026.
The Origin
A computer has two workers. Each needs a desk for whatever they are working on right now. That desk is memory.
Standard RAM is the manager's desk. Built so the manager can grab any single item instantly.
VRAM is the artist's desk. Built so the artist can pour a whole bucket of material across it at once, because painting a picture means touching millions of dots at the same time.
One desk cannot be best at both jobs. So each worker gets a desk shaped for their work, and the artist's desk is bolted directly to the artist's chair so nothing has to travel across the room.
If the artist's desk fills up, one of two things happens. Sometimes the overflow goes on the manager's desk across the room, and every time the artist needs it someone has to walk over and fetch it. That walk is slow and the picture stutters. Other times the artist just stops and says there is no room. Which one you get depends on what kind of work it is.
To get a bigger one, you buy a different artist.
The Vector
VRAM stands for Video Random Access Memory. It is memory that lives physically on the graphics card, wired straight to the GPU, and used only by the GPU.
The name is left over from the 1980s and 1990s, when "VRAM" meant a specific chip design: dual-ported DRAM that let the CPU write one frame into memory while the display hardware read the previous frame out at the same time. That design is long dead. Today "VRAM" just means "whatever memory is bolted to the GPU", regardless of the technology underneath. You will still find consumer explainers repeating the dual-port story as if it describes the GDDR7 on a current card. It does not. Treat it as history.
A CPU runs a handful of threads that jump around unpredictably. It wants any single byte back as fast as possible. DDR system memory is optimised for low latency, measured in nanoseconds.
A GPU runs tens of thousands of threads doing the same operation on adjacent data. It does not much care whether one byte arrives in 100 or 400 nanoseconds. It cares whether it can move hundreds of gigabytes per second. Graphics memory is optimised for bandwidth and placed as physically close to the GPU die as manufacturing allows.
When a workload needs more memory than the GPU can keep resident, what happens depends on the memory model. Graphics drivers page textures and buffers into system RAM and fetch them back over PCIe as needed, and CUDA Unified Memory allocations can migrate between device and host the same way. Ordinary CUDA device allocations do not spill; they fail with an out-of-memory error. When paging does happen it runs over PCIe, which peaks around 32 GB/s on 4.0 x16 and double that on 5.0, a small fraction of VRAM bandwidth. In games this shows up as stutter, texture pop-in and sudden frame-rate drops. In AI workloads it shows up either as throughput collapsing by an order of magnitude or as the process dying outright.
VRAM chips are soldered to the card's circuit board around the GPU. There is no slot. The number on the box is the number you will ever have on that card.
A discrete GPU is a separate card with its own VRAM.
An integrated GPU is built into the CPU package and has no VRAM of its own. It borrows a slice of system RAM. This is called shared graphics memory, or a unified memory architecture. Every laptop without a separate graphics chip works this way, as does every Apple Silicon Mac and most consoles. Unified memory trades peak bandwidth for capacity: a Mac Studio can hold 512 GB of unified memory, which no discrete card at any price can match.
nvidia-smi
The Nexus
This tier is for the person who wants to buy something and understand what they are paying for. It is also where the interesting trade-offs live.
Bandwidth is roughly memory clock multiplied by bits per clock multiplied by bus width. A 256-bit bus with GDDR6 at 16 Gbps effective gets about 512 GB/s. A 512-bit bus with GDDR7 at 28 Gbps gets about 1.8 TB/s. HBM3e on an H200 reaches 4.8 TB/s across six stacks. Bus width is why two cards with the same memory generation can have wildly different bandwidth, and why cheap cards with a narrow bus underperform their headline capacity.
That is the case for error correction.
The consumer versus professional split is less clean than it used to be. GDDR7 on RTX Blackwell cards has single-bit error correction built in and always on, with no toggle and no performance cost, so a GeForce RTX 50 card does correct single-bit flips. What it does not have is the end-to-end RAS stack of a datacentre part: double-bit detection with error containment, dynamic page offlining, row remapping, and driver-level reporting you can alert on. Professional and datacentre GPUs expose all of that, plus validation and support. ECC is part of why they cost several times more, but it is not the main reason.
That is unified memory in one sentence, and it has become the practical answer to the consumer VRAM ceiling. Three platforms matter right now.
Samsung, SK hynix and Micron have moved wafer capacity to HBM for AI accelerators, where margins are far higher. HBM consumes three to four times the wafer area per gigabyte of standard DRAM, and hyperscaler orders are effectively open-ended. The result is DRAM and GDDR tightness that TrendForce expects to run through 2027, with SK hynix warning 2027 could be the worst year of it and some vendors pointing at 2028 before normalisation. NAND has been caught up in the same cycle but is expected to ease sooner, in the second half of 2027.
Concrete effects as of September 2026: memory is now the largest single component cost on a high-end consumer card. The RTX 5090 has traded at roughly double its launch price through the summer. Tom's Hardware reported that NVIDIA moved memory procurement onto board partners in late 2025, leaving them to buy GDDR7 at spot prices. NVIDIA has not confirmed it. Apple pulled the 512 GB Mac Studio option for months before restoring it with the M5 Ultra. Console makers have raised prices citing memory cost. HBM is sold out for 2026.
Two 24 GB cards are not one 48 GB card. Software has to split the work deliberately. NVLink (datacentre cards, and some older consumer cards) makes that split efficient at hundreds of GB/s; over plain PCIe it is much slower. RTX 40 and 50 series consumer cards have no NVLink, so multi-card builds with them are PCIe-bound.
The Apex
This tier is for the engineer who has to make a model fit, run fast, and not fall over. It is also where the arithmetic starts to bite.
The practical rule from Hugging Face's documentation: a model with X billion parameters needs roughly 2X GB of VRAM for its weights at 16-bit precision. A 7B model needs about 14 GB. A 70B model needs about 140 GB, which is why a single 80 GB H100 cannot hold one at full precision and two are the floor.
That is weights alone. Add:
Quantisation converts weights from 16-bit to 8-bit or 4-bit. It cuts VRAM by 50 to 75% and, because inference is bandwidth-bound, usually raises tokens per second too, since fewer bytes get read per token. The cost is some accuracy loss, which varies by model and method. Common formats: GGUF (llama.cpp, Ollama), AWQ and GPTQ (vLLM and other GPU servers), and FP8 natively accelerated on Hopper, Blackwell and Rubin, with native FP4 (NVFP4) starting on Blackwell and continuing on Rubin.
A 128K-token context on Llama 3.1 70B with an FP16 KV cache costs about 40 GB of memory before the model itself. Other 70B designs and lower KV precisions come in far smaller, but the shape of the problem is the same. That is why long conversations eat serving headroom, and why chat systems truncate, compress or refuse once they hit a context limit. It is also the most misunderstood number in LLM sizing.
During generation, an LLM keeps the key and value tensors for every token in the context so it does not recompute them. That store is the KV cache. Size per token, in bytes:
2 × layers × KV heads × head dimension × bytes per element
Autoregressive decode produces one token at a time, and each step reads the entire weight tensor from VRAM. At batch size 1 the compute units mostly sit idle waiting on memory. The ceiling is:
tokens per second ≈ memory bandwidth ÷ bytes of weights read per token
Batching amortises the weight read across many sequences, which is how serving frameworks hit thousands of aggregate tokens per second on one card.
Prefill (processing the prompt) is compute-bound rather than memory-bound. Hold that thought; it matters when scheduling.
MIG is the only one of the three that gives a hardware-isolated, fixed memory partition with its own bandwidth. MPS narrowed the gap in 2026: CUDA 13.4's MPS v3 adds cgroup-integrated device-memory limits, so an allocation that exceeds the hard limit fails with an out-of-memory error rather than starving a neighbour. That is a real budget, but it is enforced in software over a shared card, and a fault in one client can still take down the others.
Here is how you see that.
nvidia-smi # point-in-time nvidia-smi dmon # streaming nvtop # interactive, across cards
The Zenith
This tier is for the architect. Everything above was about one card and one workload. Here VRAM becomes a resource an organisation has to allocate, isolate, meter, autoscale and gate, across many cards, many teams and a pipeline that runs without anyone watching.
Historically, GPU scheduling has relied on a device plugin that tells the kubelet "this node has four of something", and most clusters still work that way. DRA is replacing that model. The core went GA in Kubernetes 1.34 back in 2025; what 1.37 added this year is the piece that makes migration practical, a DRA driver answering the old nvidia.com/gpu request directly. Details further down this section.
How the plugin works
The vendor provides a device plugin (a DaemonSet) that discovers GPUs on each node and advertises them to the kubelet as an extended resource, conventionally nvidia.com/gpu (AMD uses amd.com/gpu). A pod requests an integer count of that resource in its container limits. The scheduler places the pod on a node with enough unallocated units and the device plugin mounts the device files into the container.
Three properties of this model shape everything downstream.
Making memory size visible to the scheduler
Since the resource is opaque, VRAM capacity has to be exposed as node labels so pods can select on it. NVIDIA's GPU Feature Discovery (part of the GPU Operator) labels each node with product name, memory size, compute capability, driver version and MIG configuration. A pod that needs an 80 GB card uses a node selector or affinity on the memory label. Without this, a pod requesting one GPU can land on a 16 GB T4 when it needed an 80 GB H100, and OOM on start.
Giving a pod less than a whole card
Three ways, mapping to the three sharing mechanisms from the Apex tier.
Changing MIG geometry is disruptive. The GPU Operator's MIG Manager handles the reconfiguration and restarts the Operator-managed GPU components, but user workloads have to be evacuated first, and on some driver and platform combinations a reboot is required. There is no fixed duration you can promise. I treat MIG geometry as node-pool infrastructure configuration, dedicate pools to fixed layouts, and do not let workloads change them.
The plugin's replacement: Dynamic Resource Allocation
DRA is the replacement for the device plugin model, and as of September 2026 it is real enough to plan around.
The timeline: core DRA APIs went GA in Kubernetes 1.34 (August 2025) in the resource.k8s.io/v1 group. Kubernetes 1.37, released 26 August 2026, graduated three more pieces to GA: DRA Extended Resource support, device taints and tolerations, and ResourceClaim status. Extended Resource support is the one that matters operationally. A DRA driver can now satisfy a plain nvidia.com/gpu: 1 request in a pod spec with no ResourceClaim on the workload and no device plugin running beside the driver. Tenants keep their manifests; the allocation underneath moves to DRA. That removes the migration blocker that kept most clusters on the device plugin.
The model: workloads describe what they need through ResourceClaims and ResourceClaimTemplates, drivers publish structured device attributes as ResourceSlices, and the scheduler filters and selects on them. Practically, this is "give me any GPU with at least 40 GB and compute capability 9.0" as a first-class scheduling constraint. Device taints let a driver or admin mark degraded hardware so nothing new lands on it.
NVIDIA's side: GPU Operator 26.7.0 manages the DRA driver natively through two new custom resources, NVIDIADriver and GPUCluster, instead of ClusterPolicy. A cluster runs one or the other, not both, and there is no in-place migration between them. Prerequisites are Kubernetes 1.34.2 or later, driver 580 or later, and a CDI-capable container runtime. Workloads can claim full GPUs and preconfigured MIG devices through ResourceClaims and select by attribute. The upstream DRA driver v0.5.0 adds multi-user MPS, NUMA attributes and system-mediated GPU sharing across namespaces via DRA's consumable capacity feature, though several of those remain alpha and off by default.
Related and worth watching: the PodGroup gang-scheduling API reached beta in 1.37, though like other new beta APIs it is not on by default; it sits behind the GenericWorkload feature gate and the scheduling.k8s.io/v1beta1 API has to be enabled. Multi-GPU training jobs that need all their pods or none have been a persistent gap; this is the upstream fix.
Stopping one team from eating the cluster
Growing and shrinking the GPU fleet
Cluster Autoscaler and Karpenter both understand extended resources and will provision a GPU node when a pod is pending on nvidia.com/gpu. Two things bite.
Node start time on GPU instances is long: image pull for the driver container, driver load, device plugin registration, plus a serving pod that may need to pull a 40 GB model. Budget several minutes, set readiness accordingly, or pre-warm.
Scale-down behaviour is autoscaler-specific. Cluster Autoscaler has a separate GPU utilisation threshold for scale-down, based on requested GPU resources on accelerator nodes. Karpenter uses scheduling simulation and consolidation. Neither observes actual VRAM utilisation as a pressure signal, which is why DCGM metrics belong in application autoscaling and capacity policy, not in node scale-down. Annotate long-running GPU pods as non-evictable or use a pod disruption budget.
Karpenter's node pools can express GPU instance families and let the scheduler pick the cheapest instance that satisfies the memory label. Materially cheaper than fixed node groups when workload sizes vary.
The one thing to install first
Why the pipeline needs a GPU at all
The moment a repository contains a model, a CUDA kernel, an inference service, or a data pipeline that runs on GPU, the pipeline needs GPU stages: unit tests that exercise CUDA paths, integration tests that load a model, regression tests on output quality, latency and throughput benchmarks, and image builds that bake in weights or compile kernels. On CPU these either do not work or take so long the pipeline is useless.
Three ways to give it one
What the job needs to actually run
Treating memory like a test
A mature pipeline treats VRAM the way it treats test coverage or bundle size: measured on every run, compared to a threshold, failing the build on regression.
Record these as pipeline artefacts and plot them. VRAM creep is gradual and only visible in trend.
Keeping jobs from poisoning each other
The job of infrastructure-as-code is to make "how much VRAM" a decision someone writes down, not an accident of which shape got picked.
Naming the memory class, not the shape
Define a variable or module per GPU class (for example "inference-80gb", "dev-24gb", "training-141gb-x8"), map each to the shapes that satisfy it per region, and have node pools and VMs reference the class, not the shape. Changing vendor or generation then becomes a mapping change.
Quota is a prerequisite, capacity is a rumour
GPU quota is separate from general compute quota on every major cloud and defaults to zero or close to it. Increases take days and are per region and per shape family. Treat quota as a prerequisite tracked alongside the IaC, and have plan-time checks (or a pre-apply script) fail fast when a plan exceeds it. Capacity is also not guaranteed: a region can have quota available and no physical cards, and in the current memory market that happens more than it used to. For production, use capacity reservations or committed capacity; for CI and batch, use spot with a fallback shape list.
Where the driver comes from
Rules the pipeline enforces so people do not have to
Run these in the pipeline against the plan, before apply.
What you are actually paying for
GPU hours are the most expensive line on the bill and VRAM is what you are actually paying for. Three levers:
Meter VRAM utilisation per team via DCGM and show it back. Allocated-but-idle VRAM is the usual waste pattern and it is invisible without the metric.
What the framework does with the memory
Setting the budget
Serving frameworks expose controls for their GPU and KV cache memory budgets, each in its own way. In current vLLM releases, gpu_memory_utilization is a fraction of total VRAM, defaults to 0.92, and is a per-instance limit. The framework loads weights, then allocates the remainder up to that fraction as KV cache. The arithmetic:
available KV ≈ (VRAM × utilisation fraction)
− weights
− peak activation reserve
− non-PyTorch overhead (NCCL, backend buffers)
− CUDA graph reserve
That is what vLLM's startup profiling actually measures: it loads the weights, runs a dummy forward pass to find the activation peak, measures what was allocated outside PyTorch, estimates CUDA graph memory, and hands whatever is left to the KV cache. The startup log prints each term.
Divide by KV bytes per token to get the card's total token capacity, then divide by expected tokens per request to get concurrent sequence capacity. That number is the node's practical concurrency and context-capacity limit. It is not a speed; bandwidth sets speed, KV capacity sets how many conversations of what length can be resident at once. It should be an explicit input to capacity planning, not something you discover under load.
Set maximum model length explicitly. vLLM sizes the KV pool from the memory budget, not per request, so the setting does not reserve space in advance; what it does is bound the requests the server will accept and determine how much concurrency that pool supports at that context length. Do not advertise 128K if the application only needs 32K.
How the pods should look
Scaling on the right signal
CPU-based HPA is usually the wrong signal for GPU inference; the CPU can sit nearly idle while the GPU or the KV cache is saturated. Tokenisation and preprocessing can make CPU matter in some designs, but it is rarely the limit. Scale on a metric that reflects VRAM and GPU pressure:
Scale-out lead time is node provisioning plus model load, typically five to fifteen minutes. Either keep headroom replicas or use predictive scaling on traffic patterns.
Shipping a model like software
A model version is a deployable artefact with a VRAM footprint, the same way a container image has a size. The promotion pipeline should:
Weights should be mounted, not pulled, wherever possible: a read-only persistent volume populated once per version, or a node-local cache warmed by a DaemonSet before the rollout, so replica start is seconds not minutes.
What to watch and when to wake someone
Minimum metric set per GPU, scraped from DCGM:
Per serving replica, from the framework: KV cache utilisation, running and queued requests, time to first token, tokens per second, request failures by reason (OOM versus timeout versus error).
The device plugin answers "is there a free GPU on this node". It does not answer "should this team get it before that one", "does this job need all eight or none", or "can this notebook have a quarter of a card". Those three questions are where most real clusters live, and the answers come from the scheduling and admission layer around Kubernetes: sometimes in front of the default scheduler, sometimes replacing it.
Three tools, three layers
All three coexist with the device plugin or with DRA. None of them replaces the GPU Operator.
Slicing a card in software
Apex covered time-slicing, MPS and MIG. There is another common software approach, and it is a library rather than a driver feature.
HAMi (a CNCF project) intercepts CUDA calls inside the container through a preloaded library and enforces a memory ceiling and a compute share that the pod requested in megabytes and percent. A pod asks for one vGPU with nvidia.com/gpu: 1, nvidia.com/gpumem: 8000 (MiB) and nvidia.com/gpucores: 30, sees 8 GB inside the container, and gets an out-of-memory error when it tries to take more. No MIG-capable hardware needed, no profile table, no reset to change the split, and it works on a consumer card. HAMi plugs into Kueue as a schedulable flavour, into Volcano through the vGPU plugin, and into KAI's resource isolator. A published test on an 8x RTX PRO 6000 Blackwell node under Kubernetes 1.35 turned eight cards into eighty schedulable slices and showed the OOM firing at the quota.
The decision table an architect actually needs:
A 128K-token context on Llama 3.1 70B at FP16 KV is about 40 GB. In 2026 the answer to "where does that live" is no longer "in VRAM or nowhere".
Why one card is the wrong unit
vLLM's built-in prefix cache is per replica. Ten replicas behind a load balancer each keep their own, so a request that lands on replica B gets nothing from what replica A computed a minute ago for the same 100K-token document. Agentic workloads make this worse: a coding agent runs a long multi-turn loop, each turn re-sends the whole history, and the working context outgrows any single card's KV budget. vLLM's own analysis of Codex-style traces on SWE-bench is what drove its Mooncake integration.
What replaces it
The stack that makes this operational
Prefill/decode disaggregation (Apex, "hold that thought") plus a shared KV pool is now an orchestrated pattern rather than a research setup. NVIDIA's Dynamo is the inference runtime that does the routing, Grove is the Kubernetes API that wraps prefill, decode and router pods into one gang-scheduled unit with startup ordering, and KAI places the gang on hardware that can talk to itself fast. The point for VRAM: prefill pools want compute, decode pools want bandwidth, and the KV cache moves between them over the fabric. Buying identical cards for both is the mistake this stack exists to stop.
Tensor parallelism across two GPUs on the same NVLink pair runs at NVLink speed. The same job on two GPUs that only share a PCIe root complex runs at PCIe speed, and across two that share nothing but the CPU it is worse. Nothing in nvidia.com/gpu: 2 says which pair you get.
How the scheduler learns the map
The traditional device plugin knows a little. It can publish NUMA locality and preferred allocations for the node it runs on, and the kubelet's Topology Manager can use that to keep a pod's GPUs and CPUs together inside one node. What it cannot do is make a rack-scale NVLink domain, or a GPU-plus-NIC relationship, a first-class object the cluster scheduler reasons about. Placement of multi-GPU jobs across nodes under the device plugin is affinity heuristics, not scheduler logic.
DRA fixes this at the description layer: the NVIDIA DRA driver publishes GPU attributes including NVLink topology as structured parameters the scheduler can filter on. For multi-node NVLink (GB200 and GB300 racks), the NVIDIA DRA driver manages ComputeDomains, its abstraction for "this set of nodes shares one NVLink domain, keep the whole job inside it", and labels each node with nvidia.com/gpu.clique to say which partition it belongs to.
A scheduler still has to consume that. KAI's topology-aware scheduling reads a cluster-scoped Topology object whose levels run from widest to narrowest and can include the GPU clique and the hostname; it places a gang inside the requested level. NVIDIA's own NVCF stack creates exactly that mapping from the clique labels for its topology-aware path. Kueue has its own topology-aware scheduling built on a labelled node hierarchy. Volcano has a topology policy on its PodGroup. HAMi's scheduler avoids a poorly connected GPU for multi-GPU requests and offers it to single-GPU ones.
What to actually do
The Zenith observability section says to alert on double-bit ECC, rising single-bit ECC, XID faults and thermal throttling. This section is what the platform does when those alerts fire, because a page to a human at three in the morning is the expensive answer.
Detect, quarantine, drain, fix
The shape every mature platform converges on, and the shape NVIDIA shipped as open source:
NVSentinel is NVIDIA's implementation of exactly that pipeline for Kubernetes: GPU health monitor over DCGM, syslog monitor, cloud-provider monitor, a Kubernetes object monitor that turns any CEL expression into a health event, fault quarantine, node drainer, remediation through node reboot or the cloud API, and preflight. NVIDIA runs it internally with full remediation on by default; the documentation cites deployments across AWS, GCP, Azure and OCI up to 1,100 nodes and around 40,000 GPUs. Status is beta with production use, and it needs the GPU Operator to expose DCGM as a standalone service because it talks to DCGM directly rather than through the exporter. The recommended rollout is the sane one: monitoring only, then cordon, then drain, then remediation, one step at a time.
The managed clouds have their own versions. EKS's node monitoring agent uses the same DCGM push channel and reports a full cycle from fault to a replacement node running workloads in under twelve minutes. AKS runs GPU checks in Node Problem Detector (XID errors, NVLink status, InfiniBand link flapping, ECC) and publishes node conditions; NPD detects and reports only, so remediation is yours to wire up.
The two lines that matter
"Mount, not pull" from 4.4 is the principle. This is the machinery, because a rollout that pulls a 140 GB model from one bucket to fifty nodes at once is a network incident with a model attached.
The ladder
What it changes upstream
A model artefact in the registry (4.4, step 3) now has a distribution plan attached: which volume or cache holds it, which nodes are warm, and how long a cold node takes to become warm. That number goes into the autoscaler's lead time. Without it, the headroom replicas in 4.4 are a guess.
The Reference
Every term the guide uses, and every subject it points at without covering.
Where to go next
Each of these got a box above and each deserves its own climb: