[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$fux95wn0jkcnn":3},{"_id":4,"slug":5,"title":6,"subtitle":7,"kind":8,"cards":9,"tags":58,"categories":60,"source":62,"lang":65,"author":66,"audioState":69,"stats":70,"publishedAt":73,"renderer":74},"6abcc759ca21c797c7ea118a","vram-for-local-llms-why-memory-bandwidth-sets-your-tokens-pe-87d3cf12","VRAM for local LLMs: why memory bandwidth sets your tokens per second","How much VRAM for an LLM is the wrong first question.","news",[10,13,18,23,28,33,38,43,48,53],{"headline":6,"body":11,"imageUrl":12,"sourceImageUrl":12},"How much VRAM for an LLM is the wrong first question. The better one is how fast that VRAM is, because a local model generating text reads its entire set of weights from memory for every single token. That makes memory bandwidth, in gigabytes per second, the number that decides whether your coding agent types or crawls. This is the bandwidth-first companion to my Local AI Hardware Guide (2026) from March: fewer shopping lists, more of the arithmetic behind them.","https:\u002F\u002Fmedia2.dev.to\u002Fdynamic\u002Fimage\u002Fwidth=1200,height=627,fit=cover,gravity=auto,format=auto\u002Fhttps%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F59gvru2exeetw6larphd.jpg",{"headline":14,"body":15,"imageUrl":16,"images":17},"Token generation at batch size 1 is memory-bandwidth","Token generation at batch size 1 is memory-bandwidth bound. A rough ceiling is bandwidth ÷ model size in memory: a 10 GB model on a 936 GB\u002Fs RTX 3090 tops out around 90 tokens a second.","\u002Fapi\u002Fmedia\u002Fposts\u002Fvram-for-local-llms-why-memory-bandwidth-sets-your-tokens-pe-87d3cf12\u002F1.webp",{"local":16},{"headline":19,"body":20,"imageUrl":21,"images":22},"When the model or its context does not","When the model or its context does not fit, layers spill to DDR5 (about 50 GB\u002Fs) over PCIe 4.0 (31.5 GB\u002Fs): a 20x bandwidth drop, and 42.5 tok\u002Fs becomes 3.8 in the benchmark below.","\u002Fapi\u002Fmedia\u002Fposts\u002Fvram-for-local-llms-why-memory-bandwidth-sets-your-tokens-pe-87d3cf12\u002F2.webp",{"local":21},{"headline":24,"body":25,"imageUrl":26,"images":27},"At 4-bit, weights cost about 5 GB for","At 4-bit, weights cost about 5 GB for 8B, 10 GB for 14B, 20 GB for 32B and 40 GB for 70B, before the KV cache. A 32k-token agent context adds several more gigabytes. The best VRAM per dollar is still a used RTX 3090 24GB, around $650–750, because it pairs 24 GB with a 384-bit bus. Why memory bandwidth, not VRAM size alone, sets LLM speed","\u002Fapi\u002Fmedia\u002Fposts\u002Fvram-for-local-llms-why-memory-bandwidth-sets-your-tokens-pe-87d3cf12\u002F3.webp",{"local":26},{"headline":29,"body":30,"imageUrl":31,"images":32},"A game is compute-bound: the CPU prepares draw","A game is compute-bound: the CPU prepares draw calls, the shaders fill pixels, a faster core means more frames. Autoregressive decoding, the way an LLM writes one token at a time, is the reverse. To compute the next token the GPU streams every weight of the model from VRAM through its compute units, does a little maths on each, and starts again. The cores mostly wait on the memory bus. The back-of-envelope formula:","\u002Fapi\u002Fmedia\u002Fposts\u002Fvram-for-local-llms-why-memory-bandwidth-sets-your-tokens-pe-87d3cf12\u002F4.webp",{"local":31},{"headline":34,"body":35,"imageUrl":36,"images":37},"Real runs land below the ceiling (roughly 60–75","Real runs land below the ceiling (roughly 60–75 tok\u002Fs on the 3090, 1.5–2 on DDR5), but the ratio holds: same model, same machine, twenty times slower because the weights moved. That is why a 5.8 GHz Core i9 and a liquid loop do nothing here. The CPU only feeds the GPU; a six-core Ryzen 5 is plenty. GPU memory bandwidth by tier","\u002Fapi\u002Fmedia\u002Fposts\u002Fvram-for-local-llms-why-memory-bandwidth-sets-your-tokens-pe-87d3cf12\u002F5.webp",{"local":36},{"headline":39,"body":40,"imageUrl":41,"images":42},"Bandwidth is set by the memory type and","Bandwidth is set by the memory type and the bus width; the marketing tier says nothing about it. The same 16 GB can be fast or slow: Figures are from Nvidia's spec pages for the RTX 3090 and RTX 4090, plus the PCIe 4.0 and JEDEC DDR5 standards.","\u002Fapi\u002Fmedia\u002Fposts\u002Fvram-for-local-llms-why-memory-bandwidth-sets-your-tokens-pe-87d3cf12\u002F6.webp",{"local":41},{"headline":44,"body":45,"imageUrl":46,"images":47},"The 4060 Ti row explains itself: 288 GB\u002Fs","The 4060 Ti row explains itself: 288 GB\u002Fs over a 5 GB model is a ~58 tok\u002Fs ceiling, and it measures 42.5 below. Fine for a starter card, out of its depth at 32B. The 20x cliff: what happens when a model does not fit in VRAM","\u002Fapi\u002Fmedia\u002Fposts\u002Fvram-for-local-llms-why-memory-bandwidth-sets-your-tokens-pe-87d3cf12\u002F7.webp",{"local":46},{"headline":49,"body":50,"imageUrl":51,"images":52},"Runtimes like llama.cpp, Ollama and vLLM don't refuse","Runtimes like llama.cpp, Ollama and vLLM don't refuse a model that is too big. They split it: some layers in VRAM, the rest in system RAM across the PCIe bus. The GPU finishes its layers in microseconds, then stalls. The benchmark from the episode, Llama 8B on an RTX 4060 Ti 16GB:","\u002Fapi\u002Fmedia\u002Fposts\u002Fvram-for-local-llms-why-memory-bandwidth-sets-your-tokens-pe-87d3cf12\u002F8.webp",{"local":51},{"headline":54,"body":55,"imageUrl":56,"images":57},"Remember the middle row. Moving a fifth of","Remember the middle row. Moving a fifth of the model off the card cost 91 % of the speed, because every token waits for the slowest pipe. \"It almost fits\" is not a thing: either the model and its context fit in VRAM, or you run at DDR5 speed with extra steps. How much VRAM do you need for a 32B model? Weights plus KV cache Weights are the entry fee. At 4-bit quantization (GGUF Q4_K_M or AWQ):","\u002Fapi\u002Fmedia\u002Fposts\u002Fvram-for-local-llms-why-memory-bandwidth-sets-your-tokens-pe-87d3cf12\u002F9.webp",{"local":56},[59],"dev",[61],"Technology",{"name":63,"url":64},"Dev.to","https:\u002F\u002Fdev.to\u002Faxrisi\u002Fvram-for-local-llms-why-memory-bandwidth-sets-your-tokens-per-second-h4h","en",{"handle":67,"displayName":68},"spots","Spots","queued",{"views":71,"likes":72,"saves":72,"shares":72,"completions":72,"opens":72,"skips":72,"depthSum":72},1,0,"2026-09-30T08:24:57.740Z","local"]