
VRAM for local LLMs: why memory bandwidth sets your tokens per second
How much VRAM for an LLM is the wrong first question. The better one is how fast that VRAM is, because a local model generating text reads its entire set of weights from memory for every single token. That makes memory bandwidth, in gigabytes per second, the number that decides whether your coding agent types or crawls. This is the bandwidth-first companion to my Local AI Hardware Guide (2026) from March: fewer shopping lists, more of the arithmetic behind them.

