What VRAM is required to execute LLMs locally?
When deploying a local LLM, the primary consideration is straightforward: does it fit within your GPU's capabilities? The answer hinges on model size, quantization levels, and context length. This guide provides a practical baseline for determining the optimal amount of VRAM for your needs.
Impact of quantization on VRAM
Quantization lowers the precision used for storing model weights. Reducing the bit depth shrinks the model size and decreases VRAM consumption, albeit with a minor compromise in quality.
| Quant | Bits per weight | Typical use |
|---|---|---|
| Q8_0 | 8 | Very high quality |
| Q6_K | ~6.6 | Very good quality |
| Q5_K_M | ~5.5 | Good quality and size |
| Q4_K_M | ~4.5 | Good balance of size and quality |
| Q3_K_M | ~3.5 | Lower VRAM, more quality loss |
Q4_K_M is a standard choice when VRAM is constrained. With greater VRAM availability, Q5 or Q6 allows you to run the same model with higher precision and less quantization.
Approximate VRAM by model size
The following are rough estimates for model weights. Actual VRAM requirements are higher, as the runtime, KV cache, and context also consume memory.
| Model size | Q8_0 | Q6_K | Q4_K_M | Q3_K_M |
|---|---|---|---|---|
| 4B | ~5 GB | ~4 GB | ~3 GB | ~2.5 GB |
| 8B | ~9 GB | ~7 GB | ~5.5 GB | ~4.5 GB |
| 12B | ~13 GB | ~10 GB | ~8 GB | ~6.5 GB |
| 14B | ~16 GB | ~12 GB | ~9 GB | ~7.5 GB |
| 27B | ~30 GB | ~22 GB | ~17 GB | ~13 GB |
| 32B | ~36 GB | ~27 GB | ~20 GB | ~16 GB |
| 70B | ~80 GB | ~60 GB | ~42 GB | ~34 GB |
These figures are estimates rather than strict limits. Variations in model architecture and quantization formats can alter the actual size requirements.
Capabilities of different VRAM tiers
| VRAM | Practical range | Current examples |
|---|---|---|
| 8 GB | Small models around 4B to 9B | Gemma 4 E4B, Qwen3.5 9B |
| 12 GB | Small to mid-sized models around 9B to 14B | Gemma 4 12B, Qwen3.5 9B |
| 16 GB | 12B to 27B with lower quantization | Gemma 4 26B-A4B, Qwen3.6 27B at Q4 |
| 24 GB | 27B to 35B at Q4 to Q6 | Qwen3.8 27B, Gemma 4 31B |
| 32 GB | 27B to 35B at higher quantization | Qwen3.8 27B, Gemma 4 31B |
| 48 GB | Large dense models at lower quantization | 70B-class models at Q3 to Q4 |
| 80 GB | Large dense models at higher quantization | 70B-class models at Q4 to Q6 |
These ranges apply to models whose weights can be loaded onto the GPU. Large MoE models differ: although only a subset of parameters is active per token, the model must still store its complete set of weights. Consequently, a model with 100B or more total parameters will not fit within a 100B-sized VRAM budget simply because it has fewer active parameters.
MoE models
Mixture-of-Experts models consist of multiple parameter groups known as experts. Since only specific experts are engaged for each token, inference can be more efficient compared to a dense model with the same total parameter count.
However, inactive experts remain part of the model structure. Therefore, large MoE models can demand significantly more memory than their active parameter count implies. Very large models may necessitate multiple GPUs or offloading to system RAM.
Context length also consumes VRAM
Model weights represent only a portion of the total memory requirement. As context length increases, the KV cache expands, meaning running the same model at 64K context can demand substantially more VRAM than at 4K.
- Longer context requires more VRAM.
- KV-cache precision affects memory usage.
- Batch size and concurrent users also increase memory usage.
- Reserve some VRAM for the runtime instead of saturating the GPU with model weights.
Practical tips
- Verify the actual size of the quantized model you intend to run.
- Do not equate the model file size with the exact VRAM requirement. Allow space for the KV cache and runtime.
- If a model does not fit entirely in VRAM, part of it can be offloaded to system RAM, though inference will typically be slower.
- For long-context or agentic workloads, allocate more VRAM than the model weights alone require.
- Multiple GPUs can distribute the model if a single GPU lacks sufficient VRAM.
Run it on DaDesktop
There is no need to purchase a GPU to run a local LLM. DaDesktop provides a cloud desktop equipped with the necessary VRAM, allowing you to execute the model directly without owning the hardware.
Select the VRAM tier that suits your model, load it, and begin usage. No setup, no hardware purchase, and no driver issues. Visit available GPUs to explore your options.