D
DevCalc
All tools
AI calculators
Guides
About
Explore tools
AI calculators
Infrastructure planning
LLM VRAM Calculator
Estimate GPU memory for LLM inference.
Live estimate
Model
Llama 3.1 8B
Qwen2.5 14B
Gemma 2 27B
Custom model
Precision
FP16 / BF16
INT8
INT4
Context length
8,192 tokens
Concurrent sequences
Runtime overhead
12%
Save configuration
Share result
Reset
Architecture presets reviewed Sep 3, 2026
Official model card ↗