GPUCalcPROVRAM & Cloud Cost Lab
VRAM Sizer CalculatorLive GPU MatrixFormulas & MathAbout GPUCalcContact & FeedbackPrivacy Policy
Live Multi-Cloud Pricing Matrix • Updated for RTX 5090, Blackwell B200 & DeepSeek-R1

LLM VRAM Calculator & Cloud GPU Cost Estimator

Accurately calculate GPU VRAM requirements for Inference, QLoRA, LoRA, and Full Fine-Tuning. Compare real-time spot and on-demand GPU rental prices across RunPod, Lambda Labs, Vast.ai, AWS, and GCP.

Model Architectures
15+ Preloaded
Context Window Scaling
Up to 128k Tokens
Cloud Providers Indexed
8 Major Clouds
Quantization Formats
FP32 • FP16 • FP8 • INT4
Interactive VRAM Sizer & Live Cloud Pricing Matrix Active

Popular Model Presets

Auto-Architected
B
0.5B (Edge)7B14B32B70B140B+
0.55 B/param
tokens
reqs
Minimum Required VRAMActive
40.66GB
Recommended Target:46.76 GB(+15% buffer)
Suggested Minimum GPU Tier:
1x L40S / RTX 6000 Ada (48GB)

VRAM Memory Footprint Breakdown

Sum: 40.66 GB
Model Weights
36.16 GB
KV-Cache Context
2.5 GB
Activations
0.5 GB
CUDA & Runtime
1.5 GB
Estimated Decoding Speed (Tokens/sec):
A100 (80GB): 36.5 t/sH100 (80GB): 60 t/s

Live Cloud GPU Cost & Pricing Engine

23 Available Nodes

Real-time verified pricing across RunPod, Lambda Labs, Vast.ai, AWS, GCP, and specialized clouds.

Lowest Cost Compatible OptionNVIDIA L40S on RunPod (48GB VRAM)
Spot Rate$0.65 / 1h
Rent Node
Buy vs. Rent TCO LabHardware break-even simulator
Providers:
GPU & ArchitectureProvider
Total VRAM
Spot (1h)
On-Demand (1h)
Action
NVIDIA L40SPopular
Ada LovelacePCIe 4.0
RunPod
48 GB
$0.65($0.65/hr)$0.94($0.94/hr)Rent
NVIDIA L40S
Ada LovelacePCIe 4.0
Lambda Labs
48 GB
$0.69($0.69/hr)$0.99($0.99/hr)Rent
NVIDIA RTX 6000 Ada
Ada LovelacePCIe 4.0
RunPod
48 GB
$0.85($0.85/hr)$1.25($1.25/hr)Rent
2x NVIDIA RTX 4090
Ada LovelacePCIe 4.0
RunPod
48 GB(2x 24GB)
$0.88($0.88/hr)$1.38($1.38/hr)Rent
NVIDIA A100 SXM4 (80GB)Popular
AmpereNVLink-3 (600 GB/s)
Lambda Labs
80 GB
$1.25($1.25/hr)$1.69($1.69/hr)Rent
2x NVIDIA RTX 5090
BlackwellPCIe 5.0
RunPod
64 GB(2x 32GB)
$1.38($1.38/hr)$1.98($1.98/hr)Rent
NVIDIA A100 SXM4 (80GB)
AmpereNVLink-3 (600 GB/s)
RunPod
80 GB
$1.39($1.39/hr)$1.89($1.89/hr)Rent
NVIDIA A100 SXM4 (80GB)
AmpereNVLink-3 (600 GB/s)
GCP (a2-ultragpu-1g)
80 GB
$1.75($1.75/hr)$3.67($3.67/hr)Rent
4x NVIDIA RTX 4090
Ada LovelacePCIe 4.0
RunPod
96 GB(4x 24GB)
$1.76($1.76/hr)$2.76($2.76/hr)Rent
NVIDIA H100 SXM5 (80GB)
HopperNVLink-4 (900 GB/s) + InfiniBand
Nebius AI Studio
80 GB
$2.19($2.19/hr)$2.85($2.85/hr)Rent
NVIDIA H100 SXM5 (80GB)Popular
HopperNVLink-4 (900 GB/s)
Lambda Labs
80 GB
$2.29($2.29/hr)$2.99($2.99/hr)Rent
NVIDIA H100 SXM5 (80GB)
HopperNVLink-4 (900 GB/s)
CoreWeave
80 GB
$2.35($2.35/hr)$3.15($3.15/hr)Rent
NVIDIA H100 SXM5 (80GB)Popular
HopperNVLink-4 (900 GB/s)
RunPod
80 GB
$2.49($2.49/hr)$3.29($3.29/hr)Rent
4x NVIDIA RTX 5090
BlackwellPCIe 5.0
RunPod
128 GB(4x 32GB)
$2.76($2.76/hr)$3.96($3.96/hr)Rent
NVIDIA H200 SXM5 (141GB)Popular
Hopper HBM3eNVLink-4 (900 GB/s)
Lambda Labs
141 GB
$3.49($3.49/hr)$4.29($4.29/hr)Rent
8x NVIDIA A100 SXM4 (80GB)
AmpereNVLink-3 (600 GB/s)
Lambda Labs
640 GB(8x 80GB)
$9.90($9.90/hr)$13.52($13.52/hr)Rent
8x NVIDIA A100 SXM4 (80GB)
AmpereNVLink-3 (600 GB/s)
AWS (EC2 p4de.24xlarge)
640 GB(8x 80GB)
$14.50($14.50/hr)$40.97($40.97/hr)Rent
8x NVIDIA H100 SXM5 (80GB)Popular
HopperNVLink-4 (900 GB/s) + Quantum-2 IB
Lambda Labs
640 GB(8x 80GB)
$18.30($18.30/hr)$23.92($23.92/hr)Rent
8x NVIDIA H100 SXM5 (80GB)
HopperNVLink-4 (900 GB/s) + 3.2Tbps InfiniBand
RunPod
640 GB(8x 80GB)
$19.80($19.80/hr)$26.32($26.32/hr)Rent
8x NVIDIA H200 SXM5 (141GB)
Hopper HBM3eNVLink-4 (900 GB/s) + Quantum-2 IB
Lambda Labs
1128 GB(8x 141GB)
$27.50($27.50/hr)$34.32($34.32/hr)Rent
8x NVIDIA H100 SXM5 (80GB)
HopperNVLink-4 (900 GB/s) + GPUDirect-TCPX
GCP (a3-highgpu-8g)
640 GB(8x 80GB)
$32.00($32.00/hr)$87.05($87.05/hr)Rent
8x NVIDIA H100 SXM5 (80GB)
HopperNVLink-4 (900 GB/s) + 3.2Tbps EFA
AWS (EC2 p5.48xlarge)
640 GB(8x 80GB)
$38.50($38.50/hr)$98.32($98.32/hr)Rent
8x NVIDIA Blackwell B200 (192GB)Popular
BlackwellNVLink-5 (1.8 TB/s) + Quantum-X800 IB
Nebius / CoreWeave
1536 GB(8x 192GB)
$44.00($44.00/hr)$58.00($58.00/hr)Rent
Showing all 23 nodes • Scroll table vertically to view full hardware catalog
Prices updated February 2026
Technical Knowledge Hub

Frequently Asked Questions & Transformer Math

Deep-dive technical answers on KV-cache calculation, CUDA Out-Of-Memory prevention, QLoRA fine-tuning benchmarks, and cloud GPU cost optimization.

Programmatic Model Sizing Hub

Popular LLM VRAM Requirements Guides

Explore dedicated memory calculators, architecture specifications, and GPU sizing guides for popular open-weights models.

24 Dedicated Model Hubs
70.6B ParamsMeta Llama

Llama 3.3 70B Instruct

Meta's state-of-the-art dense 70B open-weights powerhouse with 128k context and 8:1 Grouped Query Attention (GQA).

Calculate VRAM
8.03B ParamsMeta Llama

Llama 3.1 8B Instruct

The industry standard compact LLM for local reasoning, autonomous agents, and fast fine-tuning pipelines.

Calculate VRAM
671B ParamsDeepSeek

DeepSeek R1 (671B MoE)

Frontier open reasoning model featuring 671B total parameters, 37B active per token, and Multi-Head Latent Attention (MLA).

Calculate VRAM
671B ParamsDeepSeek

DeepSeek V3 (671B MoE)

Flagship 671B mixture-of-experts generalist model with state-of-the-art coding and multilingual performance.

Calculate VRAM
70.6B ParamsDeepSeek

DeepSeek R1 Distill 70B (Llama)

DeepSeek reasoning tokens distilled into the Llama 3.3 70B transformer architecture.

Calculate VRAM
32.5B ParamsDeepSeek

DeepSeek R1 Distill 32B (Qwen)

Distilled reasoning traces inside Qwen 2.5 32B architecture, matching larger 70B frontier benchmarks.

Calculate VRAM
14.7B ParamsDeepSeek

DeepSeek R1 Distill 14B (Qwen)

High-efficiency coding and reasoning model tailored for single-GPU local workstations.

Calculate VRAM
72.7B ParamsAlibaba Qwen

Qwen 2.5 72B Instruct

Top-tier multilingual, math, and code generation dense LLM with 128k sequence capability.

Calculate VRAM
32.5B ParamsAlibaba Qwen

Qwen 2.5 32B Instruct

The golden sweet-spot model delivering 70B-grade capabilities with a 32B memory footprint.

Calculate VRAM
14.7B ParamsAlibaba Qwen

Qwen 2.5 14B Instruct

Ultra-fast mid-size model excelling in function calling and agentic code execution.

Calculate VRAM
7.61B ParamsAlibaba Qwen

Qwen 2.5 7B Instruct

High-throughput lightweight model optimized for edge deployments, real-time chatbots, and embeddings.

Calculate VRAM
123B ParamsMistral AI

Mistral Large 2 (123B)

Mistral AI's flagship 123B dense model featuring 128k context, strong reasoning, and native multi-lingual support.

Calculate VRAM
24B ParamsMistral AI

Mistral Small 3 24B

Enterprise-ready 24B reasoning model engineered for fast inference and lightweight fine-tuning.

Calculate VRAM
46.7B ParamsMistral AI

Mixtral 8x7B MoE

Pioneering sparse Mixture of Experts model with 47B total parameters and 13B active per token.

Calculate VRAM
27.2B ParamsGoogle Gemma

Gemma 2 27B

Google DeepMind's highly efficient architecture combining sliding-window and global attention.

Calculate VRAM
9.24B ParamsGoogle Gemma

Gemma 2 9B

Compact powerhouse model with 256 head dimension delivering top-tier performance for its weight class.

Calculate VRAM
14.7B ParamsMicrosoft Phi

Phi-4 (14B)

Microsoft Research's state-of-the-art synthetic data trained model with exceptional reasoning.

Calculate VRAM
12B ParamsBlack Forest Labs

FLUX.1 [schnell] (12B Diffusion)

12B Rectified Flow Matching Transformer for ultra-fast, high-detail text-to-image synthesis.

Calculate VRAM
3.5B ParamsStability AI

Stable Diffusion XL 1.0 (3.5B)

Flagship 3.5B latent diffusion model with dual text encoders for photorealistic image generation.

Calculate VRAM
11B ParamsMeta Llama

Llama 3.2 11B Vision Instruct

Meta's flagship multimodal model pairing an 8B text LLM with a 3B vision encoder for high-resolution image reasoning.

Calculate VRAM
90B ParamsMeta Llama

Llama 3.2 90B Vision Instruct

Frontier 90B multimodal model with advanced visual document comprehension, chart parsing, and high-res image reasoning.

Calculate VRAM
72.7B ParamsQwen

Qwen 2.5-VL 72B Instruct

Alibaba's frontier vision-language model with dynamic resolution processing and native video comprehension.

Calculate VRAM
7.6B ParamsQwen

Qwen 2.5-VL 7B Instruct

High-efficiency open multimodal model capable of analyzing long documents, charts, and video streams at rapid speed.

Calculate VRAM
12B ParamsMistral AI

Mistral Pixtral 12B Vision

Mistral's open-source 12B multimodal powerhouse featuring a 400M vision encoder capable of arbitrary image aspect ratios.

Calculate VRAM

Popular Sizing & Hardware Recipes

Standard battle-tested deployment patterns for the most common open-source LLM workloads.

Budget Friendly$0.44/hr

Llama 3.3 70B at Home

Running 70B on 1x RTX 4090 (24GB) using INT4 / AWQ quantization with 8k context window.

VRAM: ~20.5 GB1x RTX 4090
Sweet Spot$0.69/hr

Qwen 2.5 32B FP8

Zero degradation 32B model on 1x RTX 5090 (32GB) or 1x L40S (48GB) with 32k context.

VRAM: ~31.2 GB1x RTX 5090
Training$0.22/hr

8B QLoRA Fine-Tuning

Full parameter adapter tuning for Llama 3.1 8B with rank=16 and batch size 4 on 1x RTX 3090.

VRAM: ~14.8 GB1x RTX 3090
Frontier Reasoning$27.50/hr

DeepSeek-R1 (671B)

Running 671B full MoE in FP8 on an 8x H200 (141GB) or 4x B200 (192GB) supercluster.

VRAM: ~720 GB8x H200 SXM
Architectural Precision

Mathematical Methodology & Formulas

GPU memory allocation in modern PyTorch, vLLM, and TensorRT-LLM is governed by deterministic equations across weights, KV-cache, activations, and optimizer states.

1. Model Weights Memory

M_weights = (Parameters × Bytes_per_Param) / 1024³

For FP16/BF16, each parameter requires 2 bytes. For INT4/AWQ, parameters take 0.55 bytes (including metadata, scaling factors, and zero-point overhead).

2. KV-Cache Context Memory

M_kv = 2 × n_layers × n_kv_heads × d_head × C × B × b_kv

Modern architectures (Llama 3, Mistral, Qwen) use Grouped Query Attention (GQA), reducing the number of KV heads by 4x to 8x compared to query heads, drastically lowering context expansion costs.

3. Training & AdamW Optimizer States

M_opt = Parameters_trainable × 12 Bytes

Full fine-tuning requires 4 bytes for FP32 master weights, 4 bytes for first momentum m, and 4 bytes for second momentum v, totaling 12 bytes/param plus 2 bytes for gradients.

4. CUDA Runtime & PyTorch Fragmentation Buffer

M_total = (M_weights + M_kv + M_opt + M_act + 1.5GB) × 1.15

PyTorch and CUDA runtime kernels consume ~1.5 GB of context overhead. An additional 15% headroom buffer prevents CUDA Out-Of-Memory (OOM) errors caused by memory fragmentation during dynamic batching.

Explore More Resources & Support

Learn about our methodology, get in touch for custom benchmarks, or review our privacy policies.