LLM VRAM Calculator & Cloud GPU Cost Estimator
Accurately calculate GPU VRAM requirements for Inference, QLoRA, LoRA, and Full Fine-Tuning. Compare real-time spot and on-demand GPU rental prices across RunPod, Lambda Labs, Vast.ai, AWS, and GCP.
Popular Model Presets
Auto-ArchitectedVRAM Memory Footprint Breakdown
Sum: 40.66 GBLive Cloud GPU Cost & Pricing Engine
23 Available NodesReal-time verified pricing across RunPod, Lambda Labs, Vast.ai, AWS, GCP, and specialized clouds.
| GPU & Architecture | Provider | Total VRAM | Spot (1h) | On-Demand (1h) | Action |
|---|---|---|---|---|---|
NVIDIA L40SPopular Ada Lovelace•PCIe 4.0 | RunPod | 48 GB | $0.65($0.65/hr) | $0.94($0.94/hr) | Rent |
NVIDIA L40S Ada Lovelace•PCIe 4.0 | Lambda Labs | 48 GB | $0.69($0.69/hr) | $0.99($0.99/hr) | Rent |
NVIDIA RTX 6000 Ada Ada Lovelace•PCIe 4.0 | RunPod | 48 GB | $0.85($0.85/hr) | $1.25($1.25/hr) | Rent |
2x NVIDIA RTX 4090 Ada Lovelace•PCIe 4.0 | RunPod | 48 GB(2x 24GB) | $0.88($0.88/hr) | $1.38($1.38/hr) | Rent |
NVIDIA A100 SXM4 (80GB)Popular Ampere•NVLink-3 (600 GB/s) | Lambda Labs | 80 GB | $1.25($1.25/hr) | $1.69($1.69/hr) | Rent |
2x NVIDIA RTX 5090 Blackwell•PCIe 5.0 | RunPod | 64 GB(2x 32GB) | $1.38($1.38/hr) | $1.98($1.98/hr) | Rent |
NVIDIA A100 SXM4 (80GB) Ampere•NVLink-3 (600 GB/s) | RunPod | 80 GB | $1.39($1.39/hr) | $1.89($1.89/hr) | Rent |
NVIDIA A100 SXM4 (80GB) Ampere•NVLink-3 (600 GB/s) | GCP (a2-ultragpu-1g) | 80 GB | $1.75($1.75/hr) | $3.67($3.67/hr) | Rent |
4x NVIDIA RTX 4090 Ada Lovelace•PCIe 4.0 | RunPod | 96 GB(4x 24GB) | $1.76($1.76/hr) | $2.76($2.76/hr) | Rent |
NVIDIA H100 SXM5 (80GB) Hopper•NVLink-4 (900 GB/s) + InfiniBand | Nebius AI Studio | 80 GB | $2.19($2.19/hr) | $2.85($2.85/hr) | Rent |
NVIDIA H100 SXM5 (80GB)Popular Hopper•NVLink-4 (900 GB/s) | Lambda Labs | 80 GB | $2.29($2.29/hr) | $2.99($2.99/hr) | Rent |
NVIDIA H100 SXM5 (80GB) Hopper•NVLink-4 (900 GB/s) | CoreWeave | 80 GB | $2.35($2.35/hr) | $3.15($3.15/hr) | Rent |
NVIDIA H100 SXM5 (80GB)Popular Hopper•NVLink-4 (900 GB/s) | RunPod | 80 GB | $2.49($2.49/hr) | $3.29($3.29/hr) | Rent |
4x NVIDIA RTX 5090 Blackwell•PCIe 5.0 | RunPod | 128 GB(4x 32GB) | $2.76($2.76/hr) | $3.96($3.96/hr) | Rent |
NVIDIA H200 SXM5 (141GB)Popular Hopper HBM3e•NVLink-4 (900 GB/s) | Lambda Labs | 141 GB | $3.49($3.49/hr) | $4.29($4.29/hr) | Rent |
8x NVIDIA A100 SXM4 (80GB) Ampere•NVLink-3 (600 GB/s) | Lambda Labs | 640 GB(8x 80GB) | $9.90($9.90/hr) | $13.52($13.52/hr) | Rent |
8x NVIDIA A100 SXM4 (80GB) Ampere•NVLink-3 (600 GB/s) | AWS (EC2 p4de.24xlarge) | 640 GB(8x 80GB) | $14.50($14.50/hr) | $40.97($40.97/hr) | Rent |
8x NVIDIA H100 SXM5 (80GB)Popular Hopper•NVLink-4 (900 GB/s) + Quantum-2 IB | Lambda Labs | 640 GB(8x 80GB) | $18.30($18.30/hr) | $23.92($23.92/hr) | Rent |
8x NVIDIA H100 SXM5 (80GB) Hopper•NVLink-4 (900 GB/s) + 3.2Tbps InfiniBand | RunPod | 640 GB(8x 80GB) | $19.80($19.80/hr) | $26.32($26.32/hr) | Rent |
8x NVIDIA H200 SXM5 (141GB) Hopper HBM3e•NVLink-4 (900 GB/s) + Quantum-2 IB | Lambda Labs | 1128 GB(8x 141GB) | $27.50($27.50/hr) | $34.32($34.32/hr) | Rent |
8x NVIDIA H100 SXM5 (80GB) Hopper•NVLink-4 (900 GB/s) + GPUDirect-TCPX | GCP (a3-highgpu-8g) | 640 GB(8x 80GB) | $32.00($32.00/hr) | $87.05($87.05/hr) | Rent |
8x NVIDIA H100 SXM5 (80GB) Hopper•NVLink-4 (900 GB/s) + 3.2Tbps EFA | AWS (EC2 p5.48xlarge) | 640 GB(8x 80GB) | $38.50($38.50/hr) | $98.32($98.32/hr) | Rent |
8x NVIDIA Blackwell B200 (192GB)Popular Blackwell•NVLink-5 (1.8 TB/s) + Quantum-X800 IB | Nebius / CoreWeave | 1536 GB(8x 192GB) | $44.00($44.00/hr) | $58.00($58.00/hr) | Rent |
Frequently Asked Questions & Transformer Math
Deep-dive technical answers on KV-cache calculation, CUDA Out-Of-Memory prevention, QLoRA fine-tuning benchmarks, and cloud GPU cost optimization.
Popular LLM VRAM Requirements Guides
Explore dedicated memory calculators, architecture specifications, and GPU sizing guides for popular open-weights models.
Llama 3.3 70B Instruct
Meta's state-of-the-art dense 70B open-weights powerhouse with 128k context and 8:1 Grouped Query Attention (GQA).
Llama 3.1 8B Instruct
The industry standard compact LLM for local reasoning, autonomous agents, and fast fine-tuning pipelines.
DeepSeek R1 (671B MoE)
Frontier open reasoning model featuring 671B total parameters, 37B active per token, and Multi-Head Latent Attention (MLA).
DeepSeek V3 (671B MoE)
Flagship 671B mixture-of-experts generalist model with state-of-the-art coding and multilingual performance.
DeepSeek R1 Distill 70B (Llama)
DeepSeek reasoning tokens distilled into the Llama 3.3 70B transformer architecture.
DeepSeek R1 Distill 32B (Qwen)
Distilled reasoning traces inside Qwen 2.5 32B architecture, matching larger 70B frontier benchmarks.
DeepSeek R1 Distill 14B (Qwen)
High-efficiency coding and reasoning model tailored for single-GPU local workstations.
Qwen 2.5 72B Instruct
Top-tier multilingual, math, and code generation dense LLM with 128k sequence capability.
Qwen 2.5 32B Instruct
The golden sweet-spot model delivering 70B-grade capabilities with a 32B memory footprint.
Qwen 2.5 14B Instruct
Ultra-fast mid-size model excelling in function calling and agentic code execution.
Qwen 2.5 7B Instruct
High-throughput lightweight model optimized for edge deployments, real-time chatbots, and embeddings.
Mistral Large 2 (123B)
Mistral AI's flagship 123B dense model featuring 128k context, strong reasoning, and native multi-lingual support.
Mistral Small 3 24B
Enterprise-ready 24B reasoning model engineered for fast inference and lightweight fine-tuning.
Mixtral 8x7B MoE
Pioneering sparse Mixture of Experts model with 47B total parameters and 13B active per token.
Gemma 2 27B
Google DeepMind's highly efficient architecture combining sliding-window and global attention.
Gemma 2 9B
Compact powerhouse model with 256 head dimension delivering top-tier performance for its weight class.
Phi-4 (14B)
Microsoft Research's state-of-the-art synthetic data trained model with exceptional reasoning.
FLUX.1 [schnell] (12B Diffusion)
12B Rectified Flow Matching Transformer for ultra-fast, high-detail text-to-image synthesis.
Stable Diffusion XL 1.0 (3.5B)
Flagship 3.5B latent diffusion model with dual text encoders for photorealistic image generation.
Llama 3.2 11B Vision Instruct
Meta's flagship multimodal model pairing an 8B text LLM with a 3B vision encoder for high-resolution image reasoning.
Llama 3.2 90B Vision Instruct
Frontier 90B multimodal model with advanced visual document comprehension, chart parsing, and high-res image reasoning.
Qwen 2.5-VL 72B Instruct
Alibaba's frontier vision-language model with dynamic resolution processing and native video comprehension.
Qwen 2.5-VL 7B Instruct
High-efficiency open multimodal model capable of analyzing long documents, charts, and video streams at rapid speed.
Mistral Pixtral 12B Vision
Mistral's open-source 12B multimodal powerhouse featuring a 400M vision encoder capable of arbitrary image aspect ratios.
Popular Sizing & Hardware Recipes
Standard battle-tested deployment patterns for the most common open-source LLM workloads.
Llama 3.3 70B at Home
Running 70B on 1x RTX 4090 (24GB) using INT4 / AWQ quantization with 8k context window.
Qwen 2.5 32B FP8
Zero degradation 32B model on 1x RTX 5090 (32GB) or 1x L40S (48GB) with 32k context.
8B QLoRA Fine-Tuning
Full parameter adapter tuning for Llama 3.1 8B with rank=16 and batch size 4 on 1x RTX 3090.
DeepSeek-R1 (671B)
Running 671B full MoE in FP8 on an 8x H200 (141GB) or 4x B200 (192GB) supercluster.
Mathematical Methodology & Formulas
GPU memory allocation in modern PyTorch, vLLM, and TensorRT-LLM is governed by deterministic equations across weights, KV-cache, activations, and optimizer states.
1. Model Weights Memory
For FP16/BF16, each parameter requires 2 bytes. For INT4/AWQ, parameters take 0.55 bytes (including metadata, scaling factors, and zero-point overhead).
2. KV-Cache Context Memory
Modern architectures (Llama 3, Mistral, Qwen) use Grouped Query Attention (GQA), reducing the number of KV heads by 4x to 8x compared to query heads, drastically lowering context expansion costs.
3. Training & AdamW Optimizer States
Full fine-tuning requires 4 bytes for FP32 master weights, 4 bytes for first momentum m, and 4 bytes for second momentum v, totaling 12 bytes/param plus 2 bytes for gradients.
4. CUDA Runtime & PyTorch Fragmentation Buffer
PyTorch and CUDA runtime kernels consume ~1.5 GB of context overhead. An additional 15% headroom buffer prevents CUDA Out-Of-Memory (OOM) errors caused by memory fragmentation during dynamic batching.
Explore More Resources & Support
Learn about our methodology, get in touch for custom benchmarks, or review our privacy policies.