GPUCalcPROVRAM & Cloud Cost Lab
VRAM Sizer CalculatorLive GPU MatrixFormulas & MathAbout GPUCalcContact & FeedbackPrivacy Policy
Mistral AIMoE Sparse Architectureby Mistral AI

Mixtral 8x7B MoE VRAM & Cloud GPU Sizer

Pioneering sparse Mixture of Experts model with 47B total parameters and 13B active per token. Use this interactive calculator to compute exact GPU memory requirements across FP16, FP8, INT4, QLoRA, and Full Fine-Tuning.

Inference VRAM (INT4 AWQ)
~26.72 GB
1x RTX 4090 (24GB INT4) or 1x A100 80GB
Native Precision (FP16 / BF16)
~89.79 GB
Fits on 1x 24GB GPU
QLoRA 4-bit Training
~40.51 GB
2x A100 80GB (QLoRA)
Interactive VRAM Sizer & Live Cloud Pricing Matrix Active

Popular Model Presets

Auto-Architected
B
0.5B (Edge)7B14B32B70B140B+
2 B/param
tokens
reqs
Minimum Required VRAMActive
89.79GB
Recommended Target:103.25 GB(+15% buffer)
Suggested Minimum GPU Tier:
2x A100 / H100 (80GB)

VRAM Memory Footprint Breakdown

Sum: 89.79 GB
Model Weights
86.99 GB
KV-Cache Context
1 GB
Activations
0.3 GB
CUDA & Runtime
1.5 GB

Live Cloud GPU Cost & Pricing Engine

11 Available Nodes

Real-time verified pricing across RunPod, Lambda Labs, Vast.ai, AWS, GCP, and specialized clouds.

Lowest Cost Compatible Option4x NVIDIA RTX 4090 on RunPod (96GB VRAM)
Spot Rate$1.76 / 1h
Rent Node
Buy vs. Rent TCO LabHardware break-even simulator
Providers:
GPU & ArchitectureProvider
Total VRAM
Spot (1h)
On-Demand (1h)
Action
4x NVIDIA RTX 4090
Ada LovelacePCIe 4.0
RunPod
96 GB(4x 24GB)
$1.76($1.76/hr)$2.76($2.76/hr)Rent
4x NVIDIA RTX 5090
BlackwellPCIe 5.0
RunPod
128 GB(4x 32GB)
$2.76($2.76/hr)$3.96($3.96/hr)Rent
NVIDIA H200 SXM5 (141GB)Popular
Hopper HBM3eNVLink-4 (900 GB/s)
Lambda Labs
141 GB
$3.49($3.49/hr)$4.29($4.29/hr)Rent
8x NVIDIA A100 SXM4 (80GB)
AmpereNVLink-3 (600 GB/s)
Lambda Labs
640 GB(8x 80GB)
$9.90($9.90/hr)$13.52($13.52/hr)Rent
8x NVIDIA A100 SXM4 (80GB)
AmpereNVLink-3 (600 GB/s)
AWS (EC2 p4de.24xlarge)
640 GB(8x 80GB)
$14.50($14.50/hr)$40.97($40.97/hr)Rent
8x NVIDIA H100 SXM5 (80GB)Popular
HopperNVLink-4 (900 GB/s) + Quantum-2 IB
Lambda Labs
640 GB(8x 80GB)
$18.30($18.30/hr)$23.92($23.92/hr)Rent
8x NVIDIA H100 SXM5 (80GB)
HopperNVLink-4 (900 GB/s) + 3.2Tbps InfiniBand
RunPod
640 GB(8x 80GB)
$19.80($19.80/hr)$26.32($26.32/hr)Rent
8x NVIDIA H200 SXM5 (141GB)
Hopper HBM3eNVLink-4 (900 GB/s) + Quantum-2 IB
Lambda Labs
1128 GB(8x 141GB)
$27.50($27.50/hr)$34.32($34.32/hr)Rent
8x NVIDIA H100 SXM5 (80GB)
HopperNVLink-4 (900 GB/s) + GPUDirect-TCPX
GCP (a3-highgpu-8g)
640 GB(8x 80GB)
$32.00($32.00/hr)$87.05($87.05/hr)Rent
8x NVIDIA H100 SXM5 (80GB)
HopperNVLink-4 (900 GB/s) + 3.2Tbps EFA
AWS (EC2 p5.48xlarge)
640 GB(8x 80GB)
$38.50($38.50/hr)$98.32($98.32/hr)Rent
8x NVIDIA Blackwell B200 (192GB)Popular
BlackwellNVLink-5 (1.8 TB/s) + Quantum-X800 IB
Nebius / CoreWeave
1536 GB(8x 192GB)
$44.00($44.00/hr)$58.00($58.00/hr)Rent
Showing all 11 nodes • Scroll table vertically to view full hardware catalog
Prices updated February 2026
Technical Knowledge Hub

Frequently Asked Questions & Transformer Math

Deep-dive technical answers on KV-cache calculation, CUDA Out-Of-Memory prevention, QLoRA fine-tuning benchmarks, and cloud GPU cost optimization.

Mixtral 8x7B MoE Architectural Specs & Memory Scaling

Detailed breakdown of transformer parameters, attention mechanisms, and KV-cache expansion.

Total Parameters46.7 Billion
Transformer Layers32 Layers
Attention Heads / GQA32 Q / 8 KV (4:1)
Max Native Context32,768 Tokens

Quantization Precision vs VRAM Footprint

Precision / Quant TypeBytes / ParamModel Weights VRAMCompatible Single GPU
FP32 (32-bit Float)4 B173.97 GBRequires Multi-GPU / A100 (80GB)
FP16 (16-bit Float)2 B86.99 GBRequires Multi-GPU / A100 (80GB)
FP8 (8-bit Float)1 B43.49 GB✓ Fits on 1x L40S / RTX 6000 (48GB)
INT8 (8-bit Integer)1 B43.49 GB✓ Fits on 1x L40S / RTX 6000 (48GB)
INT4 / AWQ / GPTQ0.55 B23.92 GB✓ Fits on 1x RTX 4090 (24GB)
GGUF Q4_K_M0.58 B25.23 GB✓ Fits on 1x RTX 5090 (32GB)

KV-Cache Memory Growth with Sequence Length

Context Window (Tokens)FP16 KV Cache (2 B)FP8 KV Cache (1 B - vLLM)VRAM Savings
4,096 tokens0.5 GB0.25 GB-50% reduction
8,192 tokens1 GB0.5 GB-50% reduction
16,384 tokens2 GB1 GB-50% reduction
32,768 tokens4 GB2 GB-50% reduction

Compare Similar Model Calculators

Explore VRAM footprints for related open-weights architectures.