GPUCalcPROVRAM & Cloud Cost Lab
VRAM Sizer CalculatorLive GPU MatrixFormulas & MathAbout GPUCalcContact & FeedbackPrivacy Policy
Mistral AIby Mistral AI

Mistral Large 2 (123B) VRAM & Cloud GPU Sizer

Mistral AI's flagship 123B dense model featuring 128k context, strong reasoning, and native multi-lingual support. Use this interactive calculator to compute exact GPU memory requirements across FP16, FP8, INT4, QLoRA, and Full Fine-Tuning.

Inference VRAM (INT4 AWQ)
~68 GB
2x A100 80GB (FP8) or 4x RTX 4090 (INT4)
Native Precision (FP16 / BF16)
~234.11 GB
Requires 2x-4x 80GB GPUs
QLoRA 4-bit Training
~120.03 GB
8x H100 SXM5 80GB
Interactive VRAM Sizer & Live Cloud Pricing Matrix Active

Popular Model Presets

Auto-Architected
B
0.5B (Edge)7B14B32B70B140B+
0.55 B/param
tokens
reqs
Minimum Required VRAMActive
68GB
Recommended Target:78.2 GB(+15% buffer)
Suggested Minimum GPU Tier:
1x A100 / H100 (80GB)

VRAM Memory Footprint Breakdown

Sum: 68 GB
Model Weights
63 GB
KV-Cache Context
2.75 GB
Activations
0.75 GB
CUDA & Runtime
1.5 GB
Estimated Decoding Speed (Tokens/sec):
A100 (80GB): 21 t/sH100 (80GB): 34.5 t/s

Live Cloud GPU Cost & Pricing Engine

18 Available Nodes

Real-time verified pricing across RunPod, Lambda Labs, Vast.ai, AWS, GCP, and specialized clouds.

Lowest Cost Compatible OptionNVIDIA A100 SXM4 (80GB) on Lambda Labs (80GB VRAM)
Spot Rate$1.25 / 1h
Rent Node
Buy vs. Rent TCO LabHardware break-even simulator
Providers:
GPU & ArchitectureProvider
Total VRAM
Spot (1h)
On-Demand (1h)
Action
NVIDIA A100 SXM4 (80GB)Popular
AmpereNVLink-3 (600 GB/s)
Lambda Labs
80 GB
$1.25($1.25/hr)$1.69($1.69/hr)Rent
NVIDIA A100 SXM4 (80GB)
AmpereNVLink-3 (600 GB/s)
RunPod
80 GB
$1.39($1.39/hr)$1.89($1.89/hr)Rent
NVIDIA A100 SXM4 (80GB)
AmpereNVLink-3 (600 GB/s)
GCP (a2-ultragpu-1g)
80 GB
$1.75($1.75/hr)$3.67($3.67/hr)Rent
4x NVIDIA RTX 4090
Ada LovelacePCIe 4.0
RunPod
96 GB(4x 24GB)
$1.76($1.76/hr)$2.76($2.76/hr)Rent
NVIDIA H100 SXM5 (80GB)
HopperNVLink-4 (900 GB/s) + InfiniBand
Nebius AI Studio
80 GB
$2.19($2.19/hr)$2.85($2.85/hr)Rent
NVIDIA H100 SXM5 (80GB)Popular
HopperNVLink-4 (900 GB/s)
Lambda Labs
80 GB
$2.29($2.29/hr)$2.99($2.99/hr)Rent
NVIDIA H100 SXM5 (80GB)
HopperNVLink-4 (900 GB/s)
CoreWeave
80 GB
$2.35($2.35/hr)$3.15($3.15/hr)Rent
NVIDIA H100 SXM5 (80GB)Popular
HopperNVLink-4 (900 GB/s)
RunPod
80 GB
$2.49($2.49/hr)$3.29($3.29/hr)Rent
4x NVIDIA RTX 5090
BlackwellPCIe 5.0
RunPod
128 GB(4x 32GB)
$2.76($2.76/hr)$3.96($3.96/hr)Rent
NVIDIA H200 SXM5 (141GB)Popular
Hopper HBM3eNVLink-4 (900 GB/s)
Lambda Labs
141 GB
$3.49($3.49/hr)$4.29($4.29/hr)Rent
8x NVIDIA A100 SXM4 (80GB)
AmpereNVLink-3 (600 GB/s)
Lambda Labs
640 GB(8x 80GB)
$9.90($9.90/hr)$13.52($13.52/hr)Rent
8x NVIDIA A100 SXM4 (80GB)
AmpereNVLink-3 (600 GB/s)
AWS (EC2 p4de.24xlarge)
640 GB(8x 80GB)
$14.50($14.50/hr)$40.97($40.97/hr)Rent
8x NVIDIA H100 SXM5 (80GB)Popular
HopperNVLink-4 (900 GB/s) + Quantum-2 IB
Lambda Labs
640 GB(8x 80GB)
$18.30($18.30/hr)$23.92($23.92/hr)Rent
8x NVIDIA H100 SXM5 (80GB)
HopperNVLink-4 (900 GB/s) + 3.2Tbps InfiniBand
RunPod
640 GB(8x 80GB)
$19.80($19.80/hr)$26.32($26.32/hr)Rent
8x NVIDIA H200 SXM5 (141GB)
Hopper HBM3eNVLink-4 (900 GB/s) + Quantum-2 IB
Lambda Labs
1128 GB(8x 141GB)
$27.50($27.50/hr)$34.32($34.32/hr)Rent
8x NVIDIA H100 SXM5 (80GB)
HopperNVLink-4 (900 GB/s) + GPUDirect-TCPX
GCP (a3-highgpu-8g)
640 GB(8x 80GB)
$32.00($32.00/hr)$87.05($87.05/hr)Rent
8x NVIDIA H100 SXM5 (80GB)
HopperNVLink-4 (900 GB/s) + 3.2Tbps EFA
AWS (EC2 p5.48xlarge)
640 GB(8x 80GB)
$38.50($38.50/hr)$98.32($98.32/hr)Rent
8x NVIDIA Blackwell B200 (192GB)Popular
BlackwellNVLink-5 (1.8 TB/s) + Quantum-X800 IB
Nebius / CoreWeave
1536 GB(8x 192GB)
$44.00($44.00/hr)$58.00($58.00/hr)Rent
Showing all 18 nodes • Scroll table vertically to view full hardware catalog
Prices updated February 2026
Technical Knowledge Hub

Frequently Asked Questions & Transformer Math

Deep-dive technical answers on KV-cache calculation, CUDA Out-Of-Memory prevention, QLoRA fine-tuning benchmarks, and cloud GPU cost optimization.

Mistral Large 2 (123B) Architectural Specs & Memory Scaling

Detailed breakdown of transformer parameters, attention mechanisms, and KV-cache expansion.

Total Parameters123 Billion
Transformer Layers88 Layers
Attention Heads / GQA64 Q / 8 KV (8:1)
Max Native Context1,28,000 Tokens

Quantization Precision vs VRAM Footprint

Precision / Quant TypeBytes / ParamModel Weights VRAMCompatible Single GPU
FP32 (32-bit Float)4 B458.21 GBRequires Multi-GPU / A100 (80GB)
FP16 (16-bit Float)2 B229.11 GBRequires Multi-GPU / A100 (80GB)
FP8 (8-bit Float)1 B114.55 GBRequires Multi-GPU / A100 (80GB)
INT8 (8-bit Integer)1 B114.55 GBRequires Multi-GPU / A100 (80GB)
INT4 / AWQ / GPTQ0.55 B63 GBRequires Multi-GPU / A100 (80GB)
GGUF Q4_K_M0.58 B66.44 GBRequires Multi-GPU / A100 (80GB)

KV-Cache Memory Growth with Sequence Length

Context Window (Tokens)FP16 KV Cache (2 B)FP8 KV Cache (1 B - vLLM)VRAM Savings
4,096 tokens1.38 GB0.69 GB-50% reduction
8,192 tokens2.75 GB1.38 GB-50% reduction
16,384 tokens5.5 GB2.75 GB-50% reduction
32,768 tokens11 GB5.5 GB-50% reduction
65,536 tokens22 GB11 GB-50% reduction

Compare Similar Model Calculators

Explore VRAM footprints for related open-weights architectures.