🚀 FOUNDER PASS LAUNCH:Get Full TensorRT-LLM Pro for only €4.99/mo (Locked for Life) — 814 Spots Left!Claim Pass →
NVIDIA HARDWARE & VRAM MATRIX

Global GPU Compatibility Engine

Determine memory requirements, throughput limits, and optimal quantization across all NVIDIA enterprise and consumer architectures.

1B (Edge)8B (Llama-3)70B405B (Frontier)
Blackwell$5.5/hr

NVIDIA B200 SXM (Blackwell)

Bandwidth: 8000 GB/s • TDP: 1000W
Memory Load:77.75 / 192 GB (40.5%)
Est. Throughput:104 tok/s
Hopper$3.2/hr

NVIDIA H100 SXM5 (Hopper)

Bandwidth: 3350 GB/s • TDP: 700W
Memory Load:77.75 / 80 GB (97.2%)
Est. Throughput:39 tok/s
Ampere$1.85/hr

NVIDIA A100 SXM4 80GB

Bandwidth: 2039 GB/s • TDP: 400W
Memory Load:77.75 / 80 GB (97.2%)
Est. Throughput:19 tok/s
Ada Lovelace$1.25/hr

NVIDIA L40S 48GB (Ada)

Bandwidth: 864 GB/s • TDP: 350W
Memory Load:77.75 / 48 GB (100%)
Est. Throughput:OOM
Ada Lovelace$0.55/hr

NVIDIA GeForce RTX 4090 24GB

Bandwidth: 1008 GB/s • TDP: 450W
Memory Load:77.75 / 24 GB (100%)
Est. Throughput:OOM
Ampere$0.38/hr

NVIDIA GeForce RTX 3090 24GB

Bandwidth: 936 GB/s • TDP: 350W
Memory Load:77.75 / 24 GB (100%)
Est. Throughput:OOM
NVIDIA TOPOLOGY & DISTRIBUTED PARALLELISM MAP

Multi-GPU Interconnect & Tensor Parallelism

Inspect how weights and KV Caches split across NVSwitch fabrics to achieve microsecond inter-GPU communication latency.

INTERCONNECT FABRIC:NVLink 4.0 NVSwitch
900 GB/s per GPU
GPU #0
80 GB
TP=1
GPU #1
80 GB
TP=2
GPU #2
80 GB
TP=3
GPU #3
80 GB
TP=4
GPU #4
80 GB
TP=5
GPU #5
80 GB
TP=6
GPU #6
80 GB
TP=7
GPU #7
80 GB
TP=8
Tensor Parallelism (TP)TP = 8 (All-Reduce)
Pipeline Parallelism (PP)PP = 1
Total Cluster VRAM640 GB Dedicated
Recommended Target Workload:Llama-3.1-70B FP16, Llama-3.1-70B FP8, DeepSeek-R1 INT4 AWQ