VRAM
288 GB
Memory
HBM3e
Bandwidth
8000 GB/s
TDP
1400W
Large Language Models
Training and inference for models like GPT-4, Llama 70B+
Deep Learning Training
High-performance training for neural networks
Distributed Training
Multi-node training with fast interconnects
High-Throughput Inference
Optimized for batched inference workloads
Blackwell Ultra. FP8 rate identical to B200 (5/10 PF dense/sparse, so FP16 sparse figure matches B200 convention); NVFP4 15 PF dense. Shipping since Jan 2026. Liquid-cooled in all form factors.
Estimates based on INT8 quantization. Actual fit depends on framework and batch size.
Added Jul 30, 2026
Last updated: Jul 30, 2026
Every month new models become cheaper, faster, and more capable. Inferbase ensures your application automatically benefits without changing a single API call.