VRAM
96 GB
Memory
GDDR7 ECC
Bandwidth
1597 GB/s
TDP
600W
Large Language Models
Training and inference for models like GPT-4, Llama 70B+
Deep Learning Training
High-performance training for neural networks
Enterprise Deployment
Designed for 24/7 datacenter operations
Passive dual-slot server variant. FP4 4 PF / FP8 2 PF / FP16 1 PF dense per NVIDIA datasheet. Popular single-card inference option in 1U/2U racks.
Estimates based on INT8 quantization. Actual fit depends on framework and batch size.
Added Jul 30, 2026
Last updated: Jul 30, 2026
Every month new models become cheaper, faster, and more capable. Inferbase ensures your application automatically benefits without changing a single API call.