VRAM
32 GB
Memory
GDDR7
Bandwidth
1792 GB/s
TDP
575W
Smaller Language Models
Inference for 7B-13B parameter models
FP16 dense tensor with FP16 accumulate. The highest-VRAM consumer card; common self-host choice for quantized 30B-70B models.
Estimates based on INT8 quantization. Actual fit depends on framework and batch size.
Added Jul 30, 2026
Last updated: Jul 30, 2026
Every month new models become cheaper, faster, and more capable. Inferbase ensures your application automatically benefits without changing a single API call.