VRAM
48 GB
Memory
GDDR6 ECC
Bandwidth
960 GB/s
TDP
300W
Medium Language Models
Inference for models up to 70B parameters
FP16 dense tensor per NVIDIA datasheet. Previous-gen 48GB workstation card, still common in on-prem inference builds.
Estimates based on INT8 quantization. Actual fit depends on framework and batch size.
Added Jul 30, 2026
Last updated: Jul 30, 2026
Every month new models become cheaper, faster, and more capable. Inferbase ensures your application automatically benefits without changing a single API call.