VRAM
48 GB
Memory
GDDR6
Bandwidth
768 GB/s
TDP
-
Medium Language Models
Inference for models up to 70B parameters
Enterprise Deployment
Designed for 24/7 datacenter operations
China market. Vendor-published specs: 100 TFLOPS FP16/BF16, 200 TOPS INT8, 48GB GDDR6 at 768 GB/s, MTLink multi-card interconnect. Used domestically for LLM training clusters.
Estimates based on INT8 quantization. Actual fit depends on framework and batch size.
Added Jul 30, 2026
Last updated: Jul 30, 2026
Every month new models become cheaper, faster, and more capable. Inferbase ensures your application automatically benefits without changing a single API call.