VRAM
64 GB
Memory
HBM2e
Bandwidth
1600 GB/s
TDP
550W
Medium Language Models
Inference for models up to 70B parameters
Deep Learning Training
High-performance training for neural networks
Enterprise Deployment
Designed for 24/7 datacenter operations
China market. Specs from Biren's own Hot Chips 34 disclosure: 1024 TFLOPS BF16, 2048 TOPS INT8, 64GB HBM2e at 1.6 TB/s. Volume production subsequently constrained by US export controls; successor SKUs (BR104) shipped in its place.
Estimates based on INT8 quantization. Actual fit depends on framework and batch size.
Added Jul 30, 2026
Last updated: Jul 30, 2026
Every month new models become cheaper, faster, and more capable. Inferbase ensures your application automatically benefits without changing a single API call.