VRAM
32 GB
Memory
GDDR6
Bandwidth
512 GB/s
TDP
300W
Smaller Language Models
Inference for 7B-13B parameter models
Enterprise Deployment
Designed for 24/7 datacenter operations
Open-source stack (TT-Metalium). 664 TFLOPS BLOCKFP8 / 332 BF16 per Tenstorrent; 120 Tensix cores (cards downgraded from 140 via firmware, Jan 2026). 4x QSFP-DD 800G ports pool memory across cards. $1,399 active or passive.
Estimates based on INT8 quantization. Actual fit depends on framework and batch size.
Added Jul 30, 2026
Last updated: Jul 30, 2026
Every month new models become cheaper, faster, and more capable. Inferbase ensures your application automatically benefits without changing a single API call.