Dual-slot PCIe form factor of Intel's third-generation Gaudi accelerator, with 128 GB HBM2e and 24 integrated 200 GbE ports for scale-out via standard Ethernet (no proprietary fabric needed). Targets cost-conscious LLM inference and fine-tuning.
VRAM
128 GB
Memory
HBM2e
Bandwidth
3,700 GB/s
FP16
1,835 TF
TDP
600 W
compute_cores reflects 64 TPC (Tensor Processor Cores); ai_accelerators is the count of dedicated MME (Matrix Multiplication Engines). Compute figures are dense (no sparsity).
Estimates based on INT8 quantization. Actual fit depends on framework and batch size.
The manufacturer's other cards in the catalog, largest memory first, with bandwidth and power where published.
Answered from the entry's own figures: memory, which models fit, bandwidth, power and the card's class.
Intel Gaudi 3 PCIe (HL-338) has 128 GB of HBM2e at 3,700 GB/s. At INT8 a model's weights take about one byte per parameter, so the card holds roughly a 102B-parameter model with room for context; larger models need several cards or a lower precision.
From the catalog's estimates at INT8, Llama 3.1 Nemotron 70B Instruct HF, Gemma 4 31B IT, Llama Nemotron Embed VL 1B V2 and GLM Z1 9B 0414 fit on a single Intel Gaudi 3 PCIe (HL-338) with their default context. The list on this page shows the estimated memory each takes; the capacity planner sizes any model against this card for your own context length, batch size and request rate.
3,700 GB/s. Bandwidth bounds how fast a model generates tokens, because every token reads the whole set of weights and the KV cache from memory; at this rate a 102B model at INT8 could be read on the order of 36 times a second on one card, before batching and framework efficiency.
Intel Gaudi 3 PCIe (HL-338) has a TDP of 600 W, with a maximum of 600 W, in the PCIe form factor. Budget the TDP per card plus host overhead when sizing a server's power and cooling.
Intel Gaudi 3 PCIe (HL-338) is a datacenter accelerator, meant for servers and multi-GPU serving, built on the Habana Gaudi 3 architecture and released in 2024. It supports a multi-GPU interconnect at 600 GB/s, so a model larger than one card can be split across several.
Specifications from the manufacturer's published material; a figure that is not published is shown as such, and a derived one is explained in the notes.
Added Apr 30, 2026 · Last updated Apr 30, 2026
Companion tools that draw on the same catalog and routing engine.
One gateway in front of every model, with your policies applied and every decision on record. Start with $5 of credit and 5,000 routing decisions a month, no card required.