Air-cooled CDNA 4 accelerator with 288 GB of HBM3e and 8 TB/s of memory bandwidth, designed for retrofit into existing MI300/MI325 OAM platforms. New native FP4 and FP6 datatypes target inference cost-per-token reductions over Hopper-class hardware.
VRAM
288 GB
Memory
HBM3e
Bandwidth
8,000 GB/s
FP16
1,840 TF
TDP
1000 W
Spec verification recommended: numbers reflect AMD's Advancing AI 2025 announcement; double-check against the published MI350X datasheet before relying for production sizing.
Estimates based on INT8 quantization. Actual fit depends on framework and batch size.
The manufacturer's other cards in the catalog, largest memory first, with bandwidth and power where published.
Answered from the entry's own figures: memory, which models fit, bandwidth, power and the card's class.
AMD Instinct MI350X has 288 GB of HBM3e at 8,000 GB/s. At INT8 a model's weights take about one byte per parameter, so the card holds roughly a 230B-parameter model with room for context; larger models need several cards or a lower precision.
From the catalog's estimates at INT8, Qwen3 235B A22B Thinking 2507, Minimax M2.7, Qwen3 VL 235B A22B Thinking and Minimax M2.1 fit on a single AMD Instinct MI350X with their default context. The list on this page shows the estimated memory each takes; the capacity planner sizes any model against this card for your own context length, batch size and request rate.
8,000 GB/s. Bandwidth bounds how fast a model generates tokens, because every token reads the whole set of weights and the KV cache from memory; at this rate a 230B model at INT8 could be read on the order of 35 times a second on one card, before batching and framework efficiency.
AMD Instinct MI350X has a TDP of 1000 W, with a maximum of 1000 W, in the OAM form factor. Budget the TDP per card plus host overhead when sizing a server's power and cooling.
AMD Instinct MI350X is a datacenter accelerator, meant for servers and multi-GPU serving, built on the CDNA 4 architecture and released in 2025. It supports a multi-GPU interconnect at 1,075 GB/s, so a model larger than one card can be split across several.
Specifications from the manufacturer's published material; a figure that is not published is shown as such, and a derived one is explained in the notes.
Added Apr 30, 2026 · Last updated Apr 30, 2026
Companion tools that draw on the same catalog and routing engine.
One gateway in front of every model, with your policies applied and every decision on record. Start with $5 of credit and 5,000 routing decisions a month, no card required.