VRAM
24 GB
Memory
GDDR6X
Bandwidth
1,008 GB/s
FP16
330 TF
TDP
450 W
FP16 dense tensor with FP16 accumulate. The de facto self-host workhorse for quantized models up to ~32B.
Estimates based on INT8 quantization. Actual fit depends on framework and batch size.
The manufacturer's other cards in the catalog, largest memory first, with bandwidth and power where published.
Answered from the entry's own figures: memory, which models fit, bandwidth, power and the card's class.
NVIDIA GeForce RTX 4090 has 24 GB of GDDR6X at 1,008 GB/s. At INT8 a model's weights take about one byte per parameter, so the card holds roughly a 19B-parameter model with room for context; larger models need several cards or a lower precision.
From the catalog's estimates at INT8, NVIDIA Nemotron Nano 12B V2 VL BF16, Qwen 2.5 14B Instruct, Gemma 3 12B and Moonlight 16B A3B Instruct fit on a single NVIDIA GeForce RTX 4090 with their default context. The list on this page shows the estimated memory each takes; the capacity planner sizes any model against this card for your own context length, batch size and request rate.
1,008 GB/s. Bandwidth bounds how fast a model generates tokens, because every token reads the whole set of weights and the KV cache from memory; at this rate a 19B model at INT8 could be read on the order of 52 times a second on one card, before batching and framework efficiency.
NVIDIA GeForce RTX 4090 has a TDP of 450 W, in the PCIe form factor. Budget the TDP per card plus host overhead when sizing a server's power and cooling.
NVIDIA GeForce RTX 4090 is a consumer card from the gaming line, the lowest price per card and no inter-card link, built on the Ada Lovelace architecture and released in 2022. It has no multi-GPU interconnect, so a model must fit on one card or be split over the slower PCIe bus.
Specifications from the manufacturer's published material; a figure that is not published is shown as such, and a derived one is explained in the notes.
Added Jul 30, 2026 · Last updated Jul 30, 2026
Companion tools that draw on the same catalog and routing engine.
One gateway in front of every model, with your policies applied and every decision on record. Start with $5 of credit and 5,000 routing decisions a month, no card required.