Open-weight and frontier AI models with benchmarks, context windows and modalities, or the best models per task.
A curated catalog rather than a logo wall: every model is attributed to its creator, every variant is a variant, and every claim names its source or says it has none.
This models hub covers open-source LLMs such as Llama, DeepSeek, Qwen and Mistral alongside the frontier models from OpenAI, Anthropic and Google, across chat, embedding, audio, image and video. Every model is attributed to its creator, and a quantized or distilled variant is listed as a variant.
Each entry shows LLM benchmark scores from cited suites, the context window, the modalities and the license, so you can compare AI models on evidence rather than marketing. Where a model has no measurement, the entry says so. Models the managed pool serves are marked; the rest are listed for reference.
Open any entry for its benchmark scores by suite, its capabilities, its GPU requirements for self-hosting, and the models it is most often compared with.
Companion tools that draw on the same catalog and routing engine.
What is in it, where the numbers come from, and what you can run.
A curated hub of AI models: open-source LLMs such as Llama, DeepSeek, Qwen and Mistral alongside frontier models from OpenAI, Anthropic and Google, across chat, embedding, audio, image and video. Each entry records the creator, the variant, the context window, the modalities, the license, benchmark evidence where it exists, and whether Inferbase serves the model today.
Filter the catalog by creator, modality, context window, parameter count, license or serving status, then open any entry for its benchmark scores by suite, capabilities and GPU requirements. The Compare tab puts up to four models side by side on the same fields, and the Recommender turns a description of your workload into a ranked shortlist.
The open-weight families from the curated creators: Llama from Meta, DeepSeek, Qwen, Mistral, Gemma from Google, GLM from Z.ai, Kimi from Moonshot, MiniMax, Phi from Microsoft and Nemotron from NVIDIA, each with its verified variants. Filter by Open Source to see them.
The ones marked as served: chat models from the managed pool. Filter by Inference Ready to see them. Other models are listed for reference and comparison; connect your own provider keys and the gateway serves those on your accounts as well.
From public, cited benchmark suites ingested into the catalog and rank-normalized per suite. A score is shown only where a suite measured that model; a model with no measurement for a task is shown as missing rather than estimated.
Yes. The Modalities filter narrows the catalog by what a model accepts and produces: text, image, audio and video, in either direction. The In and Out columns show the same per model, so a vision model or a text-to-speech model is visible at a glance.
Creator listings and Hugging Face are re-ingested on the catalog sync cycle and benchmarks refresh weekly. The date on the catalog page is the last data refresh.
Attribution and evidence. A model is attributed to its creator, never to a host that resells it; a quantized or distilled variant is listed as a variant; and every quality figure names its suite. The catalog is the knowledge layer the router reads when it picks a model per request.
One gateway in front of every model, with your policies applied and every decision on record. Start with $5 of credit and 5,000 routing decisions a month, no card required.
Open the full entry: benchmarks by suite, capabilities, GPU requirements, and the models it is compared with.