Model Comparison
Compare AI models side by side on benchmark scores, context window, modalities and capabilities, from the catalog’s own evidence.
Source: inferbase.ai
Compare AI models side by side on benchmark scores, context window, modalities and capabilities, from the catalog’s own evidence.
Use the search bar above to find and add a model for comparison.
Use the search bar above to find and add a model for comparison.
Every figure on a comparison comes from the catalog: published specifications, capabilities as deployed, and benchmark scores by suite.
Specifications. Factual attributes published by the model's creator: context window, maximum output, parameters, modalities, license and release date. Open-source LLMs and frontier models compare on the same fields.
Capabilities. What each deployment honours, such as function calling, structured output and extended thinking, is checked against a host's capability flag or the creator's documentation, and the entry keeps the source. A feature that has not been checked reads as unknown rather than as absent.
Benchmarks. Standardized test scores listed exactly as each source published them, grouped by source because scores from different suites are not comparable with each other. A cell reads as not available when a suite has not measured a model, rather than being filled with an estimate. Inferbase does not score the models itself.
Companion tools that draw on the same catalog and routing engine.
How a comparison is built, what the cells mean, and what you can do with one.
Search for a model in each slot, or open a model page and choose Add to Compare. Up to four models sit side by side on the same specifications, capabilities and benchmark scores. The URL of a comparison is shareable, and a signed-in export as CSV or PDF also files it under your saved comparisons.
Public, cited suites ingested into the catalog, grouped by source. Scores are shown as each source published them and are not combined across suites, because a score on one suite is not comparable with a score on another.
A benchmark is one suite measuring one thing, such as graduate-level science questions or agentic coding; a leaderboard folds many suites into a single rank. This comparison shows the suites and keeps them apart, so you can read the measurement that matters for your task rather than a blended position.
Because that suite has not measured that model. The cell stays empty rather than being estimated, so a comparison never rests on a guessed figure.
Context window and maximum output are specification rows on every comparison, as the creator published them. The figure is the advertised maximum; the context a particular host serves can be lower, and the gateway routes on what each host actually serves.
Yes. The catalog holds both, and a comparison puts them on the same fields: context window, modalities, license, capabilities as deployed, and benchmark evidence where it exists. Llama, DeepSeek, Qwen and Mistral compare against GPT, Claude and Gemini models on the same rows.
Read the suites that measure the task in the benchmark rows, the coding and agentic suites for coding, for example. For a task-first answer, the best-models rankings list the top models per task from the same evidence, and the recommender turns a use case into a shortlist.
Models the managed pool serves carry a Try now button that opens the playground with that model pinned. The rest are listed for reference; connect your own provider key and the gateway serves those on your account.
One gateway in front of every model, with your policies applied and every decision on record. Start with $5 of credit and 5,000 routing decisions a month, no card required.
A comparison sets two to four such columns side by side; a figure missing here is missing there too. Open the entry for every suite and capability.