AI model recommendations for your use case, ranked on benchmark evidence and your priorities, or the best models per task.
Your industry, use case, scale and priorities become a shortlist from the catalog, on the same evidence the router uses per request.
The recommender maps a use case to the tasks it needs, keeps the models that clear the quality bar for those tasks, and ranks them on the priorities you set: quality, cost, latency or a balance. A capable open-source LLM sits alongside the frontier models when one exists, and nothing is ranked on guesswork.
Each recommendation shows the tasks it covers, the benchmark evidence behind its rank, its context window and modalities, its license, and whether it is served on Inferbase today, so a pick can be checked in the catalog, compared side by side, or run in the playground.
Prefer a ready-made list? The best-models rankings apply the same method per task, and a pinned model in the gateway skips the choice altogether.
The fainter marks are the other two floors, permissive lower and strict higher. Price is compared among the models marked in; a preset that leaves nobody in relaxes one step, and the relaxation is on the decision record.
Companion tools that draw on the same catalog and routing engine.
Which model to use, how the shortlist is ranked, what it will not do, and where its data comes from.
Answer four questions, your industry, the use case, the scale you run at and up to two priorities, and the recommender returns a ranked shortlist of five models from the catalog. Each recommendation shows the evidence behind its rank, so the answer is a shortlist you can check rather than a single verdict.
Every visible model is scored on three parts: fit for the use case, which is the largest share and reads the benchmark suites relevant to that use case plus the context window; your priorities; and a small popularity share. A model with no benchmark evidence for the use case is left out rather than guessed at, and the same catalog data feeds the router.
Pick the use case and the recommender ranks models on the suites that measure it: code generation and code review under Software, chatbots and sentiment analysis under Customer experience, summarization and data extraction under Operations, translation and content writing under Content, and the rest across finance, healthcare, legal, research, manufacturing and retail. A use case that is not listed goes in as a custom description.
Yes. Choose up to two of cost efficiency, best quality, speed, privacy and self-hosting, or easy integration; the first counts for about two thirds of the priority share and the second for the rest. Cost is judged for the scale you chose, and privacy favours open-weight models you can host yourself.
No. Recommendations are ranked on technical fit, benchmark evidence and the priorities you set. A cheaper open model wins when it fits the job, and nothing in the ranking depends on what a model earns Inferbase.
Yes. Open-weight models are scored alongside proprietary ones, the shortlist surfaces at least one open option when a capable one exists for your use case, and an Open source only switch on the results narrows the list to models you can self-host.
Enter a custom use case in step 2, which adds a keyword match on your description to the benchmark and context scoring, or go to the model catalog to filter and compare models directly. The best-models rankings list the top models per task with the same evidence.
The catalog is re-ingested on a schedule from model registries and host APIs, and benchmark scores are refreshed weekly, so new releases reach the recommender without a code change. The date on the catalog page is the last data refresh.
One gateway in front of every model, with your policies applied and every decision on record. Start with $5 of credit and 5,000 routing decisions a month, no card required.