About Inferbase
What Inferbase is
The control plane for AI models: one gateway in front of every model, with four capabilities behind it.
Inferbase is built on one observation: no single model is the right choice for every request, and the right choice keeps moving. New models arrive every few weeks and prices change under them; a model hard-wired in at design time leaves quality and cost on the table from day one.
So the product is a control plane rather than a proxy. One gateway serves every request through one OpenAI-compatible API, and four capabilities sit behind it. The catalog knows which models exist, which variants serve, what they cost and what the benchmarks say per task. Routing classifies each request and sends it to the model that wins on the objective you set. Security sets which models a key may use, inspects prompts, responses and tool results, and governs which tools an agent may call. Observability keeps the decision, the verdicts, the tokens and the cost of every request.
Inferbase Gateway
One OpenAI-compatible API, with failover, rate limits and budgets behind it.
- knowledge
Catalog
Curated creators, verified variants, benchmark evidence where it exists, live prices.
- decision
Routing
A prompt classified, a model chosen on your objective, the decision kept.
- policy
Guardrails
Which models may serve, guardrails at three doors, which tools an agent may call.
- record
Observability
Decision, verdicts, tokens, cost and latency per request; usage and an audit log.
How it works
One call in, one decision out before the first token, the response streamed behind it. The mechanism, stage by stage.
How we think about routing
Automatic does not have to mean opaque. Three rules keep the router honest.
Every decision is auditable
Each route records the task the request was classified as, the models that qualified, the one that won and how clearly. Nothing happens in a black box.
Unknown is an honest answer
A model with no measurement for a task is treated as unknown, never assumed to match the model it was derived from or the size class it sits in.
You set the objective
You choose what to optimize for per key or per request: cost, balanced, quality or latency. Routing decides within that; a named model is served as named.
Our principles
The bar every feature and every number on this site is held to.
Accuracy with speed
Benchmarks are matched to the catalog and checked, unknowns are marked unknown, and nothing is dressed up as more certain than it is.
Cost is a first-class metric
Price is a routing dimension beside quality and latency, read live from the hosts that serve each model, never an afterthought.
Evidence before claims
A problem is quantified before anything is built for it, and a figure on this site names its measurement and date or is not stated.
Built for practitioners
Made for people shipping real systems: one API, no SDK to adopt, receipts and decisions by request id, and nothing of ours to host.
You stay in control
Routing decides within the scope you set: the objective, the pool, the preset, your own keys. A named model is always served as named.
Knowledge, not possession
The plane decides and governs on what it knows about models, whether a model runs on the managed pool or on your own accounts.
Put intelligence in the middle.
One gateway in front of every model, with your policies applied and every decision on record. Start with $5 of credit and 5,000 routing decisions a month, no card required.