Skip to main content
Company

About Inferbase

The idea

What Inferbase is

The control plane for AI models: one gateway in front of every model, with four capabilities behind it.

Inferbase is built on one observation: no single model is the right choice for every request, and the right choice keeps moving. New models arrive every few weeks and prices change under them; a model hard-wired in at design time leaves quality and cost on the table from day one.

So the product is a control plane rather than a proxy. One gateway serves every request through one OpenAI-compatible API, and four capabilities sit behind it. The catalog knows which models exist, which variants serve, what they cost and what the benchmarks say per task. Routing classifies each request and sends it to the model that wins on the objective you set. Security sets which models a key may use, inspects prompts, responses and tool results, and governs which tools an agent may call. Observability keeps the decision, the verdicts, the tokens and the cost of every request.

Your applications · SDK unchanged

Inferbase Gateway

One OpenAI-compatible API, with failover, rate limits and budgets behind it.

  1. knowledge

    Catalog

    Curated creators, verified variants, benchmark evidence where it exists, live prices.

  2. decision

    Routing

    A prompt classified, a model chosen on your objective, the decision kept.

  3. policy

    Guardrails

    Which models may serve, guardrails at three doors, which tools an agent may call.

  4. record

    Observability

    Decision, verdicts, tokens, cost and latency per request; usage and an audit log.

Providers · your keys or the managed catalog
Mechanism

How it works

One call in, one decision out before the first token, the response streamed behind it. The mechanism, stage by stage.

Tenets

How we think about routing

Automatic does not have to mean opaque. Three rules keep the router honest.

01

Every decision is auditable

Each route records the task the request was classified as, the models that qualified, the one that won and how clearly. Nothing happens in a black box.

02

Unknown is an honest answer

A model with no measurement for a task is treated as unknown, never assumed to match the model it was derived from or the size class it sits in.

03

You set the objective

You choose what to optimize for per key or per request: cost, balanced, quality or latency. Routing decides within that; a named model is served as named.

Principles

Our principles

The bar every feature and every number on this site is held to.

Accuracy with speed

Benchmarks are matched to the catalog and checked, unknowns are marked unknown, and nothing is dressed up as more certain than it is.

Cost is a first-class metric

Price is a routing dimension beside quality and latency, read live from the hosts that serve each model, never an afterthought.

Evidence before claims

A problem is quantified before anything is built for it, and a figure on this site names its measurement and date or is not stated.

Built for practitioners

Made for people shipping real systems: one API, no SDK to adopt, receipts and decisions by request id, and nothing of ours to host.

You stay in control

Routing decides within the scope you set: the objective, the pool, the preset, your own keys. A named model is always served as named.

Knowledge, not possession

The plane decides and governs on what it knows about models, whether a model runs on the managed pool or on your own accounts.

Put intelligence in the middle.

One gateway in front of every model, with your policies applied and every decision on record. Start with $5 of credit and 5,000 routing decisions a month, no card required.