Skip to main content
Product · Pricing

Pricing

Usage-based pricing for LLM routing and AI inference through one API. No subscription, no seats, a free tier to start.

Plans

Pay for what you use

Routing is priced per decision, with the choice of managed inference or your own models and providers behind it.

Free

$0to start

Try routing and inference on free credit. No card required.

  • $5 inference credit on signup
  • 5,000 free routing decisions every month
  • Smart routing and custom model pools
  • OpenAI-compatible API
  • Model catalog, comparison, and playground
  • Organizations, projects, and role-based access
  • Community support
Start free
Most popular

Pay-as-you-go

$1per 1,000 routed requests

Routing per decision; models on your own provider or on managed inference. No subscription, no seats.

Everything in Free, plus

  • $0.001 per routing decision beyond the free 5,000
  • Managed inference at provider cost plus a flat 6%
  • Pinned models are never charged a routing fee
  • No per-request minimum charge
  • $10 minimum top-up, cancel any time by not topping up
  • Higher rate limits
  • Usage dashboard with spend caps and budgets
Get started

Enterprise

Customlet's talk

For teams adopting at scale, on terms and controls agreed by contract.

Everything in Pay-as-you-go, plus

  • Volume and committed-use pricing
  • Invoicing in place of top-ups
  • Custom model and creator onboarding
  • Security review, DPA and procurement support
  • Dedicated support and onboarding
  • SSO, an uptime SLA and dedicated capacity, by contract
Contact sales

Routing is the Inferbase charge: $1 per 1,000 decisions, the first 5,000 each month free, and only when the router chooses the model. A pinned model skips it.

For the models, two ways to run them. Bring your own provider: connect your provider keys or endpoints and the tokens are billed by your provider, with nothing added by us. Or use managed inference: the gateway serves the model on its own capacity and you pay per token, at the rate above, from credit you add.

One request

One routed request, itemized

What a routed request on managed inference costs, line by line: the decision, the tokens, and a guardrail check when your policy runs one. An example in round numbers; every receipt prints the live ones.

On managed inference a routed request has two lines: the router's decision, then the tokens the chosen provider served. A guardrail policy can add a third. Every line is itemized on each request's receipt and in the monthly usage.

Decisions

$1 per 1,000 routing decisions, the first 5,000 each month free. A request that names a model skips the router and is never charged one.

Tokens

The provider's per-token price plus a flat 6%, read live from the provider and printed on every receipt. No minimum, no idle charge.

Guardrails

Off unless your policy turns them on. An unsafe-content check is a guard-model call, billed per token plus the same 6%. The injection check is free.

Your own provider

On your own keys or endpoints the tokens are billed by your provider. Inferbase charges the decision and any guardrail check.

One request · example ratesitemized
Routing decision1 of the 5,000 free each month, then $1 per 1,000
$0.001
Tokens in1,000 at $0.20 per 1M, the provider's price
$0.0002
Tokens out500 at $0.60 per 1M, the provider's price
$0.0003
Guardrail checkonly if your policy turns one on: 250 guard-model tokens at $0.20 per 1M
$0.00005
Inferbasea flat 6% on the tokens and the check
$0.000033
Total
$0.001583
while the decision is inside the free 5,000
$0.000583

The rates are examples; every receipt prints the live ones. On your own provider key the token lines are billed by the provider; a guardrail check is still charged, because the guard model runs on ours. A pinned model is never charged a decision.

Compare

Compare plans

Every control ships on every plan: pools, presets, failover, the decision audit, budgets. The plans differ on price, limits and support.

Feature comparison across the Free, Pay-as-you-go, and Enterprise plans
FeaturesFreePay-as-you-goEnterprise
Smart routing
Automatic model selection
Free routing decisions5,000 / mo5,000 / moCustom
Routing fee beyond the free allowance$0.001 / decisionVolume pricing
Custom model pools
Eligibility presets
Automatic fallback and retries
Routing decision audit trail
Inference
OpenAI-compatible endpoint
Managed inference priceAt cost + 6%At cost + 6%Negotiated
Streaming responses
$5 signup creditCustom
Minimum top-up$10Invoicing
Platform
Model catalog and comparison
Playground
Organizations and projects
Role-based access control
Scoped API keys
Usage dashboard
Spend caps and budgets
GPU capacity planner
Scale and support
Rate limit per key60 rpm defaultUp to 600 rpmCustom
SupportCommunityCommunityDedicated
SSOBy contract
Uptime SLABy contract
Dedicated capacityBy contract
Custom model onboarding
Compliance supportBy contract

Rate limits are per API key and can be set lower on any key. Enterprise terms such as SSO, an uptime SLA, dedicated capacity and compliance support are arranged by contract. Contact us for higher limits or a custom plan.

FAQ

Frequently asked questions

How routing is priced, the two ways to run models, what happens when credit runs out, and what is arranged by contract.

Put intelligence in the middle.

One gateway in front of every model, with your policies applied and every decision on record. Start with $5 of credit and 5,000 routing decisions a month, no card required.