Pricing
Usage-based pricing for LLM routing and AI inference through one API. No subscription, no seats, a free tier to start.
Pay for what you use
Routing is priced per decision, with the choice of managed inference or your own models and providers behind it.
Free
Try routing and inference on free credit. No card required.
- $5 inference credit on signup
- 5,000 free routing decisions every month
- Smart routing and custom model pools
- OpenAI-compatible API
- Model catalog, comparison, and playground
- Organizations, projects, and role-based access
- Community support
Pay-as-you-go
Routing per decision; models on your own provider or on managed inference. No subscription, no seats.
Everything in Free, plus
- $0.001 per routing decision beyond the free 5,000
- Managed inference at provider cost plus a flat 6%
- Pinned models are never charged a routing fee
- No per-request minimum charge
- $10 minimum top-up, cancel any time by not topping up
- Higher rate limits
- Usage dashboard with spend caps and budgets
Enterprise
For teams adopting at scale, on terms and controls agreed by contract.
Everything in Pay-as-you-go, plus
- Volume and committed-use pricing
- Invoicing in place of top-ups
- Custom model and creator onboarding
- Security review, DPA and procurement support
- Dedicated support and onboarding
- SSO, an uptime SLA and dedicated capacity, by contract
Routing is the Inferbase charge: $1 per 1,000 decisions, the first 5,000 each month free, and only when the router chooses the model. A pinned model skips it.
For the models, two ways to run them. Bring your own provider: connect your provider keys or endpoints and the tokens are billed by your provider, with nothing added by us. Or use managed inference: the gateway serves the model on its own capacity and you pay per token, at the rate above, from credit you add.
One routed request, itemized
What a routed request on managed inference costs, line by line: the decision, the tokens, and a guardrail check when your policy runs one. An example in round numbers; every receipt prints the live ones.
On managed inference a routed request has two lines: the router's decision, then the tokens the chosen provider served. A guardrail policy can add a third. Every line is itemized on each request's receipt and in the monthly usage.
Decisions
$1 per 1,000 routing decisions, the first 5,000 each month free. A request that names a model skips the router and is never charged one.
Tokens
The provider's per-token price plus a flat 6%, read live from the provider and printed on every receipt. No minimum, no idle charge.
Guardrails
Off unless your policy turns them on. An unsafe-content check is a guard-model call, billed per token plus the same 6%. The injection check is free.
Your own provider
On your own keys or endpoints the tokens are billed by your provider. Inferbase charges the decision and any guardrail check.
- Routing decision1 of the 5,000 free each month, then $1 per 1,000
- $0.001
- Tokens in1,000 at $0.20 per 1M, the provider's price
- $0.0002
- Tokens out500 at $0.60 per 1M, the provider's price
- $0.0003
- Guardrail checkonly if your policy turns one on: 250 guard-model tokens at $0.20 per 1M
- $0.00005
- Inferbasea flat 6% on the tokens and the check
- $0.000033
- Total
- $0.001583
- while the decision is inside the free 5,000
- $0.000583
The rates are examples; every receipt prints the live ones. On your own provider key the token lines are billed by the provider; a guardrail check is still charged, because the guard model runs on ours. A pinned model is never charged a decision.
Compare plans
Every control ships on every plan: pools, presets, failover, the decision audit, budgets. The plans differ on price, limits and support.
| Features | Free | Pay-as-you-go | Enterprise |
|---|---|---|---|
| Smart routing | |||
| Automatic model selection | |||
| Free routing decisions | 5,000 / mo | 5,000 / mo | Custom |
| Routing fee beyond the free allowance | $0.001 / decision | Volume pricing | |
| Custom model pools | |||
| Eligibility presets | |||
| Automatic fallback and retries | |||
| Routing decision audit trail | |||
| Inference | |||
| OpenAI-compatible endpoint | |||
| Managed inference price | At cost + 6% | At cost + 6% | Negotiated |
| Streaming responses | |||
| $5 signup credit | Custom | ||
| Minimum top-up | $10 | Invoicing | |
| Platform | |||
| Model catalog and comparison | |||
| Playground | |||
| Organizations and projects | |||
| Role-based access control | |||
| Scoped API keys | |||
| Usage dashboard | |||
| Spend caps and budgets | |||
| GPU capacity planner | |||
| Scale and support | |||
| Rate limit per key | 60 rpm default | Up to 600 rpm | Custom |
| Support | Community | Community | Dedicated |
| SSO | By contract | ||
| Uptime SLA | By contract | ||
| Dedicated capacity | By contract | ||
| Custom model onboarding | |||
| Compliance support | By contract | ||
Rate limits are per API key and can be set lower on any key. Enterprise terms such as SSO, an uptime SLA, dedicated capacity and compliance support are arranged by contract. Contact us for higher limits or a custom plan.
Frequently asked questions
How routing is priced, the two ways to run models, what happens when credit runs out, and what is arranged by contract.
No subscription. Smart routing is $1 per 1,000 routing decisions, with the first 5,000 each month free and pinned models never charged a routing fee. For the models, bring your own provider keys or endpoints and pay your provider directly, or use managed inference at the model host's token price plus a flat 6%. New accounts start with $5 of free inference credit after email verification.
A request is refused before any model is called, with a 402 and a message that names the reason, so nothing runs unpaid and nothing is billed without credit behind it. Add credit under Settings and Billing and requests resume. A project that reaches its monthly budget is refused the same way until the budget is raised or the month rolls over.
No. Inferbase is pay-as-you-go. You add credit when you want it and pay only for what you use. There are no seats and no recurring platform fee.
$10. New accounts also receive $5 of free inference credit after verifying their email, so you can try routing and inference before adding anything.
Both, on every plan. A new account receives $5 of inference credit after verifying its email, with no card, and the first 5,000 routing decisions each month are free whether or not you have added credit. The free tier is not time-limited; it is the allowance you keep as usage grows.
Per token, as the model provider bills it: prompt tokens and completion tokens at the provider's published price, plus a flat 6% on top. Each request's tokens and price are on its receipt, and the monthly usage view totals them by key and project. Credit is prepaid; nothing is billed after the fact.
Yes. Connect a provider key or an endpoint to a project and requests to those models run on your account, billed by your provider with nothing added by Inferbase. The Inferbase charges on such a request are the routing decision, the first 5,000 a month free, and a guardrail check if your policy runs one.
Limits are per API key, set on the key within the plan's ceiling: 60 requests per minute by default on Free, up to 600 on Pay-as-you-go, and custom on Enterprise. A request over the limit is refused with a 429 and can be retried; a project that reaches its monthly budget is refused with a 402 until the budget is raised.
Yes, on every plan. Organizations hold projects, projects hold keys and a monthly budget, and members carry roles (owner, admin, billing, member, viewer) that gate what they can change and see. Enterprise terms such as SSO, an uptime SLA, dedicated capacity and compliance support are arranged by contract; contact us and we will work out the details.
Put intelligence in the middle.
One gateway in front of every model, with your policies applied and every decision on record. Start with $5 of credit and 5,000 routing decisions a month, no card required.