Pricing

Whether you are shipping an app or running long character chat sessions. One plan, every model.

Pay as you go

Add credit, use any model.

$5

500calls

  • Pay as you go credit never expires.
  • Add credit, use any model.

Popular

$20

2,000calls

  • Pay as you go credit never expires.
  • Add credit, use any model.

$100

10,000calls

  • Pay as you go credit never expires.
  • Add credit, use any model.
How does Inferio pricing work?

You pay per token consumed. Each model has its own input, output and cache rate, billed from your account credit. There is no monthly subscription required for Pay As You Go.

Do I need a subscription?

No. Pay As You Go is fully self-serve, top up once and consume any model. Subscriptions give better value (up to 1.75x credit multiplier) if your usage is predictable.

Is the API compatible with OpenAI?

Yes. Inferio exposes an OpenAI-compatible endpoint for chat completions, plus Anthropic and Gemini compatible endpoints. Point any OpenAI SDK at our base URL and use your Inferio API key.

Which models are included?

Every plan includes access to all supported models (Claude, GPT, Gemini, DeepSeek and more). We do not gate models by tier.

What happens if a provider fails?

Inferio routes requests across multiple providers with automatic failover. If one provider is rate-limited or down, we retry against the next healthy upstream transparently.

Can I get a refund?

Unused prepaid credit is refundable for 30 days from purchase. Contact [email protected] for refund requests.

Can I use this with SillyTavern, RisuAI, or other OpenAI-compatible AI chat clients?

Yes. Every plan works as a custom OpenAI-compatible source in any major character chat client. SillyTavern, Janitor.AI, RisuAI, and Chub each have a setup guide in the docs.

How long can my conversations or character history be?

Inferio supports the full context window of every model, up to 1M tokens on Gemini and 200k on Claude. Long character conversations and long codebases both work.

[email protected]