Models & Pricing
Half the lab list. Packs, not a rate card at checkout.
Model pages list context window, endpoints and capabilities. The live catalog is the source for which ids are on the key.

Prices are 50% of the lab’s published list. Promo and off-peak stickers are ignored.
New here? Getting a key and sending a first request takes a minute in Quickstart
Inferio token prices are half the first-party lab list. Image, video, and audio follow the lab’s per-unit sticker, also at 50% off.
The Pricing page is where you buy packs and read the usage table.
Request counts use OpenCode Go’s published coding mix (tokens in, cache read, and out per typical request) and this formula: cost per request from those tokens at Inferio rates, then budget divided by that cost.
Worked example: GLM-5.3-Flash mix is 700 in / 52,000 cache / 150 out. Zhipu list is $0.15 / $0.03 / $0.50. Inferio is half. Checkout is EGP (USD as a guide). The E£612 / E£1,530 / E£3,060 ($12 / $30 / $60) columns are illustration budgets, not caps.
The live numbers sit on Pricing.
On models whose lab publishes a cached-input line, repeated prompt prefixes bill at that cache rate, then the same 50% cut. We do not invent a cache sticker.
Caching is automatic when the upstream supports it. Long stable system prompts benefit most.
Requests fail over to the next provider group when one is rate limited or down. A model leaves the catalog only when every channel is out.
A model disappearing under load is expected. It returns within minutes once a channel recovers.
Keys that pin provider groups fail over only within their pinned groups, see Group Pinning
To get pinged the moment it comes back, watch it in Notifications