Pricing
Prices are generated from the live catalog — this table is what the platform charges, to the token, with nothing else on the bill.
How pricing works
- Per-token, per-model. Input tokens, output tokens, and cached-input tokens bill at the three rates in the table above — no per-request fee, no minimum, no monthly line item.
- Cache reads bill at the cached rate. Prompt-prefix cache hits are reported in
usage.prompt_tokens_details.cached_tokensand billed at the cached column, not the input column. - Thinking tokens bill as output. On models that emit reasoning (for example the MiniMax family), hidden thinking tokens are billed as output tokens and are visible in
usage.completion_tokens_details.reasoning_tokens— the number you are charged for is the number you can audit. - Vision is available on
qwen3.8-27bandmuse-glimmer-30bat the same token rates (image parts count as input tokens).
Variant aliases (a -fast tier, for instance) appear on this table the moment they are listed; unlisted variants are never served a price row they were not promised.
What the platform commits to
Three of our beta commitments are price-relevant, stated here in those words:
- Prepaid wallet = hard cap. Your wallet cannot go negative — runaway billing is structurally impossible.
- Price you see is price you pay — no credit-load fees. Credits are charged at face value.
- No mid-term changes to purchased credits. Bought credits keep their terms.
All five commitments live on the beta commitments page.