Skip to content
All systems operationalStatus

Balance & billing

Credit packs, per-token rates, prompt caching and what a request costs.

How billing works

  • Access is prepaid. Pick a plan and buy a credit pack in the dashboard:
    • Claude (every Claude model, including Fable 5.1 and Opus 5.5) or ChatGPT (every ChatGPT model, including GPT-6 Astra): $50 for $15, $100 for $20, $200 for $30, $400 for $40, $1,000 for $70, $2,000 for $125.
    • Claude & ChatGPT (both families on one key): $50 for $20, $100 for $25, $200 for $35, $400 for $50, $1,000 for $80, $2,000 for $140.
  • Once your payment confirms, you get an email, and your API key follows by email shortly after. Your order's status is always visible in your dashboard.
  • A key only calls the model families in its plan. A request for a model outside the plan is rejected before it is routed.
  • Your key's credit is checked before each request runs. When it is used up, requests fail with insufficient_balance.
  • Each request's token usage is charged to your key's credit at the rate of the model that served it. See How routing works.
  • To top up, buy another pack from your dashboard and we add the credit to your key.
  • Payment methods: USDC & USDT, Bitcoin, ETH and 100+ coins.

Rates

USD per 1M tokens. Every model in a tier bills identically. See which models are in each tier on the Models page.

TierInput5m cache write1h cache writeCached inputOutput
Frontier$4.20$4.20$4.20$2.45$28.00
Balanced$1.30$1.30$1.30$1.30$6.00
Speed$1.30$1.30$1.30$1.30$6.00

Prompt caching

When a request reuses a prompt prefix the model has cached, those input tokens are billed at the cached-input rate instead of the input rate. Writing to the cache is billed at the input rate.

  • Claude models cache explicitly: mark breakpoints with cache_control on the Messages API, with a 5-minute or 1-hour lifetime.
  • GPT models cache long, repeated prompt prefixes automatically.

Worked example

A Frontier-tier request with 20,000 input tokens, of which 15,000 come from the cache, and 2,000 output tokens:

LineTokensRate / 1MCost
Uncached input5,000$4.20$0.02
Cached input15,000$2.45$0.04
Output2,000$28.00$0.06
Total$0.11

Without caching the same request would cost $0.14. Try your own numbers in the cost estimator.