Balance & billing
Credit packs, per-token rates, prompt caching and what a request costs.
How billing works
- Access is prepaid. Pick a plan and buy a credit pack in the dashboard:
- Claude (every Claude model, including Fable 5.1 and Opus 5.5) or ChatGPT (every ChatGPT model, including GPT-6 Astra): $50 for $15, $100 for $20, $200 for $30, $400 for $40, $1,000 for $70, $2,000 for $125.
- Claude & ChatGPT (both families on one key): $50 for $20, $100 for $25, $200 for $35, $400 for $50, $1,000 for $80, $2,000 for $140.
- Once your payment confirms, you get an email, and your API key follows by email shortly after. Your order's status is always visible in your dashboard.
- A key only calls the model families in its plan. A request for a model outside the plan is rejected before it is routed.
- Your key's credit is checked before each request runs. When it is used up, requests fail with
insufficient_balance. - Each request's token usage is charged to your key's credit at the rate of the model that served it. See How routing works.
- To top up, buy another pack from your dashboard and we add the credit to your key.
- Payment methods: USDC & USDT, Bitcoin, ETH and 100+ coins.
Rates
USD per 1M tokens. Every model in a tier bills identically. See which models are in each tier on the Models page.
| Tier | Input | 5m cache write | 1h cache write | Cached input | Output |
|---|---|---|---|---|---|
| Frontier | $4.20 | $4.20 | $4.20 | $2.45 | $28.00 |
| Balanced | $1.30 | $1.30 | $1.30 | $1.30 | $6.00 |
| Speed | $1.30 | $1.30 | $1.30 | $1.30 | $6.00 |
Prompt caching
When a request reuses a prompt prefix the model has cached, those input tokens are billed at the cached-input rate instead of the input rate. Writing to the cache is billed at the input rate.
- Claude models cache explicitly: mark breakpoints with
cache_controlon the Messages API, with a 5-minute or 1-hour lifetime. - GPT models cache long, repeated prompt prefixes automatically.
Worked example
A Frontier-tier request with 20,000 input tokens, of which 15,000 come from the cache, and 2,000 output tokens:
| Line | Tokens | Rate / 1M | Cost |
|---|---|---|---|
| Uncached input | 5,000 | $4.20 | $0.02 |
| Cached input | 15,000 | $2.45 | $0.04 |
| Output | 2,000 | $28.00 | $0.06 |
| Total | $0.11 |
Without caching the same request would cost $0.14. Try your own numbers in the cost estimator.