No subscription, no per-seat licenses, no minimum spend, no annual commitment. You fund a balance, and every request draws from it at the model's published rate.
| Model | Input / 1M tokens | Output / 1M tokens | Class |
|---|---|---|---|
| gpt-5.6-sol | $8.00 | $45.00 | Flagship |
| gpt-5.6-terra | $3.00 | $18.00 | Balanced |
| o3 | $3.00 | $12.00 | Reasoning |
These are the prices you pay us, platform fee included. New models are added regularly — the catalog in the console is always the authoritative list, and it is what your requests are priced against.
Nothing here is estimated after the fact. Every step below happens on the request itself.
Before any provider is called, the worst-case cost of the request is reserved against your balance. If your balance does not cover it, the request is refused right there — not after the tokens are spent.
The request runs. Input and output tokens are counted from the provider's own usage report, not estimated from the text.
The reservation is replaced by the real amount and the difference returns to your balance. A request that never reached a provider costs nothing at all.
You are charged for what the provider reported and nothing more. A stream that fails after the first byte is recorded as a failed request, and the reserved amount is settled rather than held.
No. The reservation happens before the provider is called, so your balance is the ceiling. There is no invoice that arrives later.
No. Funded balance stays exactly as it is until you spend it.
Not today. You pay for what you use from the first request, at the rates above.
Amounts are held in micro-dollars — millionths of a dollar — so a request that costs a fraction of a cent is recorded exactly instead of rounding to zero.
Keys carry scopes today. Per-key spend limits are on the roadmap; until then the balance is the shared ceiling for the workspace.