Charged on actual token usage; how failed requests are handled.
Chat models are billed on input + output tokens at the prices shown on the Models page. Cache hits, reasoning tokens, audio and images are priced separately and itemised on your bill.
We place a conservative hold when a request starts, then settle against actual usage when it finishes and release the excess immediately. A briefly lower balance mid-request is expected.
The test is whether upstream actually spent compute, not whether you got a 200.