Sources checked
What belongs in the estimate?
OpenAI, Anthropic, and Google publish separate pricing documentation. The applicable bill can depend on the model, caching, processing mode, tools, and modalities. Use the current provider page and your actual usage records.
For a simple text call: cost = input tokens ÷ 1,000,000 × input rate + output tokens ÷ 1,000,000 × output rate. Use the denominator shown on the provider’s price table; do not mix per-token and per-million rates.
A worked example with invented rates
Assume a hypothetical API charges $2 per million input tokens and $8 per million output tokens. These are teaching numbers, not a quotation of any provider’s current prices.
A call using 5,000 input tokens and 1,000 output tokens costs $0.010 + $0.008 = $0.018 before tools and other fees. At 1,000 identical calls, that is $18.
| Component | Calculation | Cost |
|---|---|---|
| Input | 5,000 ÷ 1,000,000 × $2 | $0.010 |
| Output | 1,000 ÷ 1,000,000 × $8 | $0.008 |
| One call | Input + output | $0.018 |
| 1,000 calls | 1,000 × $0.018 | $18.00 |
How do retries change the budget?
If each completed task needs two calls with the same usage, the example cost doubles to $0.036 per task. Additional search, storage, image, or audio charges must be added separately where applicable.
Do not assume a cache discount on every request. Determine which tokens qualify from the service’s usage reporting and pricing rules.
Track cost per successful task
Choose a reporting period, record total billed usage, and count tasks that passed your defined quality checks. Divide the spend by successful tasks; include the failed attempts in the spend.
- Record model, input, output, and any separately billed usage.
- Count all calls in a task, including retries.
- Add applicable tool and service charges.
- Compare the estimate with the provider’s bill.
- Measure completed-task quality alongside total cost.
Common questions
Is an AI subscription the same as API credit?
Do not assume so. Check the product and API billing terms for the service you use; consumer plans and developer usage can be billed separately.
Does a shorter answer always cost less?
Not necessarily overall. It may reduce output tokens, but an incomplete answer can trigger more calls or more human correction. Measure the whole workflow.
What to remember
Budget for the complete workflow, then verify the estimate against billed usage and successful outcomes.
Sources & further reading
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





