Token usage should shape cost forecasting by tracking input and output tokens separately and applying each to its corresponding price. Measured counts should replace pre-run assumptions when available, but they still should not be treated as the final billed amount.
The cited cost guidance uses this approximation:
approximate model cost = ((input tokens × input price) + (output tokens × output price)) / 1,000,000
How to check the forecast
- Record expected input and output token counts separately.
- Multiply input tokens by the applicable input price.
- Multiply output tokens by the applicable output price.
- Add both amounts and divide by 1,000,000.
- After execution, substitute measured token counts where available. The guidance describes measured usage as more representative than pre-run assumptions.
Why total tokens alone are not enough
A single total does not show how usage is distributed between the two separately priced components. It can therefore conceal changes in the composition of the cost estimate even when total token usage remains unchanged.
Keeping the counts separate also makes the forecast easier to check: readers can see which usage assumptions correspond to each part of the calculation.
What to confirm before relying on the estimate
The calculated amount is an approximate model cost, not a billing statement. Readers should confirm both the measured input and output token counts and the final billed amount in the applicable billing record.
The cited guidance does not provide a reconciliation rule explaining any difference between the approximation and the final bill. Any such difference must therefore be checked directly rather than inferred from token counts alone.