Request routing affects AI cost control because it can change which pricing context applies to a request. The relevant rate is tied to the subscription, region, deployment type, and billing agreement; using one undifferentiated rate can therefore hide differences between routes. Routing affects the inputs to the estimate, not the need to account for both input and output token usage.
How to check it
- Identify the route context. Record the subscription, region, deployment type, and billing agreement associated with the request.
- Use the applicable rate. Select the rate that applies to that context rather than carrying over a rate from a different context. The cited guidance explicitly lists these dimensions.
- Calculate both token components. The stated approximation is:
approximate model cost = ((input tokens × input price) + (output tokens × output price)) / 1,000,000
Multiply the input tokens by the input price and the output tokens by the output price, add the results, and divide by 1,000,000. - Keep comparisons aligned. Retain the route context, token counts, and prices with the estimate. A lower or higher calculated amount is not a like-for-like comparison if the pricing context is different.
What the reader must still confirm
The formula alone does not establish a route-specific amount. The cited guidance does not establish a universal numerical rate, so before relying on an estimate, the reader still needs to confirm:
- the applicable rate for the exact subscription, region, deployment type, and billing agreement;
- the input and output prices being used;
- the input and output token counts; and
- that the selected route and billing context match the request being evaluated.
The guidance provides a method for an approximate model-cost calculation, not a guaranteed total or a promised saving. The final amount remains dependent on the route-specific rate and usage inputs.