An AI capacity forecast should be treated as a chain of testable assumptions, not as a quota figure. The first assumptions to examine are whether historical prompt and completion token metrics still represent expected work and whether per-call latency variation will keep achieved throughput below quota. The cited technical guidance says historical prompt and completion token metrics are needed to estimate system-level throughput and cautions that quota alone does not guarantee that throughput.
Which assumptions need testing?
| Assumption | Why it needs testing | How to check it |
|---|---|---|
| Historical prompt token metrics represent future prompts | The throughput estimate depends on whether past prompt behavior remains relevant. | Compare the historical metrics with the prompt pattern expected during the forecast period. |
| Historical completion token metrics represent future completions | A different completion pattern could invalidate an estimate based on earlier observations. | Compare historical completion metrics with the expected completion behavior and document material differences. |
| The expected workload resembles the historical workload | Changes in workload characteristics can make historical observations poor predictors. | Record which workload the history represents and whether the planned workload has the same relevant characteristics. |
| Available quota will become achieved throughput | Per-call latency variation may prevent throughput from reaching quota. | Keep quota and estimated achieved throughput separate, then compare both with observed results. |
| Per-call latency will remain consistent with the forecast | Higher or variable latency can widen the gap between quota and throughput actually achieved. | Test the estimate against representative latency observations rather than assuming a fixed relationship. |
How to check the forecast
Start with the measurement basis. A system-level throughput estimate should identify the historical prompt and completion token metrics used to produce it. If those observations do not represent the expected workload, the resulting estimate remains conditional rather than validated.
Do not collapse quota into throughput. A forecast that treats quota as the amount of throughput that will necessarily be achieved assumes away the effect of per-call latency variation. Quota and achieved throughput should therefore appear as separate forecast outputs.
Make latency an explicit assumption. The forecast should show what happens to achieved throughput when per-call latency is less favourable than expected. The cited guidance supports testing that relationship, but the statement used here does not establish a universal latency threshold or adjustment factor.
Record unresolved assumptions. Any assumption about historical representativeness, future workload behavior, quota achievability, or latency stability should remain visible in the forecast until supporting evidence is available.
What still needs confirmation
The cited statement alone does not verify how closely the historical metrics match the planned workload, what quota constraints apply in a particular environment, or what conversion relationship will hold between quota, estimated throughput, and achieved throughput.
Each forecast therefore still requires confirmation of:
- the period and workload represented by the historical metrics;
- the expected prompt and completion token behavior;
- the applicable quota;
- per-call latency under representative conditions; and
- whether estimated throughput has been validated against achieved throughput.
No fixed buffer, conversion ratio, or guaranteed result can be inferred from the two documented points alone.