Teams should plan capacity for model changes by sizing system-level throughput and per-call latency separately. Historical prompt and completion token metrics provide the basis for the throughput estimate; per-call response times then need an independent check against the workload’s own criterion. A review should not let one result stand in for the other.
Start with the sizing distinction
The cited guidance treats throughput and latency as separate concepts:
| Sizing question | Measure | What the team is checking |
|---|---|---|
| How much token volume must the system handle? | System-level throughput, measured in tokens per minute (TPM) | The system’s volume requirement |
| How long does an individual call take? | Per-call response time, also called latency | The response time for each call |
A TPM estimate describes the amount of token processing required over time. It does not, on its own, show that individual calls meet a response-time criterion. A latency measurement likewise does not establish that the system can handle the required volume, so the two results belong in separate records.
Estimate throughput from historical usage
The guidance identifies historical prompt and completion token metrics as the metrics needed to estimate system-level throughput. Teams should keep those measures visible in the review:
- Prompt token metrics describe the token pattern on the request side.
- Completion token metrics describe the token pattern on the response side.
When a model changes, compare the historical pattern with the workload expected after the change. If that comparison is not verified, record the gap as an assumption rather than substituting a guessed value. The cited guidance does not provide a universal conversion formula or a model-specific capacity ceiling, so the estimate should state its assumptions explicitly.
A practical review sequence is:
- Capture the historical prompt and completion token metrics.
- State the resulting TPM estimate and the assumptions behind it.
- Revisit the estimate when post-change measurements become available.
- Keep the latency result in a separate record.
Check latency independently
Run calls representative of the intended workload after the model change and record per-call response times. Compare the observations with a response-time criterion and measurement method that the team has explicitly defined. If either is missing, the latency side remains unverified; a favorable TPM estimate cannot close that gap.
Confirm the unresolved items
The cited guidance supports the distinction and the historical-metric method, but it does not establish model-specific capacity ceilings, latency targets, fees, deadlines, regional availability, contractual commitments, or rollout conditions. Those items must be confirmed separately before a capacity decision is finalized.
Before sign-off, teams should verify:
- The limits that apply to the proposed model and configuration.
- The workload’s peak volume and how it is represented in the TPM estimate.
- The response-time criterion and measurement method.
- Whether historical prompt and completion token metrics remain representative after the change.
- Applicable cost, schedule, reliability, and contractual assumptions.
Until those items are verified, the throughput estimate and latency check should remain separate open workstreams. This creates a documented planning position without turning an estimate into a performance guarantee.