A capacity decision is reviewable when it connects the selected capacity to a defined workload, identifies the supporting evidence and assumptions, and makes unresolved dependencies visible. The record should not present throughput as a fixed, workload-independent figure: the sizing guidance states that throughput per PTU depends on the model and the input/output token mix in a given minute.
What Should the Record Contain?
A practical review record can use the following structure:
| Review field | What to document | Review question |
|---|---|---|
| Decision scope | The workload and capacity choice under review, including the unit and measurement basis used | Can a reviewer identify exactly what is being decided? |
| Model and version | The exact model and version to which the throughput figure applies | Could the figure be confused with a result for another model or version? |
| Workload combination | The conditions represented by the test, estimate, or operational evidence | Are the workload conditions defined rather than described only as “typical”? |
| Input/output token mix | The input and output token proportions used during a given minute | Is the token mix attached to the reported throughput? |
| Throughput evidence | The reported result, whether measured or estimated, and the conditions under which it applies | Is the evidentiary status of the figure clear? |
| Assumptions | Inputs that remain uncertain, including assumptions about workload and token mix | Can assumptions be challenged separately from evidence? |
| Decision rationale | Why the documented capacity was selected and what alternatives were considered | Does the conclusion follow from the recorded evidence? |
| Re-review conditions | Changes in model, version, workload combination, or token mix that would require the decision to be reassessed | Will an outdated decision be identified promptly? |
Why Must the Context Stay Attached to the Number?
The sizing guidance links provisioned throughput to both the model and the input/output token mix. Separate guidance also states that throughput for a given model, version, and workload combination varies across those factors.
That makes the context part of the result rather than optional background. A throughput figure separated from its model, version, workload combination, and token mix cannot safely be treated as interchangeable with a figure recorded under different conditions. It may still be useful evidence, but the record must show what it does—and does not—represent.
How Should Evidence and Assumptions Be Separated?
Measured values, estimated values, and assumptions should remain visibly distinct. A reviewable record can make that distinction clearer by:
- identifying each throughput value as observed or estimated;
- preserving the test or analysis conditions;
- recording the token-mix profile used for that result;
- documenting the source and date of the supporting evidence;
- identifying uncertainty where the workload profile may change; and
- recording which changes would trigger a new review.
This structure does not turn an estimate into evidence. It allows the reviewer to see where the conclusion is strongest and where further validation is still required.
What Must the Reader Confirm Separately?
The two cited references do not establish a mandatory documentation template, approval role, capacity threshold, or guaranteed result. Before accepting a capacity decision as complete, the reviewer still needs to confirm the exact applicable guidance, the actual workload, any relevant internal or contractual review requirements, and whether the recorded assumptions remain current.
The core record is therefore a traceable chain:
workload → model and version → input/output token mix → throughput evidence → assumptions → decision → re-review conditions
A capacity number alone does not provide that chain. The documentation becomes reviewable when another reader can follow the basis of the decision without relying on undocumented context.