Batch and real-time demand should appear in one time-based capacity view while remaining separate workload streams. Size system-level throughput, measured in tokens per minute (TPM), separately from per-call response time, also called latency, and then test whether batching requests improves response time rather than assuming it will.
Build two complementary views
| Capacity view | Measure | Question answered |
|---|---|---|
| System-level throughput | Tokens per minute (TPM) | How much token-processing demand must the system handle? |
| Per-call response time | Latency for each call | How long does an individual call take? |
Align the batch and real-time token-demand series to the same time windows, but retain the distinction between them. Do not use total throughput as a substitute for response time: the cited guidance identifies throughput and per-call latency as separate sizing concepts.
Test batching rather than assume an improvement
The cited guidance recommends testing whether batching requests improves response time. Compare a batching configuration with a defined test baseline, keeping batch and real-time calls identifiable in the results.
A throughput result alone does not establish whether response time improved. The test must examine per-call latency directly, and any improvement should apply only to the tested workload and configuration.
What still needs confirmation
The operator must confirm the actual token arrival pattern, peak periods, per-call response times, and the effect of batching in the intended deployment.
The cited guidance does not establish a universal batching ratio, latency target, throughput ceiling, or formula for converting batch and real-time demand into one capacity figure. Those details require workload-specific testing or authoritative deployment documentation. Until such evidence is available, batching remains a hypothesis to test, not a promised improvement.