Model peak demand as a time-based arrival profile and concurrency as the work actually in flight; do not treat either as a single average. Microsoft warns that per-call latency variation can keep achieved throughput below quota, so capacity plans must connect incoming demand with overlap, latency, completion time, and completed work.
Separate demand, concurrency, and throughput
| Measure | What the model should represent |
|---|---|
| Peak demand | When work arrives, how much arrives, and which workload types account for it |
| Concurrency | How many calls or jobs are in flight at the same time |
| Latency | How long each call or operation takes |
| Completion time | How long work remains unfinished from its start to its final completion |
| Achieved throughput | How much work actually finishes over time |
An average arrival rate can hide a short burst, while an average latency can hide long-running calls that overlap with later demand. Model each workload type separately if their arrival patterns or durations differ materially.
If queues, dependent tool calls, or other workflow stages exist, define whether they count as active work, queued work, or both. That prevents the concurrency figure from depending on an ambiguous operational definition.
Build the peak from overlapping work
Keep timestamps for work starting and finishing. From those records:
- Reconstruct the observed peak arrival pattern.
- Group calls or jobs by workload type.
- Measure the distribution of call durations, not only the mean.
- Count overlapping in-flight work on the same timeline.
- Compare planned concurrency with the capacity constraints that apply to that workload.
A useful average sanity check is:
average concurrency ≈ arrival rate × average time in system
However, that relationship does not reproduce bursts, queues, or long-duration work. Peak planning still requires the timeline and the observed duration distribution.
For example, a rise in call duration can increase overlap even if the arrival rate does not change. More calls may then be in flight without a corresponding increase in completed work. This is why concurrency should not be derived from demand or quota alone.
Instrument the workload before setting assumptions
Microsoft recommends observing throughput, latency, and completion times, primarily through application instrumentation. The resulting dataset should support at least these views:
- Demand by time and workload type
- Active and queued concurrency
- Start-to-finish duration for each call or job
- Completed work over time
- Failures, retries, and cancellations, where applicable
- Provider-reported rate-limit information
Instrumentation also makes it possible to distinguish a demand surge from a latency-driven capacity problem. Without start, finish, and concurrency records, teams may see lower throughput without knowing whether the cause is fewer arrivals, slower calls, longer queues, or some combination.
Check the model against quota signals
For Azure OpenAI, Microsoft documents rate-limit information in the HTTP response headers of every API call. The relevant headers expose remaining request and token allowances.
Record those allowances alongside timestamps, request types, and workload completion data. This creates a comparison between:
- The demand the application is trying to generate
- The concurrency required to sustain that demand
- The request and token allowances reported by the service
- The work the application actually completes
Do not treat quota as proof of achieved throughput. Microsoft specifically warns that per-call latency variations may prevent throughput from reaching the quota. A model should therefore test both the volume constraint and the overlap created by longer calls.
For services with equivalent controls, verify the applicable fields and semantics rather than assuming another provider uses the same headers or limits.
Validate with the observed peak, then vary latency
Replay the observed peak arrival pattern in a controlled test and collect the same instrumentation used in production. Compare the model with measured concurrency, latency, completion times, and throughput.
Then vary the duration assumptions while checking whether concurrency and completed work respond as predicted. This is particularly important when the workload has a wide range of call durations: an apparently manageable arrival rate can produce substantial overlap when calls remain active longer than expected.
Apply any operational buffer explicitly and document its basis. There is no universal headroom value that can be selected safely without workload evidence and the applicable capacity constraints.
What operators must still confirm
Before approving a capacity plan, operators should separately verify:
- Current provider-, model-, and workload-specific quotas
- The meaning and applicability of the rate-limit headers
- The measured arrival pattern and duration distribution
- Whether dependent or queued work belongs in the concurrency definition
- Current pricing, billing conditions, and contractual capacity terms
- Any required reliability or service-level commitments
No numeric quota, latency target, price, refund term, or guaranteed throughput is established by this guide. Those values require current official documentation and, where applicable, the governing agreement.