AI Bigaibig.org

How Should Fallback Behavior Be Defined for an AI Service Outage?

Fallback behavior should be defined as a workload-specific response to each failure condition, not as a blanket instruction to retry. The cited architecture guidance identifies retries and circuit breakers as mechanisms for handling transient errors such as request throttling, but it does not prescribe a universal outage policy.

Separate Recovery From Fallback

Retries and circuit breakers address different problems. A retry sends another request after a transient failure. A circuit breaker stops requests when repeated calls are unlikely to help. Neither mechanism defines what a user or downstream process should receive if the service remains unavailable, so the fallback policy must answer that question separately.

Decision point What the policy should define
Failure classification Which conditions are transient, persistent, ambiguous, or security-sensitive
Retry eligibility Which requests may be retried and which could create duplicate processing or other side effects
Retry boundaries The permitted retry limit, timing, and stopping conditions
Circuit breaking When requests should stop, how that state is observed, and what happens to waiting work
User-facing response Whether the service returns an explicit failure, an approved degraded response, or another defined outcome
Recovery What evidence is required before normal traffic resumes

An ambiguous failure needs particular care. A request that was rejected before processing may be safe to retry, while a request whose response was lost might already have produced an effect. The fallback policy should therefore account for duplicate-execution risk rather than treating every failed request alike.

For an AI workload, the fallback should also distinguish a current model result from degraded content. Stale, cached, synthetic, or manually supplied information should not be presented as a fresh model response unless that behavior is explicitly approved and detectable.

Check the Policy Before Deployment

An operator should confirm that:

  • Every defined failure condition has a corresponding action.
  • Retry behavior is bounded rather than open-ended.
  • Circuit-breaker state changes are visible to monitoring and operational staff.
  • Requests that may have already taken effect are not handled as ordinary transient failures.
  • The fallback response clearly identifies the loss or limitation of normal service.
  • Security controls remain in force when traffic is stopped or redirected.
  • Tests cover both short-lived failures and sustained unavailability.
  • Recovery tests show how normal processing resumes without creating a second failure.

The cited guidance treats reliability, security, and cost as connected architectural trade-offs. Retry traffic, fallback processing, monitoring, and recovery mechanisms should therefore be evaluated together rather than selected solely to maximize availability.

What Must Still Be Confirmed

The cited material does not specify a retry count, retry interval, circuit-breaker threshold, outage duration, recovery objective, permitted degraded mode, user notice, service commitment, or contractual remedy. Those service-specific values and obligations must be confirmed through the applicable architecture, testing evidence, and governing agreements.

Neither retries nor circuit breakers guarantee uninterrupted service or a particular result. They are components of an outage response; the complete fallback behavior must also define safe failure handling, communication, ownership, and recovery.