Overview

When a service is overloaded, it must choose between two competing policies: load shedding, which rejects excess requests at the door before they consume resources, and request completion, which guarantees every admitted request runs to its natural end. The choice determines whether overload shows up as explicit client-visible rejections or as growing queues and degraded latency for everyone still being served.

Comparison Diagram

Load SheddingRequest CompletiongateAcceptedshed (503)Rejects excess before admissionadmittedin-flightcompleteEvery admitted request runs to completion

Comparison Table

AspectLoad SheddingRequest Completion
Lifecycle stageApplied at admission, before a request enters processingApplied after admission, to work already in flight
TriggerFires when queue depth, CPU, or latency crosses an overload thresholdIs the default behavior for any request that was accepted, regardless of load
Resource cost of the decisionCheap — rejects with a fast, minimal-work responseExpensive — the request’s resources are already committed and must be paid out
Client-facing outcomeExplicit rejection (e.g. HTTP 503), client must retry laterEventual success or failure on the request’s own merits, no artificial cutoff
Effect on accepted trafficProtects tail latency for accepted requests by removing excess loadRisks rising tail latency and queueing as accepted work competes for the same resources
PrioritizationCan selectively drop low-priority or cheap-to-reject traffic firstTypically processes admitted work FIFO, with no mid-flight reordering
Failure mode if misappliedToo aggressive shedding rejects healthy capacity and wastes headroomNo shedding at all leads to resource exhaustion and cascading failure
Implementation layerLoad balancer, API gateway, or admission-control middlewareService handler or business logic that owns the request once accepted

Key Differences

  • Load shedding acts at the front door, before admission; completion policy governs work already in-flight.
  • Shedding trades a guaranteed rejection for protecting tail latency of everything else being served.
  • Completing every accepted request avoids wasting sunk cost already spent, but risks resource exhaustion under sustained overload.
  • Shedding can prioritize which traffic to drop, while completion is typically FIFO once work is admitted.
  • Over-aggressive shedding causes false rejections; refusing to shed at all invites cascading failure.

When to Use Each

Load Shedding

  • Sudden traffic spike: Shedding protects core availability by rejecting excess requests during a flash crowd instead of letting everyone queue.
  • Multi-tenant API gateway: Dropping or throttling low-priority tenants preserves SLA compliance for higher-priority ones.
  • Approaching resource exhaustion: Shedding before CPU or memory saturates avoids a total outage caused by one more accepted request.

Request Completion

  • Financial or transactional operations: Partial execution risks inconsistent state, so accepted transactions must run to completion once started.
  • Comfortable capacity headroom: When demand is well within capacity there’s no overload to shed against, so requests simply complete.
  • Already-admitted long-running jobs: Abandoning work mid-way wastes resources already spent, so finishing is cheaper than shedding late.