Load Shedding vs Request Completion: Rejecting Early vs Finishing In-Flight Work
Overview When a service is overloaded, it must choose between two competing policies: load shedding, which rejects excess requests at the door before they consume resources, and request completion, which guarantees every admitted request runs to its natural end. The choice determines whether overload shows up as explicit client-visible rejections or as growing queues and degraded latency for everyone still being served. Comparison Diagram Load SheddingRequest CompletiongateAcceptedshed (503)Rejects excess before admissionadmittedin-flightcompleteEvery admitted request runs to completion Comparison Table Aspect Load Shedding Request Completion Lifecycle stage Applied at admission, before a request enters processing Applied after admission, to work already in flight Trigger Fires when queue depth, CPU, or latency crosses an overload threshold Is the default behavior for any request that was accepted, regardless of load Resource cost of the decision Cheap — rejects with a fast, minimal-work response Expensive — the request’s resources are already committed and must be paid out Client-facing outcome Explicit rejection (e.g. HTTP 503), client must retry later Eventual success or failure on the request’s own merits, no artificial cutoff Effect on accepted traffic Protects tail latency for accepted requests by removing excess load Risks rising tail latency and queueing as accepted work competes for the same resources Prioritization Can selectively drop low-priority or cheap-to-reject traffic first Typically processes admitted work FIFO, with no mid-flight reordering Failure mode if misapplied Too aggressive shedding rejects healthy capacity and wastes headroom No shedding at all leads to resource exhaustion and cascading failure Implementation layer Load balancer, API gateway, or admission-control middleware Service handler or business logic that owns the request once accepted Key Differences Load shedding acts at the front door, before admission; completion policy governs work already in-flight. Shedding trades a guaranteed rejection for protecting tail latency of everything else being served. Completing every accepted request avoids wasting sunk cost already spent, but risks resource exhaustion under sustained overload. Shedding can prioritize which traffic to drop, while completion is typically FIFO once work is admitted. Over-aggressive shedding causes false rejections; refusing to shed at all invites cascading failure. When to Use Each Load Shedding ...