Load Shedding vs Request Completion: Rejecting Early vs Finishing In-Flight Work

Overview When a service is overloaded, it must choose between two competing policies: load shedding, which rejects excess requests at the door before they consume resources, and request completion, which guarantees every admitted request runs to its natural end. The choice determines whether overload shows up as explicit client-visible rejections or as growing queues and degraded latency for everyone still being served. Comparison Diagram Load SheddingRequest CompletiongateAcceptedshed (503)Rejects excess before admissionadmittedin-flightcompleteEvery admitted request runs to completion Comparison Table Aspect Load Shedding Request Completion Lifecycle stage Applied at admission, before a request enters processing Applied after admission, to work already in flight Trigger Fires when queue depth, CPU, or latency crosses an overload threshold Is the default behavior for any request that was accepted, regardless of load Resource cost of the decision Cheap — rejects with a fast, minimal-work response Expensive — the request’s resources are already committed and must be paid out Client-facing outcome Explicit rejection (e.g. HTTP 503), client must retry later Eventual success or failure on the request’s own merits, no artificial cutoff Effect on accepted traffic Protects tail latency for accepted requests by removing excess load Risks rising tail latency and queueing as accepted work competes for the same resources Prioritization Can selectively drop low-priority or cheap-to-reject traffic first Typically processes admitted work FIFO, with no mid-flight reordering Failure mode if misapplied Too aggressive shedding rejects healthy capacity and wastes headroom No shedding at all leads to resource exhaustion and cascading failure Implementation layer Load balancer, API gateway, or admission-control middleware Service handler or business logic that owns the request once accepted Key Differences Load shedding acts at the front door, before admission; completion policy governs work already in-flight. Shedding trades a guaranteed rejection for protecting tail latency of everything else being served. Completing every accepted request avoids wasting sunk cost already spent, but risks resource exhaustion under sustained overload. Shedding can prioritize which traffic to drop, while completion is typically FIFO once work is admitted. Over-aggressive shedding causes false rejections; refusing to shed at all invites cascading failure. When to Use Each Load Shedding ...

September 6, 2026 · 3 min · 457 words · jeonck

Deep Queues vs Backpressure: Absorbing Load vs Signaling It Back

Overview Both are strategies for handling load spikes between a fast producer and a slower consumer, but they differ in where the excess work goes. Deep queues buffer the overflow in memory or disk so the producer never has to slow down, while backpressure pushes a signal upstream so the producer itself throttles before the system gets overwhelmed. The choice determines whether your system trades memory and latency for decoupling, or throughput for bounded stability. ...

September 6, 2026 · 3 min · 440 words · jeonck