Load Shedding vs Request Completion: Rejecting Early vs Finishing In-Flight Work

Overview When a service is overloaded, it must choose between two competing policies: load shedding, which rejects excess requests at the door before they consume resources, and request completion, which guarantees every admitted request runs to its natural end. The choice determines whether overload shows up as explicit client-visible rejections or as growing queues and degraded latency for everyone still being served. Comparison Diagram Load SheddingRequest CompletiongateAcceptedshed (503)Rejects excess before admissionadmittedin-flightcompleteEvery admitted request runs to completion Comparison Table Aspect Load Shedding Request Completion Lifecycle stage Applied at admission, before a request enters processing Applied after admission, to work already in flight Trigger Fires when queue depth, CPU, or latency crosses an overload threshold Is the default behavior for any request that was accepted, regardless of load Resource cost of the decision Cheap — rejects with a fast, minimal-work response Expensive — the request’s resources are already committed and must be paid out Client-facing outcome Explicit rejection (e.g. HTTP 503), client must retry later Eventual success or failure on the request’s own merits, no artificial cutoff Effect on accepted traffic Protects tail latency for accepted requests by removing excess load Risks rising tail latency and queueing as accepted work competes for the same resources Prioritization Can selectively drop low-priority or cheap-to-reject traffic first Typically processes admitted work FIFO, with no mid-flight reordering Failure mode if misapplied Too aggressive shedding rejects healthy capacity and wastes headroom No shedding at all leads to resource exhaustion and cascading failure Implementation layer Load balancer, API gateway, or admission-control middleware Service handler or business logic that owns the request once accepted Key Differences Load shedding acts at the front door, before admission; completion policy governs work already in-flight. Shedding trades a guaranteed rejection for protecting tail latency of everything else being served. Completing every accepted request avoids wasting sunk cost already spent, but risks resource exhaustion under sustained overload. Shedding can prioritize which traffic to drop, while completion is typically FIFO once work is admitted. Over-aggressive shedding causes false rejections; refusing to shed at all invites cascading failure. When to Use Each Load Shedding ...

September 6, 2026 · 3 min · 457 words · jeonck

Idempotency vs Simplicity: Safe Retries vs Minimal Design

Overview This compares two competing goals when designing an operation or API: making it safe to repeat (idempotency) versus keeping it easy to build and reason about (simplicity). The tension matters because guarding against duplicate execution almost always adds state and logic that a minimal implementation would otherwise skip. Comparison Diagram IdempotencySimplicityClientreq #1retryServerkey store: A seenexecuted onceretries collapse to one resultClientreq #1retryServerexecutedexecuted againretries run twice, no dedupno key store, fewer moving parts Comparison Table Aspect Idempotency Simplicity Design intent Guarantee repeated execution has the same effect as one execution Minimize the number of moving parts and decisions in the implementation Handling duplicate requests Detects and ignores repeats using an idempotency key or natural key Processes each incoming request as new, with no duplicate detection State required Needs a dedup store (key, result, TTL) to remember prior executions Stateless with respect to prior calls, nothing extra to persist Behavior on client retry Safe to retry any number of times; result is unchanged Retry re-runs the operation, risking duplicate side effects Failure recovery Callers can blindly retry after timeouts without side-effect risk Callers must add their own checks before retrying after a failure Implementation cost Extra code for key generation, storage, locking, and expiry Fewer edge cases, less code, faster to build and review Testing burden Must cover concurrent duplicates, race conditions, and key expiry Test surface limited to the core logic path, no dedup scenarios Best-fit workloads Payments, distributed queues, webhooks, multi-step workflows Internal read-only endpoints, prototypes, low-stakes single-writer ops Key Differences Idempotency trades extra state for safety, simplicity trades safety for fewer parts Idempotent operations rely on a dedup key that a simple implementation has no reason to store Simplicity pushes retry-safety responsibility onto the caller instead of the server Idempotency adds testing surface for concurrency and expiry that simple code avoids entirely The right choice depends on whether duplicate side effects are tolerable for the operation When to Use Each Idempotency ...

September 6, 2026 · 3 min · 439 words · jeonck

Retries vs Load Amplification: Resilience Tactic vs Its Failure Mode

Overview Retries mask transient failures by having a client re-attempt a request that timed out or errored, trading latency for reliability. Load amplification is what happens when those same retries compound: a struggling downstream service receives multiplied traffic from many clients retrying at once, turning a partial slowdown into a full outage. The design challenge is keeping the first from causing the second. Comparison Diagram RetriesLoad AmplificationClientServicereq 1 (timeout)retry (backoff+jitter)1 extra request, delayedC1C2C3Service6x request volume, queue growing Comparison Table Aspect Retries Load Amplification Trigger A single failed or timed-out request at the client Many clients (or one client’s retries) hitting an already degraded service Nature Deliberate resilience mechanism Emergent side effect of that mechanism under stress Scope Per-request, client-local decision Fleet-wide, system-level consequence Timing pattern Delayed re-attempt, ideally with exponential backoff and jitter Requests pile up faster than the service can drain them Effect on downstream load Small, bounded increase of one extra attempt Multiplicative increase, often several times baseline traffic Worst-case outcome Slightly higher latency for the caller Cascading failure or full outage from a retry storm Mitigation Retry budgets, idempotency keys, capped attempt counts Circuit breakers, load shedding, backpressure, rate limiting Observability signal Retry count and retry rate per endpoint Request rate vs baseline, queue depth, error rate spike Key Differences Retries are a client-side decision; load amplification is a system-wide consequence that emerges when many retries overlap A single retry adds one extra attempt, but a retry storm can multiply traffic several times over in seconds Exponential backoff with jitter reduces retry-driven amplification by spreading re-attempts over time instead of synchronizing them Amplification is contained with circuit breakers and load shedding at the service, not by removing retries entirely When to Use Each Retries ...

September 6, 2026 · 2 min · 408 words · jeonck

Short vs Long Timeouts: Failing Fast vs Tolerating Slowness

Overview Timeouts define how long a caller waits for a response before giving up, and the duration you pick trades off resource protection against tolerance for slow-but-valid work. Short timeouts fail fast and protect callers from cascading slowness, while long timeouts give operations more room to complete under load or over slow links at the cost of holding resources longer. Comparison Diagram Short TimeoutLong TimeoutCallServerwait window: 200msTimeout errorfails fast, retries quicklyCallServerwait window: 30sResponse arrivestolerates slow workfrees threads/connections soonerrisk: false failures under loadholds resources longer per callrisk: cascading pileup/exhaustion Comparison Table Aspect Short Timeout Long Timeout Request initiation Caller sets an aggressive deadline immediately on send Caller allows a generous window before send returns control Behavior under normal latency Succeeds well within budget, negligible overhead Succeeds with unused slack, no functional difference Behavior under slow dependency Aborts before slow-but-valid work finishes, causing false failures Waits out transient slowness, letting valid work complete Resource holding Frees threads, sockets, and connection pool slots quickly Ties up threads, sockets, and pool slots for the full wait Failure propagation Fails fast, enabling quick retry or fallback logic Delays failure detection, slowing retries and fallback triggers System behavior under overload Sheds load quickly, protecting upstream and downstream services Risks thread/connection exhaustion and cascading backpressure Retry and circuit breaker interaction Pairs well with fast retries and quick breaker tripping Delays breaker tripping, masking degradation until timeout expires Tuning basis Set near p99 latency of a healthy, fast dependency Set to cover legitimate worst-case work like batch jobs or large payloads Key Differences A short timeout favors quick failure detection over completing genuinely slow requests A long timeout risks resource exhaustion when many calls stall simultaneously Short timeouts pair naturally with fast retries, while long timeouts delay circuit breaker activation Choosing either wrong direction turns normal latency variance into either false failures or cascading pileups The right value depends on the dependency’s actual p99 latency, not a guessed constant When to Use Each Short Timeout ...

September 6, 2026 · 3 min · 458 words · jeonck

Liveness Probe vs Readiness Probe: Kubernetes Health Checks Compared

Overview Kubernetes uses liveness and readiness probes to answer two different questions about a running container: is it alive, and is it ready to serve requests. A failed liveness probe triggers a container restart, while a failed readiness probe only affects traffic routing by pulling the pod out of Service endpoints without killing it. Comparison Diagram Liveness ProbeReadiness ProbeContainerkubelet: is it alive?Fails?yesKill & Restart ContainernostaysrunningNo effect on Service trafficContainerkubelet: is it ready?Ready?yesAdded to Service EndpointsnoRemoved(not killed)Service / Load Balancer Comparison Table Aspect Liveness Probe Readiness Probe Core question Is the process still functioning? Is the process ready to accept traffic? Action on failure kubelet kills and restarts the container Container is left running, no restart Effect on Service endpoints None directly; pod may keep receiving traffic until restart completes Pod is removed from Service endpoints, stops receiving traffic Effect on rolling deployments Not consulted for rollout progress Must pass before the pod counts as available and rollout proceeds Restart count impact Increments the container restart count on each failure Never causes a restart Typical checks used Lightweight self-check for hangs or deadlocks Checks dependency health: DB connections, cache warm-up, config load Misconfiguration risk Too-aggressive thresholds cause restart loops (CrashLoopBackOff) Too-aggressive thresholds pull healthy pods out of rotation, cutting capacity Key Differences Liveness failure causes a container restart; readiness failure only removes the pod from Service endpoints. Liveness answers “is it alive,” readiness answers “is it ready for traffic.” Readiness gates rolling deployments; liveness has no say in rollout progress. Using a dependency check as a liveness probe risks a restart loop when the real problem is an external service, not the process. When to Use Each Liveness Probe ...

August 3, 2026 · 3 min · 439 words · jeonck