Cold Start vs Warm Start: Why the First Request Feels Slower
Overview In serverless and containerized systems, a cold start happens when a request must wait for a new execution environment to be provisioned and initialized before it can run, while a warm start reuses an already-running instance and skips straight to execution. The gap between the two explains why the same function can respond in 5ms or 2 seconds depending on whether an idle instance was standing by. Comparison Diagram Cold StartReqProvisioncontainerInit runtime+ codeExecutelatency: tens of ms - several secondsWarm Startidle, pre-initialized container waitingReqExecutelatency: sub-ms - low tens of ms Comparison Table Aspect Cold Start Warm Start Trigger condition No idle instance available (scale-to-zero, scale-out, or fresh deploy) Idle, already-initialized instance is available to handle the request Environment state at invocation No running process; container or sandbox must be created from scratch Process is already running in memory from a prior invocation Steps performed Provision compute, load code, initialize runtime and dependencies, run init code, then handle request Skip provisioning and init; execute the handler directly on the existing process Typical latency added Tens of milliseconds to several seconds depending on runtime and package size Sub-millisecond to low tens of milliseconds Resource cost to provider Higher; allocates new compute, memory, and network setup Lower; reuses resources already allocated Frequency of occurrence Rare relative to total traffic but concentrated after idle periods, deploys, or scale-out Common; most requests during steady, active traffic Primary mitigation Provisioned concurrency, smaller packages, lighter runtimes, scheduled pings Sustained traffic, minimum instance counts, connection reuse Key Differences Cold start pays full provisioning overhead; warm start reuses an already-initialized process. The latency gap can span orders of magnitude — low milliseconds versus multiple seconds. Cold starts are triggered by scale-to-zero or scale-out events, not by what the request contains. Whether a start is warm depends on the platform’s idle timeout before it reclaims the instance. Avoiding cold starts usually means paying for reserved capacity to keep instances standing by. When to Use Each Cold Start ...