Latency vs Throughput: Response Time vs Processing Volume
Overview Latency and throughput are two orthogonal measures of system performance: latency is the time a single request takes to complete, while throughput is the volume of work a system finishes per unit of time. The distinction matters because architectures optimized for one can quietly degrade the other. Comparison Diagram Latency Client Server Time for ONE request to complete Throughput Client Server Total requests completed per second Comparison Table Aspect Latency Throughput Definition Time elapsed for one request to travel and complete Amount of work completed across all requests per unit time What is measured A single request’s round trip or processing delay Aggregate output of the system over an observation window Unit of measurement Milliseconds, microseconds, or seconds Requests/sec, transactions/sec, or Mbps Primary driver Network round-trip time, serialization, and processing delay Available bandwidth, parallel capacity, and resource pool size Effect of concurrency Individual request latency can rise as queueing builds up Throughput rises with more parallel workers, up to a capacity limit Behavior under overload Tail latency spikes as queues grow (p95/p99 degrade) Throughput plateaus or drops once the system saturates Typical optimization Reduce round trips, cache results, shorten the critical path Batch requests, add parallel workers, scale out capacity Measurement method Ping, request timers, percentile latency (p50/p95/p99) Requests-per-second counters, load testing, capacity benchmarks Key Differences Latency measures the time for one request; throughput measures the volume processed per unit time. Batching to raise throughput can increase tail latency for individual requests. Latency is bounded by physical round-trip time; throughput is bounded by system capacity. Under heavy load, latency spikes from queueing while throughput plateaus at a ceiling. Little’s Law links the two: average latency times concurrency roughly equals throughput. When to Use Each Latency ...