Overview

Latency and throughput are two orthogonal measures of system performance: latency is the time a single request takes to complete, while throughput is the volume of work a system finishes per unit of time. The distinction matters because architectures optimized for one can quietly degrade the other.

Comparison Diagram

LatencyClientServerTime for ONE request to completeThroughputClientServerTotal requests completed per second

Comparison Table

AspectLatencyThroughput
DefinitionTime elapsed for one request to travel and completeAmount of work completed across all requests per unit time
What is measuredA single request’s round trip or processing delayAggregate output of the system over an observation window
Unit of measurementMilliseconds, microseconds, or secondsRequests/sec, transactions/sec, or Mbps
Primary driverNetwork round-trip time, serialization, and processing delayAvailable bandwidth, parallel capacity, and resource pool size
Effect of concurrencyIndividual request latency can rise as queueing builds upThroughput rises with more parallel workers, up to a capacity limit
Behavior under overloadTail latency spikes as queues grow (p95/p99 degrade)Throughput plateaus or drops once the system saturates
Typical optimizationReduce round trips, cache results, shorten the critical pathBatch requests, add parallel workers, scale out capacity
Measurement methodPing, request timers, percentile latency (p50/p95/p99)Requests-per-second counters, load testing, capacity benchmarks

Key Differences

  • Latency measures the time for one request; throughput measures the volume processed per unit time.
  • Batching to raise throughput can increase tail latency for individual requests.
  • Latency is bounded by physical round-trip time; throughput is bounded by system capacity.
  • Under heavy load, latency spikes from queueing while throughput plateaus at a ceiling.
  • Little’s Law links the two: average latency times concurrency roughly equals throughput.

When to Use Each

Latency

  • Real-time interactive systems: Gaming and video calls need low per-action latency for the interaction to feel responsive.
  • High-frequency trading: Microsecond-level latency directly determines execution price and profitability.
  • User-facing API responses: Perceived application responsiveness is driven by how fast a single request returns.

Throughput

  • Batch ETL pipelines: Success is measured by total records processed per hour, not any single record’s speed.
  • Bulk data transfer: Large backups or file syncs care about total bytes moved per second, not per-packet delay.
  • Log and event ingestion: Analytics pipelines need to sustain high sustained volume across many producers.