Overview
Latency and throughput are two orthogonal measures of system performance: latency is the time a single request takes to complete, while throughput is the volume of work a system finishes per unit of time. The distinction matters because architectures optimized for one can quietly degrade the other.
Comparison Diagram
Comparison Table
| Aspect | Latency | Throughput |
|---|---|---|
| Definition | Time elapsed for one request to travel and complete | Amount of work completed across all requests per unit time |
| What is measured | A single request’s round trip or processing delay | Aggregate output of the system over an observation window |
| Unit of measurement | Milliseconds, microseconds, or seconds | Requests/sec, transactions/sec, or Mbps |
| Primary driver | Network round-trip time, serialization, and processing delay | Available bandwidth, parallel capacity, and resource pool size |
| Effect of concurrency | Individual request latency can rise as queueing builds up | Throughput rises with more parallel workers, up to a capacity limit |
| Behavior under overload | Tail latency spikes as queues grow (p95/p99 degrade) | Throughput plateaus or drops once the system saturates |
| Typical optimization | Reduce round trips, cache results, shorten the critical path | Batch requests, add parallel workers, scale out capacity |
| Measurement method | Ping, request timers, percentile latency (p50/p95/p99) | Requests-per-second counters, load testing, capacity benchmarks |
Key Differences
- Latency measures the time for one request; throughput measures the volume processed per unit time.
- Batching to raise throughput can increase tail latency for individual requests.
- Latency is bounded by physical round-trip time; throughput is bounded by system capacity.
- Under heavy load, latency spikes from queueing while throughput plateaus at a ceiling.
- Little’s Law links the two: average latency times concurrency roughly equals throughput.
When to Use Each
Latency
- Real-time interactive systems: Gaming and video calls need low per-action latency for the interaction to feel responsive.
- High-frequency trading: Microsecond-level latency directly determines execution price and profitability.
- User-facing API responses: Perceived application responsiveness is driven by how fast a single request returns.
Throughput
- Batch ETL pipelines: Success is measured by total records processed per hour, not any single record’s speed.
- Bulk data transfer: Large backups or file syncs care about total bytes moved per second, not per-packet delay.
- Log and event ingestion: Analytics pipelines need to sustain high sustained volume across many producers.