Overview
Batch processing collects data over a period and runs computation on the whole bounded set at once, while stream processing handles each event as it arrives, continuously. The choice determines whether your system optimizes for throughput and simplicity or low latency on fresh results.
Comparison Diagram
Comparison Table
| Aspect | Batch Processing | Stream Processing |
|---|---|---|
| Data ingestion | Data collected and stored until job triggers | Events consumed individually as they arrive |
| Data scope per run | Bounded, finite dataset (a file, a partition, a day’s data) | Unbounded, continuous sequence of events |
| Processing trigger | Scheduled interval or manual kickoff (hourly, nightly) | Continuous, triggered by each event or micro-window |
| Latency to result | Minutes to hours, depending on schedule | Milliseconds to seconds |
| State management | Recomputed fresh from full dataset each run | Maintained incrementally across the event stream |
| Fault recovery | Rerun the failed job against the same input | Checkpointing and replay from an offset in the log |
| Ordering guarantees | Whole dataset available, so ordering enforced within the job | Ordering must be explicitly handled (per-key, watermarks) |
| Resource usage pattern | Spiky: idle, then a burst of compute at run time | Steady, sustained compute and memory footprint |
Key Differences
- Batch operates on a bounded dataset, stream operates on an unbounded sequence of events
- Batch trades latency for simplicity and throughput; stream trades complexity for freshness
- Stream systems need watermarks to handle late or out-of-order events, which batch avoids entirely
- Failure recovery in batch means rerunning the job; stream relies on checkpoint and replay semantics
- Batch pipelines are easier to reason about and test since input is fixed and reproducible
When to Use Each
Batch Processing
- Nightly reporting: Aggregating a full day’s transactions into summary tables is naturally bounded and doesn’t need sub-second results.
- Large-scale ETL: Reprocessing historical data for schema migrations or backfills benefits from batch’s simplicity and reproducibility.
- Cost-sensitive workloads: Running compute in scheduled bursts on cheap spot capacity is cheaper than keeping a stream cluster always on.
Stream Processing
- Fraud detection: Flagging a suspicious transaction needs a decision within seconds, not after the next batch window.
- Real-time dashboards: Live metrics like active users or system health require continuously updated aggregates.
- Event-driven microservices: Reacting to user actions or IoT sensor data as it happens keeps downstream systems in sync with minimal delay.