Overview
A local cache stores data in the memory of a single application process, giving the fastest possible reads but no visibility into what other instances hold. A shared cache lives in a separate service that every instance queries over the network, trading a bit of latency for one consistent view of cached data across the whole fleet. The choice shapes how you handle invalidation, scaling, and failure in a multi-instance deployment.
Comparison Diagram
Comparison Table
| Aspect | Local Cache | Shared Cache |
|---|---|---|
| Storage location | In-process heap memory of the running application | External service (e.g. Redis, Memcached) reachable over the network |
| Read/write path | Direct memory access within the process, no serialization | Network round trip plus serialization/deserialization per call |
| Latency | Nanoseconds to low microseconds | Sub-millisecond to a few milliseconds depending on network |
| Consistency across instances | Each instance holds its own copy, values can diverge | Single copy visible identically to every instance |
| Invalidation | Must be broadcast to every instance individually (e.g. pub/sub) | Delete or update once, immediately visible to all readers |
| Capacity | Bounded by each process’s own memory, duplicated per instance | Centralized pool sized independently of app instance count |
| Scaling and failure | Scales automatically with app instances but restarts always cold; no external dependency | Scales as its own tier; an outage or restart affects every consuming instance at once |
Key Differences
- Local caches keep data in the same process memory as the app, so reads never leave the machine.
- Shared caches route every access through a network hop to a separate service.
- Only shared caches guarantee single source consistency across all instances.
- Local caches need explicit fan-out for invalidation; shared caches update once for everyone.
- A shared cache outage is a shared failure, while a local cache failure is isolated to one instance.
When to Use Each
Local Cache
- Hot, tiny lookup tables: Reference data like enum labels or feature flags fits comfortably in memory and benefits from nanosecond access.
- Per-request memoization: Caching a value only needed for the lifetime of a single request avoids any network overhead entirely.
- Minimizing infrastructure: No extra service to deploy, monitor, or pay for when the working set is small and per-instance duplication is acceptable.
Shared Cache
- Session or user state: Data must be visible identically no matter which app instance handles the next request from that user.
- Expensive shared computations: Caching a costly query or API result once avoids redundant work across every instance in the fleet.
- Large or fast-growing datasets: A centralized service can be sized and scaled independently instead of being duplicated in every process’s memory.