Choreography vs Orchestration: Who Drives the Workflow

Overview Both patterns coordinate a multi-step business process across independent services, but they differ in where the coordination logic lives. In choreography, each service reacts to events and decides its own next move with no central brain; in orchestration, a dedicated controller tells every service what to do and in what order. Comparison Diagram ChoreographyOrchestrationOrder SvcPayment SvcShipping SvceventeventeventOrchestratorPayment SvcOrder SvcShipping Svcno central controllercommands out, responses back Comparison Table Aspect Choreography Orchestration Trigger Any service publishes an event when something happens A client or event calls the orchestrator to start the process Coordination logic Distributed across each service’s event handlers Centralized in one orchestrator component Communication style Asynchronous events broadcast to whoever is listening Explicit commands and replies directed at specific services Step sequencing Emergent from chained event subscriptions Explicitly defined as a workflow or state machine Failure handling Each service listens for failure events and compensates locally Orchestrator detects failure and drives compensating transactions Adding a new step Add a listener; no existing service needs to change Update the orchestrator’s workflow definition Observability Hard to see the full process; requires distributed tracing Process state is visible in one place, easy to audit Coupling Low coupling between services, higher coupling to event schema Services decoupled from each other, but coupled to the orchestrator Key Differences Choreography spreads decision-making across services via events; orchestration centralizes it in a single controller. Choreography scales extensibility easily but makes the overall process hard to trace. Orchestration makes the workflow explicit and easy to audit, at the cost of a single point of coordination. Compensation logic lives in each service under choreography, but is driven centrally under orchestration. Orchestration introduces a dependency on the orchestrator itself as new coupling, even as it decouples the services from each other. When to Use Each Choreography ...

September 6, 2026 · 2 min · 413 words · jeonck

Queue vs Event Log: Consume-Once Delivery vs Replayable Stream

Overview A message queue and an event log both move data from producers to consumers, but they differ in what happens after a message is read. A queue treats delivery as a one-time handoff where each message is consumed once and then removed, while an event log keeps every event in an ordered, replayable sequence that multiple independent readers can consume at their own pace. This distinction drives how each handles multiple consumers, failure recovery, and historical reprocessing. ...

September 6, 2026 · 2 min · 426 words · jeonck

Shared Database vs Database Per Service: Data Ownership in Microservices

Overview This compares two data architecture patterns for microservices: a shared database where multiple services read and write the same schema, versus database per service where each service owns an isolated data store. The choice determines how tightly services are coupled, how transactions and queries span service boundaries, and how independently teams can deploy and scale. Comparison Diagram Shared DatabaseDatabase Per ServiceService AService BService CSharedDBService ADB AService BDB BService CDB CSingle point of coupling & contentionIsolated data, independent scaling Comparison Table Aspect Shared Database Database Per Service Schema ownership One schema shared and often co-owned by multiple teams Each service exclusively owns and evolves its own schema Write path Any service can write directly to shared tables Writes go only through the owning service’s API Cross-service queries Simple SQL joins across tables in one database Requires API calls, data replication, or an aggregation layer Distributed transactions Native ACID transactions across affected tables Needs sagas or eventual consistency to span services Schema migrations Any change risks breaking other services using the table Migrations are local and safe to run independently Independent scaling Database becomes a shared bottleneck under load Each store can be scaled or tuned to its own service’s needs Technology choice All services locked into one database engine Each service can pick the best-fit database technology Failure isolation A database outage or lock contention affects every service An outage is contained to the owning service’s data Key Differences Shared database allows cheap cross-table joins but couples every consuming service to one schema Database per service enforces service autonomy at the cost of needing sagas for cross-service transactions Schema changes in a shared database require coordinating multiple teams, while per-service schemas change independently A shared database creates a single failure domain; per-service databases contain outages to one service Polyglot persistence — choosing different database engines per need — is only possible with database per service When to Use Each Shared Database ...

September 6, 2026 · 3 min · 445 words · jeonck

Last-Write-Wins vs Merging: Resolving Conflicting Writes

Overview When two replicas accept concurrent writes to the same key, a system must reconcile them. Last-Write-Wins picks a single winner by timestamp and discards the rest, while Merging combines both writes into a new value using domain-specific or CRDT logic. Comparison Diagram Last-Write-WinsMergingWrite At=10, val=1Write Bt=12, val=2val = 2(highest timestamp wins)Write A silently discardedWrite At=10, val=1Write Bt=12, val=2{A:1, B:2}(both values combined)No data lost, app may reconcile Comparison Table Aspect Last-Write-Wins Merging Conflict trigger Fires when two writes to the same key arrive with overlapping validity, regardless of content Fires the same way, but treats both writes as valid inputs rather than competitors Resolution mechanism Compares timestamps (or version numbers) and keeps the highest one Applies a merge function, CRDT join, or three-way diff to combine both values Data/metadata required A reliable clock or monotonic counter per write Version vectors, causal history, or a semantically defined merge operation Application involvement None — resolution is automatic and content-agnostic Requires the app or data structure to define what ‘combining’ means Outcome for the losing write Discarded entirely, no trace remains Incorporated into the final merged state, nothing is dropped Consistency guarantee Deterministic convergence, but the winner may be arbitrary relative to causality Deterministic convergence that also respects the semantics of both updates Performance overhead Minimal — a single comparison per conflict Higher — merge logic, extra metadata, and sometimes multi-way comparisons Failure mode Silent data loss under clock skew or concurrent writes at the same timestamp Unresolvable merge conflicts that surface to the application or user Key Differences LWW resolves conflicts purely by comparing timestamps, keeping only one write. Merging combines concurrent writes using a merge function or CRDT join instead of picking a single winner. LWW can cause silent data loss when clocks skew or writes race within the same tick. Merging needs semantic knowledge of the data type to combine values correctly. LWW adds negligible overhead per write; merging trades that simplicity for correctness under concurrency. When to Use Each Last-Write-Wins ...

September 6, 2026 · 3 min · 447 words · jeonck

Primary vs Replica Reads: Strong Consistency vs Read Scaling

Overview In a replicated database, read queries can be routed to the primary node or to one of the read replicas. The choice trades guaranteed data freshness for the ability to scale read throughput and reduce load on the write path. Comparison Diagram ClientPrimaryhandles all writesReplicaread-only copyasync replication (lag)writeread (fresh)read (maybe stale)single node, no lagscales out horizontally Comparison Table Aspect Primary Reads Replica Reads Read target Always the single primary node Any of one or more read replicas Consistency guarantee Read-your-writes, strongly consistent Eventual consistency, may lag behind writes Replication lag exposure None, reads the current write state directly Exposed to lag, from milliseconds to seconds Contention with writes Reads compete with writes for CPU, locks, and I/O Writes on primary don’t directly compete with replica reads Read throughput scaling Bounded by single node capacity Scales horizontally by adding more replicas Latency profile Consistent, no wait for replication to catch up Can be lower if replica is geographically closer, but variable under lag Failover behavior Node failure requires promotion and brief write/read outage Load balancer can reroute to another healthy replica Typical use case Financial transactions, read-after-write flows, admin views Analytics, reporting, public APIs, dashboards Key Differences Primary reads guarantee read-your-writes consistency; replica reads may return stale data due to lag. Replica reads scale horizontally by adding read replicas; primary reads are bottlenecked by a single node. Primary reads compete with write traffic for resources; replica reads isolate read load via replication. Replication lag on replicas ranges from milliseconds to seconds depending on network and write volume. Failover on the primary causes brief unavailability; replica failures are masked by load balancing across peers. When to Use Each Primary Reads ...

September 6, 2026 · 2 min · 391 words · jeonck

Cache Freshness vs Hit Rate: Correctness vs Efficiency

Overview Cache freshness measures whether the data returned by a cache still matches the current state of its source of truth, while hit rate measures how often requests are answered directly from the cache instead of falling through to the origin. The two metrics pull in opposite directions: optimizing for freshness means shorter TTLs and more origin traffic, while optimizing for hit rate means longer TTLs and a higher chance of serving stale data. Tuning a cache well means choosing the right balance point for that specific data’s tolerance for staleness. ...

September 6, 2026 · 3 min · 470 words · jeonck

Local vs Shared Cache: Per-Instance Memory vs Centralized Cache Service

Overview A local cache stores data in the memory of a single application process, giving the fastest possible reads but no visibility into what other instances hold. A shared cache lives in a separate service that every instance queries over the network, trading a bit of latency for one consistent view of cached data across the whole fleet. The choice shapes how you handle invalidation, scaling, and failure in a multi-instance deployment. ...

September 6, 2026 · 3 min · 444 words · jeonck

Range vs Hash Partitioning: Ordered Splits vs Scattered Buckets

Overview Range and hash partitioning are two strategies for splitting a table’s rows across multiple partitions or nodes based on a partition key. Range partitioning assigns rows to contiguous key intervals (like date ranges), preserving order for efficient range scans but risking uneven load. Hash partitioning runs the key through a hash function to scatter rows evenly, trading away ordering for balanced, predictable distribution. Comparison Diagram RANGE PARTITIONINGHASH PARTITIONINGincoming keysincoming keys1-3334-6667-1001-3334-6667-100hash(key)P1P2P3P1P2P3Ordered, contiguous rangesScattered, uniform spreadEasy to extend: add a boundaryCostly to resize: rehash keysRisk: skew on hot rangesRisk: no range pruning Comparison Table Aspect Range Partitioning Hash Partitioning Partition key requirement Needs an orderable key with defined boundaries (dates, IDs) Any key works; only needs to be hashable Row-to-partition mapping Explicit boundary rules assign rows to intervals Hash function output (often mod N) selects the bucket Data distribution Can be skewed if key values aren’t uniformly spread Near-uniform if the hash function distributes well Range/scan queries Prunes to only the partitions covering the range Must fan out and scan every partition Point/equality lookups Requires a boundary search to find the right partition Direct O(1) computation locates the partition Adding or removing partitions Cheap: append or split a boundary at the edge Expensive: reshuffles most existing keys unless using consistent hashing Hotspot behavior Sequential writes (recent dates, auto-increment IDs) pile onto one partition Spreads writes evenly but destroys any physical data locality Key Differences Range partitioning preserves order, letting the query planner prune partitions; hash partitioning optimizes purely for even distribution Growing the partition count is a cheap boundary edit in range partitioning but forces a rehash of most keys in hash partitioning Range schemes are exposed to skew when writes cluster in a narrow key window; hash schemes avoid this at the cost of locality Point lookups under hashing are a direct hash computation, while range lookups need a boundary search through ordered intervals When to Use Each Range Partitioning ...

September 6, 2026 · 3 min · 438 words · jeonck

Sync vs Async Replication: When the Write Actually Commits

Overview Synchronous and asynchronous replication differ in exactly one moment: when the primary tells the client a write succeeded. Sync replication waits for the replica to confirm before acknowledging, while async replication acknowledges immediately and copies the data afterward. That single timing difference cascades into everything else — latency, throughput, and how much data you can lose on failover. Comparison Diagram Sync ReplicationAsync ReplicationClientPrimaryReplica1. write2. replicate3. ack4. commitClient waits for step 3ClientPrimaryReplica1. write2. commit3. replicate laterClient returns at step 2 Comparison Table Aspect Sync Replication Async Replication Write acknowledgment Waits for replica confirmation before committing Commits on primary alone, replicates after Commit latency Includes network round-trip to replica Bound only by primary’s local write Data consistency Replica is always up to date at commit time Replica can lag behind primary momentarily Throughput under load Degrades as replica distance or count grows Unaffected by replica speed or distance Replica or network failure Writes block or fail until replica responds Writes continue uninterrupted on primary Failover data loss None — replica always has the committed write Possible — unreplicated writes are lost Replication lag monitoring Not applicable — lag is structurally zero Critical — must track and alert on lag Key Differences Commit timing is the root difference: sync waits, async doesn’t Sync trades latency for a zero-data-loss guarantee on failover Async trades durability for consistently fast local commits Multi-region setups favor async since round-trip time would make sync commits too slow Async requires active lag monitoring that sync simply doesn’t need When to Use Each Sync Replication ...

September 6, 2026 · 2 min · 362 words · jeonck

Replication vs Sharding: Copying Data vs Splitting Data

Overview Both are techniques for scaling a database beyond a single node, but they solve different problems: replication copies the entire dataset onto multiple nodes to boost availability and read capacity, while sharding splits the dataset into disjoint partitions across nodes to boost storage and write capacity. Large-scale systems typically use both together — sharding for horizontal scale, replication within each shard for durability. Comparison Diagram Replication Sharding Primary A B C D Replica A A B C D Replica B A B C D full dataset, copied to every node Router key lookup Shard 1 keys A-M Shard 2 keys N-Z dataset split into disjoint subsets Comparison Table Aspect Replication Sharding Primary goal Increase availability and read capacity Increase storage and write capacity Data distribution Full dataset copied to every node Dataset split into disjoint partitions across nodes Write path Writes go to primary, then propagate to replicas Writes routed to the single shard owning the key Read path Any replica (or primary) can serve any read Read must be routed to the shard holding the key Node failure impact Data survives since other copies exist That shard’s data becomes unavailable unless also replicated Consistency concern Replication lag between primary and replicas Cross-shard transactions and joins are hard to coordinate Scaling ceiling Bounded by primary’s write throughput Bounded by cross-shard coordination and key hotspots Operational overhead Failover and leader election Shard key design, rebalancing, and resharding Key Differences Replication duplicates the same data everywhere; sharding partitions it so each node holds only a slice Replication scales reads and durability; sharding scales writes and total storage Sharding introduces a routing layer that must know which shard owns a given key Losing a replica is harmless, but losing an unreplicated shard causes real data loss Production systems commonly combine both: shard for scale, replicate each shard for resilience When to Use Each Replication ...

September 6, 2026 · 2 min · 420 words · jeonck