<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Stream-Processing on IT Comparison</title><link>https://comparison.metacog.co.kr/tags/stream-processing/</link><description>Recent content in Stream-Processing on IT Comparison</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Sun, 06 Sep 2026 10:02:40 +0900</lastBuildDate><atom:link href="https://comparison.metacog.co.kr/tags/stream-processing/index.xml" rel="self" type="application/rss+xml"/><item><title>Batch vs Stream Processing: Bounded Data Dumps vs Continuous Event Flow</title><link>https://comparison.metacog.co.kr/posts/2026-09-06-batch-vs-stream-processing-bounded-data-dumps-vs-continuous/</link><pubDate>Sun, 06 Sep 2026 10:02:40 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-09-06-batch-vs-stream-processing-bounded-data-dumps-vs-continuous/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;Batch processing collects data over a period and runs computation on the whole bounded set at once, while stream processing handles each event as it arrives, continuously. The choice determines whether your system optimizes for &lt;strong class="kw"&gt;throughput and simplicity&lt;/strong&gt; or &lt;strong class="kw"&gt;low latency&lt;/strong&gt; on fresh results.&lt;/p&gt;
&lt;h2 id="comparison-diagram"&gt;Comparison Diagram&lt;/h2&gt;
&lt;div class="compare-diagram"&gt;
&lt;svg viewBox="0 0 640 360" xmlns="http://www.w3.org/2000/svg"&gt;&lt;text x="160" y="36" text-anchor="middle" font-size="18" style="fill:var(--primary)"&gt;Batch&lt;/text&gt;&lt;text x="480" y="36" text-anchor="middle" font-size="18" style="fill:var(--primary)"&gt;Stream&lt;/text&gt;&lt;rect x="40" y="60" width="240" height="90" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="160" y="95" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;Data accumulates&lt;/text&gt;&lt;text x="160" y="115" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;into a bounded set&lt;/text&gt;&lt;circle cx="70" cy="130" r="5" style="fill:var(--compare-a)"/&gt;&lt;circle cx="110" cy="130" r="5" style="fill:var(--compare-a)"/&gt;&lt;circle cx="150" cy="130" r="5" style="fill:var(--compare-a)"/&gt;&lt;circle cx="190" cy="130" r="5" style="fill:var(--compare-a)"/&gt;&lt;circle cx="230" cy="130" r="5" style="fill:var(--compare-a)"/&gt;&lt;path d="M160 150 L160 190" style="stroke:var(--compare-a)" stroke-width="2" marker-end="url(#arrowA)"/&gt;&lt;rect x="80" y="195" width="160" height="55" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="160" y="217" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;Scheduled job&lt;/text&gt;&lt;text x="160" y="235" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;processes all at once&lt;/text&gt;&lt;path d="M160 250 L160 285" style="stroke:var(--compare-a)" stroke-width="2" marker-end="url(#arrowA)"/&gt;&lt;rect x="70" y="290" width="180" height="45" rx="4" style="fill:none;stroke:var(--compare-a)" stroke-width="1.5" stroke-dasharray="4 3"/&gt;&lt;text x="160" y="317" text-anchor="middle" font-size="13" style="fill:var(--secondary)"&gt;Result: high latency&lt;/text&gt;&lt;rect x="400" y="60" width="200" height="280" rx="4" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;text x="500" y="85" text-anchor="middle" font-size="13" style="fill:var(--content)"&gt;Events flow continuously&lt;/text&gt;&lt;circle cx="430" cy="115" r="5" style="fill:var(--compare-b)"/&gt;&lt;path d="M430 115 L490 140" style="stroke:var(--compare-b)" stroke-width="1.5" marker-end="url(#arrowB)"/&gt;&lt;rect x="460" y="130" width="80" height="24" rx="3" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1"/&gt;&lt;text x="500" y="146" text-anchor="middle" font-size="10" style="fill:var(--content)"&gt;process&lt;/text&gt;&lt;circle cx="430" cy="175" r="5" style="fill:var(--compare-b)"/&gt;&lt;path d="M430 175 L490 190" style="stroke:var(--compare-b)" stroke-width="1.5" marker-end="url(#arrowB)"/&gt;&lt;rect x="460" y="180" width="80" height="24" rx="3" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1"/&gt;&lt;text x="500" y="196" text-anchor="middle" font-size="10" style="fill:var(--content)"&gt;process&lt;/text&gt;&lt;circle cx="430" cy="235" r="5" style="fill:var(--compare-b)"/&gt;&lt;path d="M430 235 L490 240" style="stroke:var(--compare-b)" stroke-width="1.5" marker-end="url(#arrowB)"/&gt;&lt;rect x="460" y="230" width="80" height="24" rx="3" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1"/&gt;&lt;text x="500" y="246" text-anchor="middle" font-size="10" style="fill:var(--content)"&gt;process&lt;/text&gt;&lt;circle cx="430" cy="290" r="5" style="fill:var(--compare-b)"/&gt;&lt;path d="M430 290 L490 285" style="stroke:var(--compare-b)" stroke-width="1.5" marker-end="url(#arrowB)"/&gt;&lt;rect x="460" y="278" width="80" height="24" rx="3" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1"/&gt;&lt;text x="500" y="294" text-anchor="middle" font-size="10" style="fill:var(--content)"&gt;process&lt;/text&gt;&lt;text x="500" y="325" text-anchor="middle" font-size="13" style="fill:var(--secondary)"&gt;Result: low latency&lt;/text&gt;&lt;defs&gt;&lt;marker id="arrowA" markerWidth="8" markerHeight="8" refX="4" refY="4" orient="auto"&gt;&lt;path d="M0,0 L8,4 L0,8 Z" style="fill:var(--compare-a)"/&gt;&lt;/marker&gt;&lt;marker id="arrowB" markerWidth="8" markerHeight="8" refX="4" refY="4" orient="auto"&gt;&lt;path d="M0,0 L8,4 L0,8 Z" style="fill:var(--compare-b)"/&gt;&lt;/marker&gt;&lt;/defs&gt;&lt;/svg&gt;
&lt;/div&gt;
&lt;h2 id="comparison-table"&gt;Comparison Table&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Batch Processing&lt;/th&gt;
&lt;th&gt;Stream Processing&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Data ingestion&lt;/td&gt;
&lt;td&gt;Data collected and stored until job triggers&lt;/td&gt;
&lt;td&gt;Events consumed individually as they arrive&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Data scope per run&lt;/td&gt;
&lt;td&gt;Bounded, finite dataset (a file, a partition, a day&amp;rsquo;s data)&lt;/td&gt;
&lt;td&gt;Unbounded, continuous sequence of events&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Processing trigger&lt;/td&gt;
&lt;td&gt;Scheduled interval or manual kickoff (hourly, nightly)&lt;/td&gt;
&lt;td&gt;Continuous, triggered by each event or micro-window&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Latency to result&lt;/td&gt;
&lt;td&gt;Minutes to hours, depending on schedule&lt;/td&gt;
&lt;td&gt;Milliseconds to seconds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;State management&lt;/td&gt;
&lt;td&gt;Recomputed fresh from full dataset each run&lt;/td&gt;
&lt;td&gt;Maintained incrementally across the event stream&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fault recovery&lt;/td&gt;
&lt;td&gt;Rerun the failed job against the same input&lt;/td&gt;
&lt;td&gt;Checkpointing and replay from an offset in the log&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ordering guarantees&lt;/td&gt;
&lt;td&gt;Whole dataset available, so ordering enforced within the job&lt;/td&gt;
&lt;td&gt;Ordering must be explicitly handled (per-key, watermarks)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Resource usage pattern&lt;/td&gt;
&lt;td&gt;Spiky: idle, then a burst of compute at run time&lt;/td&gt;
&lt;td&gt;Steady, sustained compute and memory footprint&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="key-differences"&gt;Key Differences&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Batch operates on a &lt;strong class="kw"&gt;bounded dataset&lt;/strong&gt;, stream operates on an &lt;strong class="kw"&gt;unbounded sequence&lt;/strong&gt; of events&lt;/li&gt;
&lt;li&gt;Batch trades latency for &lt;strong class="kw"&gt;simplicity and throughput&lt;/strong&gt;; stream trades complexity for &lt;strong class="kw"&gt;freshness&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Stream systems need &lt;strong class="kw"&gt;watermarks&lt;/strong&gt; to handle late or out-of-order events, which batch avoids entirely&lt;/li&gt;
&lt;li&gt;Failure recovery in batch means &lt;strong class="kw"&gt;rerunning the job&lt;/strong&gt;; stream relies on &lt;strong class="kw"&gt;checkpoint and replay&lt;/strong&gt; semantics&lt;/li&gt;
&lt;li&gt;Batch pipelines are easier to reason about and test since input is &lt;strong class="kw"&gt;fixed and reproducible&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="when-to-use-each"&gt;When to Use Each&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Batch Processing&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Lambda Architecture vs Kappa Architecture: Batch-Plus-Speed vs Single-Stream Pipelines</title><link>https://comparison.metacog.co.kr/posts/2026-08-04-lambda-architecture-vs-kappa-architecture-batch-plus-speed-v/</link><pubDate>Tue, 04 Aug 2026 05:21:28 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-04-lambda-architecture-vs-kappa-architecture-batch-plus-speed-v/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;Lambda Architecture and Kappa Architecture are two patterns for building big-data pipelines that need both real-time and historical results. &lt;strong class="kw"&gt;Lambda&lt;/strong&gt; runs parallel batch and speed layers that get merged at query time, while &lt;strong class="kw"&gt;Kappa&lt;/strong&gt; pushes everything through a single, replayable stream pipeline. The choice determines how much duplicate logic you maintain and how reprocessing actually works.&lt;/p&gt;
&lt;h2 id="comparison-diagram"&gt;Comparison Diagram&lt;/h2&gt;
&lt;div class="compare-diagram"&gt;
&lt;svg viewBox="0 0 640 360" xmlns="http://www.w3.org/2000/svg"&gt;&lt;defs&gt;&lt;marker id="arrowA" viewBox="0 0 10 10" refX="8" refY="5" markerWidth="6" markerHeight="6" orient="auto-start-reverse"&gt;&lt;path d="M0,0 L10,5 L0,10 z" style="fill:var(--compare-a)"/&gt;&lt;/marker&gt;&lt;marker id="arrowB" viewBox="0 0 10 10" refX="8" refY="5" markerWidth="6" markerHeight="6" orient="auto-start-reverse"&gt;&lt;path d="M0,0 L10,5 L0,10 z" style="fill:var(--compare-b)"/&gt;&lt;/marker&gt;&lt;/defs&gt;&lt;line x1="320" y1="20" x2="320" y2="340" style="stroke:var(--border)" stroke-width="1" stroke-dasharray="4,4"/&gt;&lt;text x="160" y="28" text-anchor="middle" font-size="16" font-weight="bold" style="fill:var(--primary)"&gt;Lambda Architecture&lt;/text&gt;&lt;text x="480" y="28" text-anchor="middle" font-size="16" font-weight="bold" style="fill:var(--primary)"&gt;Kappa Architecture&lt;/text&gt;&lt;rect x="20" y="160" width="70" height="36" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="55" y="182" text-anchor="middle" font-size="10" style="fill:var(--content)"&gt;Data Source&lt;/text&gt;&lt;rect x="130" y="70" width="130" height="50" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="195" y="92" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;Batch Layer&lt;/text&gt;&lt;text x="195" y="106" text-anchor="middle" font-size="9" style="fill:var(--secondary)"&gt;full recompute&lt;/text&gt;&lt;rect x="130" y="240" width="130" height="50" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="195" y="262" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;Speed Layer&lt;/text&gt;&lt;text x="195" y="276" text-anchor="middle" font-size="9" style="fill:var(--secondary)"&gt;recent, approximate&lt;/text&gt;&lt;rect x="266" y="158" width="46" height="40" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="289" y="174" text-anchor="middle" font-size="8" style="fill:var(--content)"&gt;Serving&lt;/text&gt;&lt;text x="289" y="186" text-anchor="middle" font-size="8" style="fill:var(--content)"&gt;Layer&lt;/text&gt;&lt;line x1="92" y1="168" x2="128" y2="98" style="stroke:var(--compare-a)" stroke-width="1.5" marker-end="url(#arrowA)"/&gt;&lt;line x1="92" y1="188" x2="128" y2="258" style="stroke:var(--compare-a)" stroke-width="1.5" marker-end="url(#arrowA)"/&gt;&lt;line x1="261" y1="110" x2="267" y2="172" style="stroke:var(--compare-a)" stroke-width="1.5" marker-end="url(#arrowA)"/&gt;&lt;line x1="261" y1="250" x2="267" y2="186" style="stroke:var(--compare-a)" stroke-width="1.5" marker-end="url(#arrowA)"/&gt;&lt;line x1="313" y1="178" x2="317" y2="178" style="stroke:var(--compare-a)" stroke-width="1.5" marker-end="url(#arrowA)"/&gt;&lt;text x="195" y="322" text-anchor="middle" font-size="9" style="fill:var(--secondary)"&gt;reprocess = rerun batch job&lt;/text&gt;&lt;rect x="340" y="160" width="70" height="36" rx="4" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;text x="375" y="182" text-anchor="middle" font-size="10" style="fill:var(--content)"&gt;Event Log&lt;/text&gt;&lt;rect x="450" y="130" width="150" height="90" rx="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;text x="525" y="168" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;Stream Processing&lt;/text&gt;&lt;text x="525" y="182" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;Layer&lt;/text&gt;&lt;text x="525" y="200" text-anchor="middle" font-size="9" style="fill:var(--secondary)"&gt;single codebase&lt;/text&gt;&lt;line x1="411" y1="178" x2="448" y2="175" style="stroke:var(--compare-b)" stroke-width="1.5" marker-end="url(#arrowB)"/&gt;&lt;line x1="601" y1="175" x2="628" y2="175" style="stroke:var(--compare-b)" stroke-width="1.5" marker-end="url(#arrowB)"/&gt;&lt;path d="M460,220 C 420,290 380,290 358,198" fill="none" style="stroke:var(--compare-b)" stroke-width="1.5" stroke-dasharray="4,3" marker-end="url(#arrowB)"/&gt;&lt;text x="470" y="322" text-anchor="middle" font-size="9" style="fill:var(--secondary)"&gt;reprocess = replay the log&lt;/text&gt;&lt;text x="316" y="195" text-anchor="middle" font-size="9" style="fill:var(--secondary)"&gt;query&lt;/text&gt;&lt;text x="614" y="195" text-anchor="middle" font-size="9" style="fill:var(--secondary)"&gt;query&lt;/text&gt;&lt;/svg&gt;
&lt;/div&gt;
&lt;h2 id="comparison-table"&gt;Comparison Table&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Lambda Architecture&lt;/th&gt;
&lt;th&gt;Kappa Architecture&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Ingestion path&lt;/td&gt;
&lt;td&gt;Raw events are forked to both a batch store and a stream processor at once&lt;/td&gt;
&lt;td&gt;Raw events are written once to an immutable, replayable log (e.g. Kafka)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Processing model&lt;/td&gt;
&lt;td&gt;Two independent codebases — a batch job and a stream job — implement the same logic twice&lt;/td&gt;
&lt;td&gt;One stream-processing codebase handles both real-time and historical computation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Historical reprocessing&lt;/td&gt;
&lt;td&gt;The batch layer periodically recomputes results over the entire raw dataset&lt;/td&gt;
&lt;td&gt;Reprocessing means replaying the log from an earlier offset through the same stream job&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;State and storage&lt;/td&gt;
&lt;td&gt;Separate batch views and speed views are maintained independently, often in different stores&lt;/td&gt;
&lt;td&gt;A single serving store is continuously updated by the stream processor&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Result merging&lt;/td&gt;
&lt;td&gt;The query layer merges or reconciles batch and speed views at read time&lt;/td&gt;
&lt;td&gt;No merge step — the stream processor&amp;rsquo;s output is the only view&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Consistency behavior&lt;/td&gt;
&lt;td&gt;Speed layer results are approximate until the batch layer overwrites them, so the two can disagree temporarily&lt;/td&gt;
&lt;td&gt;One computation path avoids batch/speed drift, but correctness depends entirely on the stream engine&amp;rsquo;s guarantees&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Operational overhead&lt;/td&gt;
&lt;td&gt;Higher — two parallel pipelines to build, deploy, and monitor, with duplicated logic&lt;/td&gt;
&lt;td&gt;Lower pipeline count, but requires a log system with long enough retention to support full replays&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best-fit scenario&lt;/td&gt;
&lt;td&gt;Batch and speed logic genuinely differ, or the org already has mature batch infrastructure&lt;/td&gt;
&lt;td&gt;Team wants one canonical pipeline and has a stream engine that can absorb both live and replay traffic&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="key-differences"&gt;Key Differences&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Lambda splits ingestion across a &lt;strong class="kw"&gt;batch layer&lt;/strong&gt; and a speed layer, while Kappa sends everything through one stream.&lt;/li&gt;
&lt;li&gt;Reprocessing history in Lambda means rerunning a full batch job; in Kappa it means &lt;strong class="kw"&gt;replaying the log&lt;/strong&gt; through the same stream code.&lt;/li&gt;
&lt;li&gt;Lambda&amp;rsquo;s query layer must reconcile two separate views; Kappa exposes a single &lt;strong class="kw"&gt;serving store&lt;/strong&gt; with no merge step.&lt;/li&gt;
&lt;li&gt;Kappa&amp;rsquo;s design hinges on a durable, long-retention &lt;strong class="kw"&gt;event log&lt;/strong&gt; capable of full replays, which Lambda does not require.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="when-to-use-each"&gt;When to Use Each&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Lambda Architecture&lt;/strong&gt;&lt;/p&gt;</description></item></channel></rss>