<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Big-Data-Architecture on IT Comparison</title><link>https://comparison.metacog.co.kr/tags/big-data-architecture/</link><description>Recent content in Big-Data-Architecture on IT Comparison</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Tue, 04 Aug 2026 05:21:28 +0900</lastBuildDate><atom:link href="https://comparison.metacog.co.kr/tags/big-data-architecture/index.xml" rel="self" type="application/rss+xml"/><item><title>Lambda Architecture vs Kappa Architecture: Batch-Plus-Speed vs Single-Stream Pipelines</title><link>https://comparison.metacog.co.kr/posts/2026-08-04-lambda-architecture-vs-kappa-architecture-batch-plus-speed-v/</link><pubDate>Tue, 04 Aug 2026 05:21:28 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-04-lambda-architecture-vs-kappa-architecture-batch-plus-speed-v/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;Lambda Architecture and Kappa Architecture are two patterns for building big-data pipelines that need both real-time and historical results. &lt;strong class="kw"&gt;Lambda&lt;/strong&gt; runs parallel batch and speed layers that get merged at query time, while &lt;strong class="kw"&gt;Kappa&lt;/strong&gt; pushes everything through a single, replayable stream pipeline. The choice determines how much duplicate logic you maintain and how reprocessing actually works.&lt;/p&gt;
&lt;h2 id="comparison-diagram"&gt;Comparison Diagram&lt;/h2&gt;
&lt;div class="compare-diagram"&gt;
&lt;svg viewBox="0 0 640 360" xmlns="http://www.w3.org/2000/svg"&gt;&lt;defs&gt;&lt;marker id="arrowA" viewBox="0 0 10 10" refX="8" refY="5" markerWidth="6" markerHeight="6" orient="auto-start-reverse"&gt;&lt;path d="M0,0 L10,5 L0,10 z" style="fill:var(--compare-a)"/&gt;&lt;/marker&gt;&lt;marker id="arrowB" viewBox="0 0 10 10" refX="8" refY="5" markerWidth="6" markerHeight="6" orient="auto-start-reverse"&gt;&lt;path d="M0,0 L10,5 L0,10 z" style="fill:var(--compare-b)"/&gt;&lt;/marker&gt;&lt;/defs&gt;&lt;line x1="320" y1="20" x2="320" y2="340" style="stroke:var(--border)" stroke-width="1" stroke-dasharray="4,4"/&gt;&lt;text x="160" y="28" text-anchor="middle" font-size="16" font-weight="bold" style="fill:var(--primary)"&gt;Lambda Architecture&lt;/text&gt;&lt;text x="480" y="28" text-anchor="middle" font-size="16" font-weight="bold" style="fill:var(--primary)"&gt;Kappa Architecture&lt;/text&gt;&lt;rect x="20" y="160" width="70" height="36" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="55" y="182" text-anchor="middle" font-size="10" style="fill:var(--content)"&gt;Data Source&lt;/text&gt;&lt;rect x="130" y="70" width="130" height="50" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="195" y="92" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;Batch Layer&lt;/text&gt;&lt;text x="195" y="106" text-anchor="middle" font-size="9" style="fill:var(--secondary)"&gt;full recompute&lt;/text&gt;&lt;rect x="130" y="240" width="130" height="50" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="195" y="262" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;Speed Layer&lt;/text&gt;&lt;text x="195" y="276" text-anchor="middle" font-size="9" style="fill:var(--secondary)"&gt;recent, approximate&lt;/text&gt;&lt;rect x="266" y="158" width="46" height="40" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="289" y="174" text-anchor="middle" font-size="8" style="fill:var(--content)"&gt;Serving&lt;/text&gt;&lt;text x="289" y="186" text-anchor="middle" font-size="8" style="fill:var(--content)"&gt;Layer&lt;/text&gt;&lt;line x1="92" y1="168" x2="128" y2="98" style="stroke:var(--compare-a)" stroke-width="1.5" marker-end="url(#arrowA)"/&gt;&lt;line x1="92" y1="188" x2="128" y2="258" style="stroke:var(--compare-a)" stroke-width="1.5" marker-end="url(#arrowA)"/&gt;&lt;line x1="261" y1="110" x2="267" y2="172" style="stroke:var(--compare-a)" stroke-width="1.5" marker-end="url(#arrowA)"/&gt;&lt;line x1="261" y1="250" x2="267" y2="186" style="stroke:var(--compare-a)" stroke-width="1.5" marker-end="url(#arrowA)"/&gt;&lt;line x1="313" y1="178" x2="317" y2="178" style="stroke:var(--compare-a)" stroke-width="1.5" marker-end="url(#arrowA)"/&gt;&lt;text x="195" y="322" text-anchor="middle" font-size="9" style="fill:var(--secondary)"&gt;reprocess = rerun batch job&lt;/text&gt;&lt;rect x="340" y="160" width="70" height="36" rx="4" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;text x="375" y="182" text-anchor="middle" font-size="10" style="fill:var(--content)"&gt;Event Log&lt;/text&gt;&lt;rect x="450" y="130" width="150" height="90" rx="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;text x="525" y="168" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;Stream Processing&lt;/text&gt;&lt;text x="525" y="182" text-anchor="middle" font-size="11" style="fill:var(--content)"&gt;Layer&lt;/text&gt;&lt;text x="525" y="200" text-anchor="middle" font-size="9" style="fill:var(--secondary)"&gt;single codebase&lt;/text&gt;&lt;line x1="411" y1="178" x2="448" y2="175" style="stroke:var(--compare-b)" stroke-width="1.5" marker-end="url(#arrowB)"/&gt;&lt;line x1="601" y1="175" x2="628" y2="175" style="stroke:var(--compare-b)" stroke-width="1.5" marker-end="url(#arrowB)"/&gt;&lt;path d="M460,220 C 420,290 380,290 358,198" fill="none" style="stroke:var(--compare-b)" stroke-width="1.5" stroke-dasharray="4,3" marker-end="url(#arrowB)"/&gt;&lt;text x="470" y="322" text-anchor="middle" font-size="9" style="fill:var(--secondary)"&gt;reprocess = replay the log&lt;/text&gt;&lt;text x="316" y="195" text-anchor="middle" font-size="9" style="fill:var(--secondary)"&gt;query&lt;/text&gt;&lt;text x="614" y="195" text-anchor="middle" font-size="9" style="fill:var(--secondary)"&gt;query&lt;/text&gt;&lt;/svg&gt;
&lt;/div&gt;
&lt;h2 id="comparison-table"&gt;Comparison Table&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Lambda Architecture&lt;/th&gt;
&lt;th&gt;Kappa Architecture&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Ingestion path&lt;/td&gt;
&lt;td&gt;Raw events are forked to both a batch store and a stream processor at once&lt;/td&gt;
&lt;td&gt;Raw events are written once to an immutable, replayable log (e.g. Kafka)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Processing model&lt;/td&gt;
&lt;td&gt;Two independent codebases — a batch job and a stream job — implement the same logic twice&lt;/td&gt;
&lt;td&gt;One stream-processing codebase handles both real-time and historical computation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Historical reprocessing&lt;/td&gt;
&lt;td&gt;The batch layer periodically recomputes results over the entire raw dataset&lt;/td&gt;
&lt;td&gt;Reprocessing means replaying the log from an earlier offset through the same stream job&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;State and storage&lt;/td&gt;
&lt;td&gt;Separate batch views and speed views are maintained independently, often in different stores&lt;/td&gt;
&lt;td&gt;A single serving store is continuously updated by the stream processor&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Result merging&lt;/td&gt;
&lt;td&gt;The query layer merges or reconciles batch and speed views at read time&lt;/td&gt;
&lt;td&gt;No merge step — the stream processor&amp;rsquo;s output is the only view&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Consistency behavior&lt;/td&gt;
&lt;td&gt;Speed layer results are approximate until the batch layer overwrites them, so the two can disagree temporarily&lt;/td&gt;
&lt;td&gt;One computation path avoids batch/speed drift, but correctness depends entirely on the stream engine&amp;rsquo;s guarantees&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Operational overhead&lt;/td&gt;
&lt;td&gt;Higher — two parallel pipelines to build, deploy, and monitor, with duplicated logic&lt;/td&gt;
&lt;td&gt;Lower pipeline count, but requires a log system with long enough retention to support full replays&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best-fit scenario&lt;/td&gt;
&lt;td&gt;Batch and speed logic genuinely differ, or the org already has mature batch infrastructure&lt;/td&gt;
&lt;td&gt;Team wants one canonical pipeline and has a stream engine that can absorb both live and replay traffic&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="key-differences"&gt;Key Differences&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Lambda splits ingestion across a &lt;strong class="kw"&gt;batch layer&lt;/strong&gt; and a speed layer, while Kappa sends everything through one stream.&lt;/li&gt;
&lt;li&gt;Reprocessing history in Lambda means rerunning a full batch job; in Kappa it means &lt;strong class="kw"&gt;replaying the log&lt;/strong&gt; through the same stream code.&lt;/li&gt;
&lt;li&gt;Lambda&amp;rsquo;s query layer must reconcile two separate views; Kappa exposes a single &lt;strong class="kw"&gt;serving store&lt;/strong&gt; with no merge step.&lt;/li&gt;
&lt;li&gt;Kappa&amp;rsquo;s design hinges on a durable, long-retention &lt;strong class="kw"&gt;event log&lt;/strong&gt; capable of full replays, which Lambda does not require.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="when-to-use-each"&gt;When to Use Each&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Lambda Architecture&lt;/strong&gt;&lt;/p&gt;</description></item></channel></rss>