<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Throughput on IT Comparison</title><link>https://comparison.metacog.co.kr/tags/throughput/</link><description>Recent content in Throughput on IT Comparison</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Sun, 06 Sep 2026 09:34:39 +0900</lastBuildDate><atom:link href="https://comparison.metacog.co.kr/tags/throughput/index.xml" rel="self" type="application/rss+xml"/><item><title>Latency vs Throughput: Response Time vs Processing Volume</title><link>https://comparison.metacog.co.kr/posts/2026-09-06-latency-vs-throughput-response-time-vs-processing-volume/</link><pubDate>Sun, 06 Sep 2026 09:34:39 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-09-06-latency-vs-throughput-response-time-vs-processing-volume/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;Latency and throughput are two orthogonal measures of system performance: &lt;strong class="kw"&gt;latency&lt;/strong&gt; is the time a single request takes to complete, while &lt;strong class="kw"&gt;throughput&lt;/strong&gt; is the volume of work a system finishes per unit of time. The distinction matters because architectures optimized for one can quietly degrade the other.&lt;/p&gt;
&lt;h2 id="comparison-diagram"&gt;Comparison Diagram&lt;/h2&gt;
&lt;div class="compare-diagram"&gt;
&lt;svg viewBox="0 0 640 360" xmlns="http://www.w3.org/2000/svg"&gt;
&lt;text x="320" y="28" text-anchor="middle" font-size="20" font-weight="bold" style="fill:var(--primary)"&gt;Latency&lt;/text&gt;
&lt;rect x="40" y="50" width="60" height="40" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;
&lt;text x="70" y="75" text-anchor="middle" font-size="12" style="fill:var(--content)"&gt;Client&lt;/text&gt;
&lt;rect x="540" y="50" width="60" height="40" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;
&lt;text x="570" y="75" text-anchor="middle" font-size="12" style="fill:var(--content)"&gt;Server&lt;/text&gt;
&lt;line x1="100" y1="70" x2="540" y2="70" style="stroke:var(--border)" stroke-width="2" stroke-dasharray="4 4"/&gt;
&lt;circle cx="320" cy="70" r="9" style="fill:var(--compare-a);stroke:var(--compare-a)"/&gt;
&lt;path d="M330,64 L342,70 L330,76 Z" style="fill:var(--compare-a)"/&gt;
&lt;line x1="100" y1="120" x2="540" y2="120" style="stroke:var(--compare-a)" stroke-width="1.5"/&gt;
&lt;line x1="100" y1="113" x2="100" y2="127" style="stroke:var(--compare-a)" stroke-width="1.5"/&gt;
&lt;line x1="540" y1="113" x2="540" y2="127" style="stroke:var(--compare-a)" stroke-width="1.5"/&gt;
&lt;path d="M100,120 L108,116 L108,124 Z" style="fill:var(--compare-a)"/&gt;
&lt;path d="M540,120 L532,116 L532,124 Z" style="fill:var(--compare-a)"/&gt;
&lt;text x="320" y="145" text-anchor="middle" font-size="13" style="fill:var(--secondary)"&gt;Time for ONE request to complete&lt;/text&gt;
&lt;line x1="20" y1="180" x2="620" y2="180" style="stroke:var(--border)" stroke-width="1" stroke-dasharray="2 4"/&gt;
&lt;text x="320" y="208" text-anchor="middle" font-size="20" font-weight="bold" style="fill:var(--primary)"&gt;Throughput&lt;/text&gt;
&lt;rect x="40" y="225" width="60" height="40" rx="4" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;
&lt;text x="70" y="250" text-anchor="middle" font-size="12" style="fill:var(--content)"&gt;Client&lt;/text&gt;
&lt;rect x="540" y="225" width="60" height="40" rx="4" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;
&lt;text x="570" y="250" text-anchor="middle" font-size="12" style="fill:var(--content)"&gt;Server&lt;/text&gt;
&lt;rect x="100" y="235" width="440" height="20" rx="10" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;
&lt;circle cx="140" cy="245" r="6" style="fill:var(--compare-b);stroke:var(--compare-b)"/&gt;
&lt;circle cx="200" cy="245" r="6" style="fill:var(--compare-b);stroke:var(--compare-b)"/&gt;
&lt;circle cx="260" cy="245" r="6" style="fill:var(--compare-b);stroke:var(--compare-b)"/&gt;
&lt;circle cx="320" cy="245" r="6" style="fill:var(--compare-b);stroke:var(--compare-b)"/&gt;
&lt;circle cx="380" cy="245" r="6" style="fill:var(--compare-b);stroke:var(--compare-b)"/&gt;
&lt;circle cx="440" cy="245" r="6" style="fill:var(--compare-b);stroke:var(--compare-b)"/&gt;
&lt;circle cx="500" cy="245" r="6" style="fill:var(--compare-b);stroke:var(--compare-b)"/&gt;
&lt;path d="M545,245 L557,239 L557,251 Z" style="fill:var(--compare-b)"/&gt;
&lt;text x="320" y="300" text-anchor="middle" font-size="13" style="fill:var(--secondary)"&gt;Total requests completed per second&lt;/text&gt;
&lt;/svg&gt;
&lt;/div&gt;
&lt;h2 id="comparison-table"&gt;Comparison Table&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Latency&lt;/th&gt;
&lt;th&gt;Throughput&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Definition&lt;/td&gt;
&lt;td&gt;Time elapsed for one request to travel and complete&lt;/td&gt;
&lt;td&gt;Amount of work completed across all requests per unit time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;What is measured&lt;/td&gt;
&lt;td&gt;A single request&amp;rsquo;s round trip or processing delay&lt;/td&gt;
&lt;td&gt;Aggregate output of the system over an observation window&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Unit of measurement&lt;/td&gt;
&lt;td&gt;Milliseconds, microseconds, or seconds&lt;/td&gt;
&lt;td&gt;Requests/sec, transactions/sec, or Mbps&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Primary driver&lt;/td&gt;
&lt;td&gt;Network round-trip time, serialization, and processing delay&lt;/td&gt;
&lt;td&gt;Available bandwidth, parallel capacity, and resource pool size&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Effect of concurrency&lt;/td&gt;
&lt;td&gt;Individual request latency can rise as queueing builds up&lt;/td&gt;
&lt;td&gt;Throughput rises with more parallel workers, up to a capacity limit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Behavior under overload&lt;/td&gt;
&lt;td&gt;Tail latency spikes as queues grow (p95/p99 degrade)&lt;/td&gt;
&lt;td&gt;Throughput plateaus or drops once the system saturates&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typical optimization&lt;/td&gt;
&lt;td&gt;Reduce round trips, cache results, shorten the critical path&lt;/td&gt;
&lt;td&gt;Batch requests, add parallel workers, scale out capacity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Measurement method&lt;/td&gt;
&lt;td&gt;Ping, request timers, percentile latency (p50/p95/p99)&lt;/td&gt;
&lt;td&gt;Requests-per-second counters, load testing, capacity benchmarks&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="key-differences"&gt;Key Differences&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Latency measures the &lt;strong class="kw"&gt;time&lt;/strong&gt; for one request; throughput measures the &lt;strong class="kw"&gt;volume&lt;/strong&gt; processed per unit time.&lt;/li&gt;
&lt;li&gt;Batching to raise throughput can increase &lt;strong class="kw"&gt;tail latency&lt;/strong&gt; for individual requests.&lt;/li&gt;
&lt;li&gt;Latency is bounded by physical &lt;strong class="kw"&gt;round-trip time&lt;/strong&gt;; throughput is bounded by system &lt;strong class="kw"&gt;capacity&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Under heavy load, latency &lt;strong class="kw"&gt;spikes&lt;/strong&gt; from queueing while throughput &lt;strong class="kw"&gt;plateaus&lt;/strong&gt; at a ceiling.&lt;/li&gt;
&lt;li&gt;&lt;strong class="kw"&gt;Little&amp;rsquo;s Law&lt;/strong&gt; links the two: average latency times concurrency roughly equals throughput.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="when-to-use-each"&gt;When to Use Each&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Latency&lt;/strong&gt;&lt;/p&gt;</description></item></channel></rss>