<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Autoscaling on IT Comparison</title><link>https://comparison.metacog.co.kr/tags/autoscaling/</link><description>Recent content in Autoscaling on IT Comparison</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Mon, 03 Aug 2026 05:24:29 +0900</lastBuildDate><atom:link href="https://comparison.metacog.co.kr/tags/autoscaling/index.xml" rel="self" type="application/rss+xml"/><item><title>Horizontal Pod Autoscaler vs Vertical Pod Autoscaler: Scaling Out vs Scaling Up</title><link>https://comparison.metacog.co.kr/posts/2026-08-03-horizontal-pod-autoscaler-vs-vertical-pod-autoscaler-scaling/</link><pubDate>Mon, 03 Aug 2026 05:24:29 +0900</pubDate><guid>https://comparison.metacog.co.kr/posts/2026-08-03-horizontal-pod-autoscaler-vs-vertical-pod-autoscaler-scaling/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;Both controllers watch metrics and adjust Kubernetes workloads automatically, but they scale in different dimensions. The &lt;strong class="kw"&gt;Horizontal Pod Autoscaler&lt;/strong&gt; adds or removes pod replicas to handle load, while the &lt;strong class="kw"&gt;Vertical Pod Autoscaler&lt;/strong&gt; resizes the CPU and memory requests/limits of existing pods.&lt;/p&gt;
&lt;h2 id="comparison-diagram"&gt;Comparison Diagram&lt;/h2&gt;
&lt;div class="compare-diagram"&gt;
&lt;svg viewBox="0 0 640 360" xmlns="http://www.w3.org/2000/svg"&gt;&lt;line x1="320" y1="50" x2="320" y2="340" style="stroke:var(--border)" stroke-width="1.5" stroke-dasharray="4 4"/&gt;&lt;text x="160" y="32" text-anchor="middle" style="fill:var(--primary)" font-size="16" font-weight="bold"&gt;Horizontal Pod Autoscaler&lt;/text&gt;&lt;text x="480" y="32" text-anchor="middle" style="fill:var(--primary)" font-size="16" font-weight="bold"&gt;Vertical Pod Autoscaler&lt;/text&gt;&lt;text x="160" y="58" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;before&lt;/text&gt;&lt;rect x="130" y="68" width="60" height="48" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="160" y="96" text-anchor="middle" style="fill:var(--content)" font-size="11"&gt;pod&lt;/text&gt;&lt;line x1="160" y1="122" x2="160" y2="148" style="stroke:var(--compare-a)" stroke-width="2" marker-end="url(#arrowA)"/&gt;&lt;text x="160" y="166" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;after (load increases)&lt;/text&gt;&lt;rect x="70" y="178" width="50" height="44" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;rect x="135" y="178" width="50" height="44" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;rect x="200" y="178" width="50" height="44" rx="4" style="fill:var(--compare-a-soft);stroke:var(--compare-a)" stroke-width="1.5"/&gt;&lt;text x="95" y="204" text-anchor="middle" style="fill:var(--content)" font-size="10"&gt;pod&lt;/text&gt;&lt;text x="160" y="204" text-anchor="middle" style="fill:var(--content)" font-size="10"&gt;pod&lt;/text&gt;&lt;text x="225" y="204" text-anchor="middle" style="fill:var(--content)" font-size="10"&gt;pod&lt;/text&gt;&lt;text x="160" y="246" text-anchor="middle" style="fill:var(--primary)" font-size="13" font-weight="bold"&gt;Scale OUT&lt;/text&gt;&lt;text x="160" y="264" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;more replicas, same pod size&lt;/text&gt;&lt;text x="480" y="58" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;before&lt;/text&gt;&lt;rect x="455" y="70" width="50" height="36" rx="4" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;text x="480" y="92" text-anchor="middle" style="fill:var(--content)" font-size="11"&gt;pod&lt;/text&gt;&lt;line x1="480" y1="122" x2="480" y2="150" style="stroke:var(--compare-b)" stroke-width="2" marker-end="url(#arrowB)"/&gt;&lt;text x="480" y="168" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;after (load increases)&lt;/text&gt;&lt;rect x="430" y="180" width="100" height="96" rx="6" style="fill:var(--compare-b-soft);stroke:var(--compare-b)" stroke-width="1.5"/&gt;&lt;text x="480" y="224" text-anchor="middle" style="fill:var(--content)" font-size="11"&gt;pod&lt;/text&gt;&lt;text x="480" y="242" text-anchor="middle" style="fill:var(--content)" font-size="10"&gt;CPU/mem ↑&lt;/text&gt;&lt;text x="480" y="300" text-anchor="middle" style="fill:var(--primary)" font-size="13" font-weight="bold"&gt;Scale UP&lt;/text&gt;&lt;text x="480" y="318" text-anchor="middle" style="fill:var(--secondary)" font-size="11"&gt;bigger pod, same replica count&lt;/text&gt;&lt;defs&gt;&lt;marker id="arrowA" markerWidth="8" markerHeight="8" refX="4" refY="4" orient="auto"&gt;&lt;path d="M0,0 L8,4 L0,8 Z" style="fill:var(--compare-a)"/&gt;&lt;/marker&gt;&lt;marker id="arrowB" markerWidth="8" markerHeight="8" refX="4" refY="4" orient="auto"&gt;&lt;path d="M0,0 L8,4 L0,8 Z" style="fill:var(--compare-b)"/&gt;&lt;/marker&gt;&lt;/defs&gt;&lt;/svg&gt;
&lt;/div&gt;
&lt;h2 id="comparison-table"&gt;Comparison Table&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Horizontal Pod Autoscaler&lt;/th&gt;
&lt;th&gt;Vertical Pod Autoscaler&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;What it adjusts&lt;/td&gt;
&lt;td&gt;Number of pod replicas in a Deployment/ReplicaSet/StatefulSet&lt;/td&gt;
&lt;td&gt;CPU and memory requests/limits on the pod&amp;rsquo;s containers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Metrics source&lt;/td&gt;
&lt;td&gt;Metrics Server or custom/external metrics (CPU, memory, custom queries) via metrics.k8s.io API&lt;/td&gt;
&lt;td&gt;Historical and current usage sampled by the VPA recommender component&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Trigger condition&lt;/td&gt;
&lt;td&gt;Observed metric crosses a target threshold averaged across pods&lt;/td&gt;
&lt;td&gt;Recommender detects requests are consistently over- or under-provisioned&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Action taken&lt;/td&gt;
&lt;td&gt;Creates or deletes pod replicas to match target replica count&lt;/td&gt;
&lt;td&gt;Evicts and recreates pods with new resource requests (or just recommends, depending on updateMode)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Disruption to running pods&lt;/td&gt;
&lt;td&gt;None — existing pods are untouched, new ones are added or removed&lt;/td&gt;
&lt;td&gt;Pod restart required to apply new resource values, causing brief downtime unless using in-place resize&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best fit for workload type&lt;/td&gt;
&lt;td&gt;Stateless, horizontally scalable services behind a Service/load balancer&lt;/td&gt;
&lt;td&gt;Single-instance or hard-to-replicate workloads, or right-sizing before enabling HPA&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Conflict risk&lt;/td&gt;
&lt;td&gt;Can fight with VPA if both manage CPU on the same workload&lt;/td&gt;
&lt;td&gt;Should not manage CPU/memory targeted by HPA on the same workload simultaneously&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Configuration object&lt;/td&gt;
&lt;td&gt;HorizontalPodAutoscaler resource with min/max replicas and target metrics&lt;/td&gt;
&lt;td&gt;VerticalPodAutoscaler resource with updateMode (Off, Initial, Recreate, Auto)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="key-differences"&gt;Key Differences&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;HPA changes &lt;strong class="kw"&gt;replica count&lt;/strong&gt;, VPA changes &lt;strong class="kw"&gt;resource requests&lt;/strong&gt; on existing pods.&lt;/li&gt;
&lt;li&gt;VPA updates typically require a &lt;strong class="kw"&gt;pod restart&lt;/strong&gt; to take effect, while HPA scaling adds/removes pods without disrupting the rest.&lt;/li&gt;
&lt;li&gt;Running both on the &lt;strong class="kw"&gt;same metric&lt;/strong&gt; (like CPU) causes conflicting decisions unless carefully scoped.&lt;/li&gt;
&lt;li&gt;VPA is often used in &lt;strong class="kw"&gt;recommendation-only mode&lt;/strong&gt; to right-size requests before HPA takes over scaling.&lt;/li&gt;
&lt;li&gt;HPA assumes the workload is &lt;strong class="kw"&gt;stateless and replicable&lt;/strong&gt;; VPA fits singleton or stateful workloads that can&amp;rsquo;t simply be duplicated.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="when-to-use-each"&gt;When to Use Each&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Horizontal Pod Autoscaler&lt;/strong&gt;&lt;/p&gt;</description></item></channel></rss>