Overview

Both controllers watch metrics and adjust Kubernetes workloads automatically, but they scale in different dimensions. The Horizontal Pod Autoscaler adds or removes pod replicas to handle load, while the Vertical Pod Autoscaler resizes the CPU and memory requests/limits of existing pods.

Comparison Diagram

Horizontal Pod AutoscalerVertical Pod Autoscalerbeforepodafter (load increases)podpodpodScale OUTmore replicas, same pod sizebeforepodafter (load increases)podCPU/mem ↑Scale UPbigger pod, same replica count

Comparison Table

AspectHorizontal Pod AutoscalerVertical Pod Autoscaler
What it adjustsNumber of pod replicas in a Deployment/ReplicaSet/StatefulSetCPU and memory requests/limits on the pod’s containers
Metrics sourceMetrics Server or custom/external metrics (CPU, memory, custom queries) via metrics.k8s.io APIHistorical and current usage sampled by the VPA recommender component
Trigger conditionObserved metric crosses a target threshold averaged across podsRecommender detects requests are consistently over- or under-provisioned
Action takenCreates or deletes pod replicas to match target replica countEvicts and recreates pods with new resource requests (or just recommends, depending on updateMode)
Disruption to running podsNone — existing pods are untouched, new ones are added or removedPod restart required to apply new resource values, causing brief downtime unless using in-place resize
Best fit for workload typeStateless, horizontally scalable services behind a Service/load balancerSingle-instance or hard-to-replicate workloads, or right-sizing before enabling HPA
Conflict riskCan fight with VPA if both manage CPU on the same workloadShould not manage CPU/memory targeted by HPA on the same workload simultaneously
Configuration objectHorizontalPodAutoscaler resource with min/max replicas and target metricsVerticalPodAutoscaler resource with updateMode (Off, Initial, Recreate, Auto)

Key Differences

  • HPA changes replica count, VPA changes resource requests on existing pods.
  • VPA updates typically require a pod restart to take effect, while HPA scaling adds/removes pods without disrupting the rest.
  • Running both on the same metric (like CPU) causes conflicting decisions unless carefully scoped.
  • VPA is often used in recommendation-only mode to right-size requests before HPA takes over scaling.
  • HPA assumes the workload is stateless and replicable; VPA fits singleton or stateful workloads that can’t simply be duplicated.

When to Use Each

Horizontal Pod Autoscaler

  • Stateless web/API services: Traffic-driven services behind a load balancer scale cleanly by adding more identical replicas.
  • Bursty, unpredictable load: HPA reacts quickly by spinning replicas up or down as request volume changes.
  • Queue-depth driven workers: Custom metrics like queue length map naturally to ‘how many workers do I need’ rather than ‘how big should each worker be’.

Vertical Pod Autoscaler

  • Single-instance or stateful services: Workloads like a primary database or leader-elected process can’t be horizontally duplicated, so resizing is the only scaling lever.
  • Right-sizing resource requests: Run VPA in recommendation mode to find correct CPU/memory requests and eliminate guesswork or over-provisioning.
  • Memory-bound workloads: Some processes need more memory per instance rather than more instances, which only vertical scaling addresses.