Overview
Both controllers watch metrics and adjust Kubernetes workloads automatically, but they scale in different dimensions. The Horizontal Pod Autoscaler adds or removes pod replicas to handle load, while the Vertical Pod Autoscaler resizes the CPU and memory requests/limits of existing pods.
Comparison Diagram
Comparison Table
| Aspect | Horizontal Pod Autoscaler | Vertical Pod Autoscaler |
|---|---|---|
| What it adjusts | Number of pod replicas in a Deployment/ReplicaSet/StatefulSet | CPU and memory requests/limits on the pod’s containers |
| Metrics source | Metrics Server or custom/external metrics (CPU, memory, custom queries) via metrics.k8s.io API | Historical and current usage sampled by the VPA recommender component |
| Trigger condition | Observed metric crosses a target threshold averaged across pods | Recommender detects requests are consistently over- or under-provisioned |
| Action taken | Creates or deletes pod replicas to match target replica count | Evicts and recreates pods with new resource requests (or just recommends, depending on updateMode) |
| Disruption to running pods | None — existing pods are untouched, new ones are added or removed | Pod restart required to apply new resource values, causing brief downtime unless using in-place resize |
| Best fit for workload type | Stateless, horizontally scalable services behind a Service/load balancer | Single-instance or hard-to-replicate workloads, or right-sizing before enabling HPA |
| Conflict risk | Can fight with VPA if both manage CPU on the same workload | Should not manage CPU/memory targeted by HPA on the same workload simultaneously |
| Configuration object | HorizontalPodAutoscaler resource with min/max replicas and target metrics | VerticalPodAutoscaler resource with updateMode (Off, Initial, Recreate, Auto) |
Key Differences
- HPA changes replica count, VPA changes resource requests on existing pods.
- VPA updates typically require a pod restart to take effect, while HPA scaling adds/removes pods without disrupting the rest.
- Running both on the same metric (like CPU) causes conflicting decisions unless carefully scoped.
- VPA is often used in recommendation-only mode to right-size requests before HPA takes over scaling.
- HPA assumes the workload is stateless and replicable; VPA fits singleton or stateful workloads that can’t simply be duplicated.
When to Use Each
Horizontal Pod Autoscaler
- Stateless web/API services: Traffic-driven services behind a load balancer scale cleanly by adding more identical replicas.
- Bursty, unpredictable load: HPA reacts quickly by spinning replicas up or down as request volume changes.
- Queue-depth driven workers: Custom metrics like queue length map naturally to ‘how many workers do I need’ rather than ‘how big should each worker be’.
Vertical Pod Autoscaler
- Single-instance or stateful services: Workloads like a primary database or leader-elected process can’t be horizontally duplicated, so resizing is the only scaling lever.
- Right-sizing resource requests: Run VPA in recommendation mode to find correct CPU/memory requests and eliminate guesswork or over-provisioning.
- Memory-bound workloads: Some processes need more memory per instance rather than more instances, which only vertical scaling addresses.