
Kubernetes is the system most teams now use to run containers in production. In the CNCF annual survey announced on January 20, 2026, 82% of container users reported running Kubernetes in production, up from 66% in 2023. Operating it well is now a core engineering skill, and this post teaches it in three layers.
The first layer explains what Kubernetes is, the problem it solves, and the objects you work with every day. The second shows how a cluster is built and what happens between typing kubectl apply and a running container. The third is a production baseline for a stateless HTTP service: seven manifests, each explained, plus the 2026 changes that can break a cluster you have not touched in a year.
Kubernetes 1.37 shipped on August 26, 2026, and the project now supports three minor versions at once. Every feature below carries its maturity level and the version that set it. Where a claim comes from a release note, a blog post or a survey, it links to that source. Where the baseline depends on something your cluster must provide, such as a Gateway API controller or a CNI plugin that enforces NetworkPolicy, the text says so.
What Kubernetes is, and the problem it solves
In the project's own words, Kubernetes is a portable, extensible, open source platform for managing containerized workloads and services, with both declarative configuration and automation. The name comes from the Greek word for helmsman or pilot, and K8s abbreviates it by counting the eight letters between the K and the s. Google open sourced the project in 2014, and it combines over 15 years of Google's experience running production workloads at scale with ideas from the community.
To see why it exists, follow how deployment changed. Teams first ran applications on physical servers, where one application could take most of a machine's resources and starve the others. Virtual machines fixed that by isolating applications, but each VM runs a full operating system. Containers relax that isolation and share the host operating system, so they are lighter and portable across clouds and Linux distributions.
Containers solve packaging, not operations. In production something must restart a container that dies, spread containers across many machines, send traffic only to the healthy ones, and roll out new versions without downtime. The overview puts it plainly: if a container goes down, another container needs to start, and it would be easier if a system handled that. Kubernetes is that system.
The overview lists what you get from it, summarized here.
Capability | What it does |
|---|---|
Service discovery and load balancing | Exposes a container by DNS name or IP address and spreads traffic across copies |
Storage orchestration | Mounts the storage system you choose, from local disks to public cloud storage |
Automated rollouts and rollbacks | Moves the actual state to the state you describe, at a controlled rate |
Automatic bin packing | Places containers on nodes using the CPU and memory you declare for each |
Self-healing | Restarts failed containers and withholds traffic until a container is ready |
Secret and configuration management | Updates secrets and configuration without rebuilding images |
Batch execution | Runs batch and CI workloads and replaces failed containers if you want |
Horizontal scaling | Scales up or down by command, in a UI, or automatically on CPU usage |
It also helps to know what Kubernetes leaves out. It is not an all-inclusive platform as a service: it does not build your code, run your CI/CD pipeline, or ship databases, message buses and caches as built-in services, and it does not choose your logging or monitoring tools. It is also not orchestration in the strict sense of running step A, then B, then C. It is a set of independent control processes that keep driving the current state toward the state you declared.
The core idea: you declare, controllers converge
Everything in Kubernetes is an object, a persistent record in the cluster. The objects documentation calls an object a record of intent: once you create it, the system works constantly to make it exist. Almost every object has a spec, which is the state you want, and a status, which Kubernetes fills in with the state it observes.
Take a Deployment whose spec asks for three replicas. Kubernetes reads the spec and starts three instances. If one fails, the status no longer matches the spec, and Kubernetes starts a replacement. The component doing this work is a controller, which runs a control loop, a non-terminating loop that regulates the state of a system. The docs compare it to a thermostat that switches equipment on and off to bring a room toward the temperature you set.
The practical consequence is that you rarely issue instructions such as "start this container". You change the declared state, and controllers work out the steps. That is why this post is built from YAML files you apply, not scripts you run.
The objects you use every day
A handful of object types cover most applications. The table lists them in the order you meet them.
Object | What it is | How you use it |
|---|---|---|
Pod | The smallest deployable unit: one or more tightly coupled containers with shared storage and network | The unit that actually runs, normally created through a workload resource and not by hand |
ReplicaSet | Keeps a stable set of replica Pods running | Rarely touched directly, because a Deployment manages it for you |
Deployment | Declarative updates for Pods and ReplicaSets, usually for apps that keep no state | Running and updating a stateless service |
Service | A stable network endpoint for a changing set of Pods | Giving clients one name and address to call |
ConfigMap | Non-confidential configuration stored as key-value pairs | Settings consumed as environment variables, arguments or files |
Secret | A small amount of sensitive data such as a password, token or key | Credentials kept out of images and Pod specs |
Namespace | A scope that isolates groups of resources inside one cluster | Separating teams or projects, since names must be unique only within a namespace |
Node | A worker machine, virtual or physical, that runs Pods | The capacity underneath everything, managed for you by the control plane |
Other workload resources cover other shapes of work. A StatefulSet gives each Pod a sticky identity for applications that need persistent storage or a stable network identity. A DaemonSet runs node-local facilities such as a networking helper. A Job runs a one-off task to completion.
Pods are ephemeral. The Service documentation says you should not expect any individual Pod to be reliable or durable, because Pods are created and destroyed to match the desired state. Each Pod gets its own IP address, but the set of healthy Pods keeps changing. That is the reason Services exist: clients need one stable name in front of Pods that come and go.
One caution about Secrets. By default they are stored unencrypted in etcd, and anyone with API access can read them, as can anyone allowed to create a Pod in the namespace. The Secret documentation tells you to enable encryption at rest and to restrict access with RBAC. A ConfigMap provides no secrecy at all.
Here is a minimal Deployment, the object you will write most often.
# hello.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: hello-node
spec:
replicas: 2
selector:
matchLabels:
app: hello-node
template:
metadata:
labels:
app: hello-node
spec:
containers:
- name: agnhost
image: registry.k8s.io/e2e-test-images/agnhost:2.53
command: ["/agnhost", "netexec", "--http-port=8080"]
ports:
- containerPort: 8080Read it from the top. apiVersion and kind name the API and the object type, and metadata.name names the object. Everything under spec is desired state: two replicas of a Pod built from the template. The selector tells the Deployment which Pods it owns by matching the template's labels, and the Deployment documentation says the API rejects a Deployment whose two do not match. Labels are plain key-value pairs that Services, Deployments and policies use to find Pods.
How a cluster is built
A cluster has a control plane and one or more worker nodes. The components overview says the control plane components manage the overall state of the cluster, and the node components run on every node, maintaining running Pods and providing the runtime environment.
Component | Runs on | Job |
|---|---|---|
kube-apiserver | Control plane | Serves the Kubernetes HTTP API as the front end of the control plane, and scales horizontally by running more instances |
etcd | Control plane | Consistent, highly available key-value store for all cluster data, which you need a backup plan for |
kube-scheduler | Control plane | Watches for newly created Pods with no node and selects a node for each |
kube-controller-manager | Control plane | Runs the controller processes, compiled into one binary and run as one process |
cloud-controller-manager | Control plane, optional | Integrates with the underlying cloud provider |
kubelet | Every node | Makes sure the containers described in PodSpecs are running and healthy, and ignores containers Kubernetes did not create |
kube-proxy | Every node, optional | Maintains the network rules that implement Services |
Container runtime | Every node | The software that actually runs containers |
flowchart LR
U["You or a CI job: kubectl"] --> API
subgraph CP["Control plane"]
API["kube-apiserver"] <--> ETCD[("etcd")]
SCH["kube-scheduler"] -- "watch Pods, bind to nodes" --> API
CM["kube-controller-manager"] -- "watch objects, reconcile" --> API
end
subgraph NODE["Worker node"]
KL["kubelet"] -- "watch Pods, report status" --> API
KL --> RT["container runtime"]
RT --> PODS["Pods"]
KP["kube-proxy"] -. "routes Service traffic" .-> PODS
endEvery arrow runs through the API server. The Kubernetes API documentation says users, the different parts of your cluster and external components all communicate with one another through it. The scheduler, the controllers and the kubelets never call each other. Each watches the API for work and writes its results back, so a component that restarts can read the current state again.
Three consequences shape every manifest in this post. The scheduler reads requests, so wrong requests mean wrong placement. The kubelet enforces limits, so wrong limits mean CPU throttling or out-of-memory kills. Controllers retry forever, so a bad rollout never fixes itself unless a probe tells the truth.
What happens when you run kubectl apply
Follow one command through the system. The CLI turns your file into API calls, as the objects documentation describes, and from there the work passes between components without any of them calling another directly.
sequenceDiagram
participant U as kubectl
participant A as kube-apiserver
participant E as etcd
participant C as Controllers
participant S as kube-scheduler
participant K as kubelet on a node
U->>A: create the Deployment
A->>E: store the object
C->>A: see the Deployment, create a ReplicaSet
C->>A: see the ReplicaSet, create Pods
S->>A: see Pods with no node, bind each to a node
K->>A: see Pods bound to this node
K->>K: ask the container runtime to start the containers
K->>A: report Pod statusA chain of hand-offs turns that file into running containers. The API server accepts and stores the Deployment. The Deployment controller creates a ReplicaSet, and the ReplicaSet creates the Pods, which the Deployment documentation describes as happening in the background. The scheduler sees Pods without a node and picks one for each. The kubelet on that node starts the containers through the runtime and reports status back.
Try it on a local cluster
You need a cluster and kubectl. The learning environment page lists kind, which runs cluster nodes as Docker containers, and minikube, which runs a single-node cluster on your machine, among other options. Install one of them and kubectl, then confirm that kubectl get nodes shows a node in the Ready state.
kubectl apply -f hello.yaml
kubectl get deployments
kubectl get pods
kubectl scale deployment/hello-node --replicas=4
kubectl get pods
kubectl delete pod <one-of-the-pod-names>
kubectl get pods
kubectl get events
kubectl port-forward deployment/hello-node 8080:8080Run them in order. The first creates the Deployment and, through it, two Pods. Scaling to four changes the desired state, and Kubernetes starts two more. Deleting a Pod does not lower the count: the ReplicaSet notices the gap and creates a replacement, the behavior the docs list as replica replacement under self-healing. kubectl get events shows the hand-offs from the previous section as they happened.
The port-forward command tunnels local port 8080 to a Pod, so you can open http://localhost:8080 and see the test image echo your request back, the same image the official Hello Minikube tutorial uses. It does not return until you stop it with Ctrl+C. When you are done, run kubectl delete -f hello.yaml to remove the Deployment and its Pods.
Where Kubernetes stands in October 2026
That demo skipped everything production needs: pinned versions, resource limits, probes, security controls and a safe way to take traffic. The rest of this post adds them, starting with the versions you can choose from.
The project maintains release branches for the three most recent minor releases and supports each patch series for roughly fourteen months: twelve in the standard period, then two in maintenance mode, when new releases cover only CVE fixes, dependency issues and critical core component issues. As of October 10, 2026, the releases page lists the following.
Minor | Latest patch listed | End of life |
|---|---|---|
1.37 | 1.37.1 (September 15, 2026) | October 28, 2027 |
1.36 | 1.36.5 | June 28, 2027 |
1.35 | 1.35.9 | February 28, 2027 |
1.34 | 1.34.12 | October 27, 2026 |
The next patch releases for all four lines are scheduled for October 13, 2026, so the patch numbers above will be stale within days. The fourth row, 1.34, is in maintenance mode, which began on August 27, 2026, and its end of life is under three weeks away. Plan the move to 1.35 or later now, and upgrade one minor version at a time, because kubeadm does not support skipping minor versions.
The last two releases were large. Kubernetes 1.36 (Haru), released April 22, 2026, had 70 enhancements: 18 Stable, 25 Beta and 25 Alpha. Kubernetes 1.37 (Garhwal) had 67: 16 Stable, 23 Beta, 27 Alpha and one deprecation or removal. You do not need most of them. You need the handful that change how you write manifests.
The baseline, one manifest at a time
The target is one service called api in a namespace called shop, behind a shared Gateway, with three replicas spread across zones. Replace the names, the host, the image digest and the GatewayClass. Everything else is meant to be applied as written.
flowchart LR
C["Client"] --> GW["Gateway: HTTPS 443"]
GW --> HR["HTTPRoute: host and path rules"]
HR --> SVC["Service: ClusterIP"]
SVC --> PODS
subgraph PODS["Pods spread across three zones"]
direction TB
P1["Pod, zone a"]
P2["Pod, zone b"]
P3["Pod, zone c"]
end
DEP["Deployment"] -- "rolls out" --> PODS
HPA["HPA: 3 to 12 replicas"] -. "sets replica count" .-> DEP
PDB["PodDisruptionBudget: 1 at a time"] -. "limits evictions" .-> PODSSave each block below as its own file, in order, using the file name in the first comment line.
1. Namespace: enforce Pod Security Admission
# 00-namespace.yaml
apiVersion: v1
kind: Namespace
metadata:
name: shop
labels:
pod-security.kubernetes.io/enforce: restricted
pod-security.kubernetes.io/enforce-version: latest
pod-security.kubernetes.io/warn: restricted
pod-security.kubernetes.io/audit: restricted
shared-gateway-access: "true"Pod Security Admission has been stable since 1.25. The restricted profile rejects Pods that run as root, allow privilege escalation, keep default Linux capabilities or lack a seccomp profile. Setting enforce, warn and audit together means you see violations in three places. The last label is used by the Gateway in step 4.
2. Deployment
# 01-deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: api
namespace: shop
labels:
app.kubernetes.io/name: api
spec:
# replicas is omitted on purpose: the HPA in step 6 owns it
revisionHistoryLimit: 3
strategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 0
maxSurge: 1
selector:
matchLabels:
app.kubernetes.io/name: api
template:
metadata:
labels:
app.kubernetes.io/name: api
spec:
hostUsers: false
automountServiceAccountToken: false
terminationGracePeriodSeconds: 45
securityContext:
runAsNonRoot: true
runAsUser: 10001
runAsGroup: 10001
seccompProfile:
type: RuntimeDefault
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: ScheduleAnyway
labelSelector:
matchLabels:
app.kubernetes.io/name: api
containers:
- name: api
# Pin by digest. Replace this placeholder with your real image digest.
image: registry.example.com/shop/api@sha256:0000000000000000000000000000000000000000000000000000000000000000
ports:
- name: http
containerPort: 8080
resources:
requests:
cpu: 250m
memory: 256Mi
limits:
memory: 256Mi
resizePolicy:
- resourceName: cpu
restartPolicy: NotRequired
- resourceName: memory
restartPolicy: RestartContainer
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: ["ALL"]
startupProbe:
httpGet:
path: /healthz
port: http
periodSeconds: 2
failureThreshold: 30
readinessProbe:
httpGet:
path: /readyz
port: http
periodSeconds: 5
failureThreshold: 2
livenessProbe:
httpGet:
path: /healthz
port: http
periodSeconds: 10
failureThreshold: 3
lifecycle:
preStop:
sleep:
seconds: 10
volumeMounts:
- name: tmp
mountPath: /tmp
volumes:
- name: tmp
emptyDir:
sizeLimit: 64MiFour settings carry most of the weight. hostUsers: false puts the Pod in its own user namespace, so root inside the container maps to an unprivileged user on the node. User namespaces reached GA in 1.36 and are Linux-only. The concept page lists the node requirements: Linux 6.3 or later for idmap mounts on tmpfs, runc 1.2 or later or crun 1.9 or later, and containerd 2.0 or later or CRI-O 1.25 or later. Check your nodes before relying on the setting.
maxUnavailable: 0 with maxSurge: 1 makes a rollout add one new Pod and wait for it to be Ready before removing an old one. The topology spread constraint keeps replicas balanced across zones. ScheduleAnyway means a zone outage does not leave Pods Pending. Switch to DoNotSchedule only if imbalance is worse than reduced capacity for your service.
The three probes answer three different questions. The startup probe allows a slow boot of up to 60 seconds (30 failures at 2 second intervals) before the other probes begin. Readiness decides whether the Pod receives traffic, and should check dependencies you cannot serve without. Liveness should check only that the process is not wedged. Never make liveness depend on a database, or a database outage becomes a restart storm.
The preStop sleep and the 45 second grace period solve a race. Endpoint removal and SIGTERM happen at about the same time, so for a few seconds a terminating Pod can still receive requests. Sleeping 10 seconds before SIGTERM gives proxies time to stop sending traffic. Your application must still handle SIGTERM by finishing in-flight requests inside the remaining 35 seconds.
The sleep hook action is Stable since 1.34, when the PodLifecycleSleepAction feature gate became locked on, so every supported minor has it. The lifecycle hooks page lists it beside the exec and HTTP handlers. The kubelet runs it, so your image needs no sleep binary.
Memory request equals memory limit, and there is no CPU limit. Memory cannot be compressed, so an equal limit makes OOM behavior predictable. CPU can be throttled, and a CPU limit often throttles bursts the node could have served. The Pod still gets a CPU guarantee through its request. Revisit this choice if you run a multi-tenant cluster that needs hard caps.
3. Service
# 02-service.yaml
apiVersion: v1
kind: Service
metadata:
name: api
namespace: shop
spec:
type: ClusterIP
selector:
app.kubernetes.io/name: api
ports:
- name: http
port: 80
targetPort: httpThe Service gives the Pods one stable virtual IP and a DNS name, api.shop.svc. It is ClusterIP on purpose. Traffic from outside the cluster enters through the Gateway, not through a Service of type LoadBalancer per application, and not through the deprecated externalIPs field covered later.
4. Gateway and HTTPRoute
# 03-gateway-and-route.yaml
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
name: public
namespace: gateway-system
spec:
gatewayClassName: example-gateway-class
listeners:
- name: https
protocol: HTTPS
port: 443
hostname: api.example.com
tls:
mode: Terminate
certificateRefs:
- kind: Secret
name: api-example-com-tls
allowedRoutes:
namespaces:
from: Selector
selector:
matchLabels:
shared-gateway-access: "true"
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: api
namespace: shop
spec:
parentRefs:
- name: public
namespace: gateway-system
sectionName: https
hostnames:
- api.example.com
rules:
- matches:
- path:
type: PathPrefix
value: /
backendRefs:
- name: api
port: 80Gateway API splits ownership. The platform team owns the Gateway and its listeners, and application teams own the routes that attach to it. Gateway and HTTPRoute have been GA since Gateway API v1.0. Gateway API v1.6.0, released June 30, 2026, also moved TCPRoute and UDPRoute to the Standard channel at API version v1.
Gateway API is not built into the API server. You install its CRDs and a controller that implements your GatewayClass, and example-gateway-class is a placeholder for that controller's class name. The allowedRoutes selector is why step 1 labeled the shop namespace: a route from an unlabeled namespace is not allowed to attach.
The manifest assumes two objects already exist: the gateway-system namespace, and a kubernetes.io/tls Secret named api-example-com-tls inside it. Per the Gateway CRD at v1.6.0, a certificate reference into another namespace is invalid unless a ReferenceGrant allows it, and the listener then reports ResolvedRefs as False with the reason RefNotPermitted. Create both before you apply this file.
5. PodDisruptionBudget
# 04-pdb.yaml
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: api
namespace: shop
spec:
maxUnavailable: 1
selector:
matchLabels:
app.kubernetes.io/name: apiA PodDisruptionBudget limits voluntary evictions, such as a node drain during an upgrade. With it, a drain can take one api Pod at a time and must wait for the replacement to be Ready. It does not protect against crashes or node failures. Prefer maxUnavailable over minAvailable, so the budget keeps working when the HPA changes the replica count.
6. HorizontalPodAutoscaler
# 05-hpa.yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: api
namespace: shop
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: api
minReplicas: 3
maxReplicas: 12
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
behavior:
scaleDown:
stabilizationWindowSeconds: 300
policies:
- type: Percent
value: 25
periodSeconds: 60Leaving spec.replicas out of the Deployment makes the HPA the only writer of the replica count. If both write, every kubectl apply resets it. On the first apply the Deployment starts at one replica until the HPA raises it to minReplicas, so apply both in the same command. CPU utilization is measured against the request, which is why requests must be honest.
Resource metrics come from the metrics.k8s.io API, usually served by metrics-server. That API graduated to Stable in 1.37 after nearly nine years in Beta. The scale-down policy above removes at most 25% of replicas per minute, and only after five minutes of consistently lower demand, which avoids flapping on spiky traffic.
7. NetworkPolicy
# 06-networkpolicy.yaml
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: default-deny
namespace: shop
spec:
podSelector: {}
policyTypes:
- Ingress
- Egress
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: allow-gateway-to-api
namespace: shop
spec:
podSelector:
matchLabels:
app.kubernetes.io/name: api
policyTypes:
- Ingress
ingress:
- from:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: gateway-system
ports:
- protocol: TCP
port: 8080
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: allow-dns-egress
namespace: shop
spec:
podSelector: {}
policyTypes:
- Egress
egress:
- to:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: kube-system
ports:
- protocol: UDP
port: 53
- protocol: TCP
port: 53NetworkPolicy objects are enforced only if your CNI plugin implements them. On a plugin that does not, these manifests are accepted and silently do nothing, so test with a throwaway Pod that should be blocked. Default-deny egress also blocks your database and every external API, so add one allow rule per destination before you ship. Some Gateway controllers run their data plane in a namespace other than gateway-system; adjust the selector to match.
Apply and verify
kubectl apply --dry-run=server -f 00-namespace.yaml
kubectl apply -f 00-namespace.yaml
kubectl apply --dry-run=server -f 01-deployment.yaml -f 02-service.yaml -f 04-pdb.yaml -f 05-hpa.yaml -f 06-networkpolicy.yaml
kubectl apply -f 01-deployment.yaml -f 02-service.yaml -f 04-pdb.yaml -f 05-hpa.yaml -f 06-networkpolicy.yaml
kubectl apply -f 03-gateway-and-route.yaml
kubectl -n shop rollout status deploy/api
kubectl -n shop get deploy,hpa,pdb,pods -o wide
kubectl -n shop get httproute api -o jsonpath='{.status.parents[*].conditions[*].type}{"\n"}'The server-side dry run catches schema errors and admission rejections, including Pod Security violations, without creating anything. The last command prints the route's status condition types; look for Accepted and ResolvedRefs, and read each condition's status field when either is missing.
Resize a running Pod without restarting it
In-place Pod resize reached Stable in 1.35, after starting as Alpha in 1.27. A container's resources are now the desired state, the Pod status shows what is actually applied, and you change the desired values through the Pod's resize subresource. Per the task documentation, the server must be at 1.33 or later and kubectl at 1.32 or later.
kubectl -n shop patch pod <pod-name> --subresource resize --patch \
'{"spec":{"containers":[{"name":"api","resources":{"requests":{"cpu":"500m"}}}]}}'
kubectl -n shop get pod <pod-name> \
-o jsonpath='{.status.containerStatuses[0].resources}{"\n"}'The resizePolicy in step 2 decides the cost. With CPU set to NotRequired, the container keeps running while the kubelet applies the change. With memory set to RestartContainer, a memory change restarts that container, because many runtimes cannot change their heap bounds on the fly. If one patch changes both, the container restarts, as the documentation states.
A resize edits one Pod. The Deployment template still holds the old values, so replacement Pods start with them. Use resize for emergency relief and for autoscalers, since the Vertical Pod Autoscaler's InPlaceOrRecreate mode, which the same post reports as Beta and built on this feature, and change the template for anything lasting.
Scale a worker to zero
HorizontalPodAutoscaler scale to zero reached Beta in 1.37 and is on by default: the HPAScaleToZero feature gate is enabled on both the kube-apiserver and the kube-controller-manager. It works only with an object metric or an external metric. CPU and memory come from running Pods, and at zero replicas there is nothing left to measure, so there would be no signal to scale back up.
# queue-worker-hpa.yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: queue-worker
namespace: shop
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: queue-worker
minReplicas: 0
maxReplicas: 10
metrics:
- type: External
external:
metric:
name: queue_consumer_lag
selector:
matchLabels:
name: worker_tasks
target:
type: AverageValue
averageValue: "30"This HPA asks for one replica per 30 queued tasks and drops to zero when the queue is empty. The external metric must be served through the External Metrics API by an adapter such as Prometheus Adapter, and you can confirm Kubernetes can read it with kubectl get --raw before creating the HPA.
The target type matters. The release post's example uses Value, but in the controller source at v1.37.1 a Value target multiplies the metric-to-target ratio by the number of ready Pods, so a queue that stays long keeps multiplying the replica count. AverageValue computes the queue length divided by 30, rounded up, and the HPA walkthrough uses it for its own queue example.
Start the Deployment with at least one replica. The release post says manually setting a Deployment to zero has always paused autoscaling, and a workload at zero is resumed only if the HPA itself scaled it down and recorded the ScaledToZero condition.
Do not use this for the api service. The release notes are explicit that Kubernetes Services do not buffer requests while no Pods are ready, so request-driven HTTP workloads need a separate buffering layer. Scale to zero fits work that can wait in a durable queue, where a cold start costs seconds of latency and nobody is waiting on a socket.
What to fix before you upgrade
Four 2026 changes can bite a cluster that has been quiet. This script covers the two you can detect from the API. The other two need a look at your nodes and your release notes.
#!/usr/bin/env bash
# preflight.sh: run with cluster-admin credentials
set -euo pipefail
echo "== 1. Is ingress-nginx still running?"
kubectl get pods --all-namespaces --selector app.kubernetes.io/name=ingress-nginx
echo "== 2. Services that still set spec.externalIPs"
kubectl get svc --all-namespaces -o json \
| jq -r '.items[]
| select((.spec.externalIPs // []) | length > 0)
| "\(.metadata.namespace)/\(.metadata.name)"'
echo "== 3. Client and server versions"
kubectl versionIngress NGINX is retired
The Kubernetes project said best-effort maintenance of Ingress NGINX would continue until March 2026. After that there would be no releases, no bug fixes and no security updates, although existing deployments keep working. That date has passed, and the GitHub repository now shows it as archived on March 24, 2026 and read-only. A statement from the Steering and Security Response Committees, citing internal Datadog research, put its use at about 50% of cloud native environments.
The statement is blunt about the cost of waiting: staying on a retired controller leaves you and your users vulnerable to attack, and none of the alternatives is a direct drop-in replacement. Install a Gateway API controller next to Ingress NGINX, translate your routes (the project's ingress2gateway tool reached 1.0 in March 2026), test, then move traffic over gradually.
Review the translated output by hand. The ingress2gateway 1.0 post shows the tool warning that it does not support the configuration-snippet annotation, and the retirement notice names the "snippets" annotations, which add arbitrary NGINX configuration directives, as an option now considered a serious security flaw. Rebuild what each snippet did with Gateway API features instead of copying it across.
Service externalIPs is deprecated
The .spec.externalIPs field was formally deprecated in 1.36 because it assumes every user in the cluster is trusted, which enables the attacks described in CVE-2020-8554. A future minor release will drop the behavior from kube-proxy and require conformant implementations not to support it.
If the preflight script prints anything, plan a replacement. The same post calls a hand-assigned LoadBalancer IP the easiest option and also the worst, because users who may set it can still replicate the CVE. It lists two better routes: a load balancer controller such as MetalLB, or a Gateway whose address you set in .spec.addresses. As a precaution, enable the DenyServiceExternalIPs admission controller so nobody adds new uses.
cgroup v1 nodes will not start
Support for cgroup v2 has been stable since 1.25, and cgroup v1 went into maintenance mode in 1.31. Starting with 1.35, the kubelet option failCgroupV1 defaults to true, so the kubelet refuses to start on a cgroup v1 node unless you set it to false, a temporary override while removal is tracked in KEP-5573. On each node, stat -fc %T /sys/fs/cgroup prints cgroup2fs for v2 and tmpfs for v1. Fix every v1 node before the upgrade, not during it.
SELinux volume labeling changes
In 1.37 the SELinuxMount and SELinuxChangePolicy features reached Stable and are enabled by default. Volumes then mount with one SELinux context instead of being relabeled recursively, but only for CSI drivers that opt in with .spec.seLinuxMount: true. Pods with different SELinux labels that share a volume on one node can therefore fail to start.
The release notes advise setting .spec.seLinuxChangePolicy to Recursive on a Pod to keep the old behavior, and they say clusters without SELinux see no effect. If your nodes enforce SELinux, as on the RHEL family, read the project's post on SELinux volume label changes and their implications in 1.37 before upgrading, and test any workload where Pods share a volume.
Where AI workloads fit
The same CNCF survey found that 66% of organizations hosting generative AI models use Kubernetes for some or all of their inference. It also found that only 7% of organizations deploy models daily and that 44% do not yet run AI or ML workloads on Kubernetes. Most teams are early, which makes the platform choices you make now easy to get right and expensive to redo.
The baseline in this post is the right foundation for an inference service, with one change: GPUs and other devices should be requested through Dynamic Resource Allocation, not through the old device-plugin counters. DRA reached GA in 1.34, and in 1.37 the ResourceClaim .status.devices field graduated to Stable, which lets drivers report device-specific status. Read the DRA concept pages before designing GPU scheduling.
A seven-day rollout plan
Day | Do this | Done when |
|---|---|---|
1 | Run | You know which of the four 2026 changes apply to you |
2 | Apply steps 1 to 3 and 5 to 6 in a staging namespace |
|
3 | Install a Gateway API controller in staging, create | The HTTPRoute shows |
4 | Apply step 7 and add one egress rule per real dependency | A blocked test Pod fails and the app still works |
5 | Enable | The Pod starts and passes its probes |
6 | Drain a node and watch the PodDisruptionBudget | Drain completes with at most one |
7 | Upgrade staging one minor version, then production | Both clusters pass |
Closing notes
None of this requires a new tool. Every manifest above is plain Kubernetes, and every feature labeled Stable is safe to depend on. Features labeled Beta, such as HPA scale to zero, are on by default in 1.37, but they can still change, so pin your cluster version and read the release notes before each upgrade.
Managed Kubernetes services often lag the upstream release by weeks or months. Run kubectl version and check your provider's release calendar before assuming a 1.37 feature is available to you. The upstream support dates in this post apply to the community project, and your provider may set its own.
Comments (0)
Join the discussion securely with your ZyVOP account.
Continue on ZyVOP to CommentNo comments yet. Be the first to comment!