InfraNullBook a readiness review →

Upgrade planning

An EKS upgrade checklist for teams without a platform engineer

An EKS upgrade looks small when it is described as changing the Kubernetes version. In production, the work is a chain of compatibility and availability decisions: APIs, controllers, add-ons, compute, storage and application behavior. Teams without a dedicated platform engineer need a checklist that exposes those dependencies rather than hiding them under “update cluster.”

Use this guide to structure the work and assign owners. It is not a universal command sequence: the safe order depends on your current versions and provisioning model. Before the change, check the AWS EKS update guide and Kubernetes version calendar. Support dates here were verified October 7, 2026.

1. Record the starting point and destination

Inventory the cluster version, AWS account and Region, support policy, node versions, provisioning method and installed controllers. Record the versions of VPC CNI, CoreDNS, kube-proxy and CSI drivers, plus how each is managed. Include ingress, certificate management, monitoring and admission webhooks. A Helm release list is useful, but it may miss components installed by another mechanism.

Choose a target version with a supported dependency path. EKS control-plane upgrades proceed one minor version at a time. If the starting point is several releases behind, write down every step and the checks between them. Do not apply a node or add-on version intended for the final destination without checking compatibility with intermediate control planes.

Name the application owner and the person authorized to approve the production window. The platform work and application tests need to meet in one plan. A cluster update is not complete merely because the AWS status returns to active.

2. Find removed APIs and old clients

Inspect EKS upgrade insights and relevant API usage, then examine rendered workload manifests. Search infrastructure repositories, Helm charts and generated resources for APIs removed in the target release. Operators and controllers can create resources dynamically, so source manifests alone do not describe all calls made to the control plane.

Consider observation coverage. A job that runs monthly or a disaster-recovery script may not appear in a short metrics window. Check the code and release notes of clients that talk to Kubernetes, especially automation outside the cluster. A clean insight result is useful evidence, but it does not prove every rarely used integration is compatible.

Fix deprecated or removed API usage before the control-plane change where possible. Test the replacement manifests against representative workloads. Record which checks are based on observed traffic and which are based on source inspection so the team understands the remaining uncertainty.

3. Build an add-on and controller matrix

For each add-on, record current version, supported target version, prerequisites and the proposed update point. Check AWS compatibility information for EKS-managed add-ons. For Helm-installed components, use the maintainer’s supported Kubernetes matrix and upgrade notes. Similar names do not make two installation mechanisms interchangeable.

Pay attention to VPC CNI configuration, DNS behavior, kube-proxy and storage drivers. A cluster can accept an update while workloads later encounter networking or volume problems. Capture custom settings before updating so a reconciliation does not silently discard intentional configuration. Where the plan involves changing defaults, make that a visible part of testing.

Admission webhooks and operators deserve their own checks. A failing webhook can block new resources, and a controller may react differently after an API change. Validate certificate handling, endpoint reachability and resource reconciliation, not just whether the controller pod is running.

4. Plan node groups or Karpenter explicitly

Record whether compute is managed node groups, self-managed nodes or Karpenter-provisioned capacity. Check the supported node version and AMI path for the proposed control plane. Review bootstrap assumptions, image architecture and instance availability. Replacing a node also tests IAM permissions, image pulls, networking and storage attachment.

For Karpenter, review its installed version, supported Kubernetes versions and the NodePool and NodeClass configuration relevant to your installation. Controller and custom-resource changes can have their own migration steps. Do not substitute a generic managed-node-group procedure for a different provisioning system.

Test a replacement node before draining critical capacity. Confirm that workloads can schedule and become useful there. Track the final node inventory: a successfully updated control plane can coexist with old compute longer than intended if node replacement is omitted from the acceptance criteria.

5. Check disruption budgets and available capacity

PodDisruptionBudgets constrain voluntary disruptions; they do not create replicas or spare nodes. Inspect each budget alongside replica counts, readiness probes, topology constraints and termination behavior. A budget can stall a drain, while a missing or permissive budget can allow more disruption than the application tolerates.

Verify capacity for rescheduling before the window. Account for affinity rules, persistent volumes and limits in the relevant availability zones. A nominal spare node may not satisfy the workload’s constraints. Stateful applications require special attention to attachment timing, graceful shutdown and whether the application tolerates a member being replaced.

Agree how to handle a blocked drain. Disabling a disruption safeguard should be a deliberate, customer-approved decision with a understood service impact. It should not be the automatic response to a slow upgrade.

6. Test behavior and define acceptance

Use a representative non-production environment to exercise real application paths. Cover ingress, authentication, reads and writes, background jobs, DNS and storage. Check alert routing and telemetry so the team can see failures during the change. Pod readiness is an input to validation, not a complete application test.

Define acceptance criteria before execution. Include healthy workloads, successful business-critical checks, intended add-on and node versions, and no unexplained error trend. Specify who decides to proceed between steps and who communicates an interruption. Record test results rather than relying on a general impression that everything looked fine.

7. Write recovery limits in plain language

AWS announced EKS version rollback on July 1, 2026: eligible clusters can revert to the previous minor version within 7 days of an upgrade, at no extra charge, after readiness checks for API compatibility, version skew, add-on compatibility and cluster health. EKS Auto Mode rolls back worker nodes first. Rollback does not restore application data or automatically undo every add-on or manifest change. The bounded window is not a substitute for preflight testing; check eligibility in the EKS user guide. See the AWS announcement and EKS user guide.

Check backup and restore evidence before relying on recovery. A scheduled backup that has never been restored leaves important assumptions untested. Any drill that creates resources or uses production data needs its own authorization and boundaries. The restore-drill guide explains how to turn those assumptions into evidence.

8. Finish with ownership and support dates

As of this article’s October 2026 date, EKS 1.33 standard support has ended and 1.34 standard support ends December 2, 2026. Published rates are $0.10 per cluster-hour in standard support and $0.60 in extended support, according to AWS pricing. Cost is a planning input, not a reason to skip testing.

Assign owners to the remaining blockers and reserve a post-change verification window. If you need help establishing the plan, start with the $1,500 read-only readiness review. Upgrade assistance is scoped and quoted after the review, when the dependency chain is understood.

A clear next step

Know what needs attention.
Then decide what to change.

A one-week, read-only EKS Production Readiness Review. $1,500. A prioritized plan your team can use.