InfraNullBook a readiness review →

Recovery readiness

Your backups aren’t tested until you’ve restored them: a restore drill for EKS.

A successful backup job says that a backup operation completed. It does not say your application can be recovered. An EKS restore can fail because of missing volume data, inaccessible encryption keys, stale credentials, unavailable images or dependencies outside Kubernetes. The only useful way to test the whole path is to restore and validate the application under controlled conditions.

This guide is for teams planning a restore drill, not a promise that one tool covers every workload. InfraNull’s early-access platform does not claim to provide backups. A readiness review can inspect backup configuration and existing recovery evidence; a drill that creates resources or restores data is a separate, explicitly authorized activity.

Define the failure you are rehearsing

Start with a scenario your team needs to handle. An accidentally deleted namespace is different from losing a cluster, an availability zone or access to an AWS account. Identify what survives in the scenario and what must be recreated. A drill that keeps using the original cluster’s secrets or network may not test the failure you intended.

Set a recovery point objective and a recovery time objective with the application owner. The first describes acceptable data loss; the second describes acceptable time to restore useful service. They are goals to test, not numbers to claim in advance. Make the clock boundaries clear: does recovery time include discovering the problem, provisioning infrastructure and changing traffic?

Agree the drill scope, participants, data classification and budget. Restoring production data into a less protected environment can create a new risk while testing an old one. Use redacted or representative data where appropriate, and preserve the access boundaries required by the data you actually restore.

Map everything the application needs

Kubernetes object backups and persistent data are different layers. A Deployment manifest describes how to start pods; it does not contain the database state on a volume. Inspect what your backup mechanism captures, what it excludes and how application consistency is achieved. A completed volume snapshot may require application-specific recovery steps.

List data outside the cluster, including managed databases, object storage and queues. Include images, registry access, certificates, DNS, IAM roles, external secrets and encryption keys. Kubernetes resources may reference those services without backing them up. Restoring the resource objects can succeed while the application remains unable to connect.

Document infrastructure creation separately from workload restoration. A new EKS cluster may need networking, access configuration, add-ons and storage drivers before restored workloads can function. The AWS EKS documentation describes the managed control plane; it does not make your application data or external dependencies automatically recoverable.

Choose an isolated destination

Prepare a destination that is separate enough to avoid touching production by mistake. Use explicit account, Region, cluster and namespace identifiers in the procedure. Verify the active context before every restore command. Labels in a terminal window are not a substitute for checking the actual destination.

Prevent restored workers from performing production side effects. A recovered scheduled job could send emails, charge a customer or consume a live queue if it inherits the original configuration. Disable or redirect integrations intentionally, and make the test endpoints visible to everyone involved. Isolation must include application behavior, not just a new namespace name.

Confirm capacity, storage classes, network reachability and key permissions. A snapshot may be present but unusable if the destination cannot access its encryption key or if a storage driver cannot provision the required volume. For workloads on EBS-backed storage, consult the EKS EBS CSI driver guide and validate the specific restore mechanism you use.

Write a procedure another person can follow

The runbook should identify the backup to restore, prerequisites, commands or tooling, expected intermediate states and validation steps. Record tool versions and the source of configuration. Avoid relying on one engineer’s shell history or access to a laptop. The test is stronger when someone other than the author can execute the documented path.

Identify ordering constraints. Custom resources may need their CRDs and controllers available; workloads may need storage and secrets before becoming useful. Restoring every object at once does not guarantee those dependencies resolve safely. Follow the backup tool’s documented procedure and record any manual intervention required for your application.

Capture the start time and retain useful evidence as the drill proceeds. Save status outputs, errors and decisions without copying sensitive payloads into a report. An unexpected manual step is a finding: it belongs in the runbook and in the timing, even if an experienced operator solves it quickly.

Validate the data, not only the pods

Check the application paths that establish useful recovery. Can a user authenticate, retrieve an expected record, write new data and see that write through the normal interface? Do background jobs resume correctly? Are reads and writes consistent with the selected recovery point? Agree these checks with the people who understand the application’s behavior.

Use records or test markers that let the team distinguish restored state from newly generated state. Validate important relationships and application-level invariants, not just row counts. A database process that starts successfully may still be missing recent transactions or data stored elsewhere. Record the observed recovery point and compare it with the objective.

Check observability too. Logs, metrics and alerts should reveal the recovered system’s health. If the application works but the team cannot detect failures in the new environment, operational recovery is incomplete. Keep the drill’s test traffic and alerts distinguishable from production so the exercise does not confuse incident responders.

Test traffic and recovery decisions

If the scenario includes changing traffic to a replacement cluster, rehearse the routing procedure within the authorized scope. Consider DNS caching, ingress configuration, certificates and client behavior. A local request to a pod does not test the public path that users depend on. Document what remains untested if the exercise stops before a real traffic switch.

Define decision points and cleanup responsibilities. The team needs to know when to stop a failing drill, how to leave production unaffected and who owns the temporary resources. Retained snapshots, volumes, load balancers and clusters can continue to cost money after the test. Confirm cleanup without deleting the evidence or backups still required by policy.

Turn the drill into evidence and follow-up work

Compare measured recovery time and observed data loss with the objectives. Record what passed, what failed, what was not tested and which manual steps were necessary. The report should state the scenario and assumptions so nobody later interprets a namespace restore as proof of full account recovery.

Assign owners to the gaps and repeat the relevant portion after fixes. Major changes to storage, application architecture or cluster provisioning can invalidate old evidence. Recovery is a capability your team maintains, not a checkbox completed once. A recent, representative drill is more useful than an unsupported claim that backups are “working.”

If the backup inventory or runbook is unclear, the EKS Production Readiness Review can identify configuration and evidence gaps for $1,500 over an agreed one-week, read-only scope. It does not execute a restore or certify recovery. You can also use the upgrade checklist to connect recovery preparation with an upcoming cluster change.

A clear next step

Know what needs attention.
Then decide what to change.

A one-week, read-only EKS Production Readiness Review. $1,500. A prioritized plan your team can use.