InfraNullBook a readiness review →

One week · Read-only · $1,500 USD

An EKS health check you can act on.

Find the production risks that deserve attention before you spend time on fixes. The EKS Production Readiness Review turns cluster configuration and operational evidence into a prioritized plan.

What we inspect

Upgrade readiness

Kubernetes support dates, deprecated API exposure, add-on versions, node groups or Karpenter configuration, and the dependencies that can block an update.

Security & access

IAM and cluster access boundaries, endpoint exposure, workload permissions and available logging. We flag risks and evidence gaps rather than claim a complete security audit.

Recovery & operations

Backup configuration and restore evidence, disruption budgets, observability coverage and control-plane cost exposure. Configuration alone does not prove recovery.

What you receive

The report connects each recommendation to evidence from the agreed scope. Your team can use it to plan work, explain priorities and decide where further testing is needed. It includes these sections:

  • Scope and inventory. Clusters, accounts, Regions, Kubernetes versions, support deadlines and the components inspected. We record the collection date and access limitations so readers know which environment the report describes and where coverage stops.
  • Prioritized findings. Each finding describes the observed state, supporting evidence, likely operational impact and recommended next action. We distinguish confirmed issues from questions requiring application owner input. Priority reflects urgency and consequences, rather than a generic score that treats every configuration difference as equally important.
  • Upgrade readiness plan. Deprecated API exposure, add-on and controller dependencies, node provisioning considerations and suggested preflight checks. The plan identifies blockers to resolve before a version change and checks to perform afterward. It is a preparation document; production execution needs a separately agreed change plan.
  • Security and access observations. Findings on the inspected IAM and Kubernetes access boundaries, endpoint configuration, workload permissions and logging coverage. We describe the evidence behind a concern and the scope of the check, without presenting the review as a penetration test or comprehensive security certification.
  • Recovery and operations evidence. Backup configuration, available restore records, disruption safeguards and observability gaps. A configured backup and a demonstrated restore appear as different kinds of evidence. Where recovery time or data loss has never been measured, the report states the uncertainty and proposes a drill to address it.
  • Cost exposure and action register. Control-plane support charges with explicit assumptions, followed by a practical sequence of recommended work. We discuss owners, prerequisites and verification criteria during the walkthrough, so a team can turn findings into tickets rather than start another investigation from scratch.

What a week looks like

The one-week review starts once the agreed access and evidence are available. Before booking, we confirm cluster count, account boundaries and the questions your team needs answered. If the environment requires a larger scope, we resolve that before the engagement rather than compressing an incomplete inspection into the same promise. The schedule below shows how the review moves from context to verified recommendations.

  1. Day 0 — Kickoff & access

    Agree the cluster inventory, business-critical workloads, upcoming deadlines and known concerns. Your team approves a scoped read-only role and confirms the access expiry or revocation process. We also identify the application contact who can explain deployment behavior and existing recovery procedures. Share relevant runbooks and previous test records through the agreed channel; credentials and production data are not a substitute for scoped access.

  2. Days 1–3 — Evidence collection

    Inspect cluster configuration, support status, add-ons, node provisioning, access boundaries and the operational evidence in scope. Compare observed API usage with available workload manifests and record the observation window. Review backup settings alongside actual restore records, and examine disruption budgets in the context of replicas and scheduling constraints. Missing logs, inaccessible resources or infrequent jobs become explicit coverage gaps, with questions for your team rather than assumed answers.

  3. Days 4–5 — Analysis & engineer verification

    Connect the evidence to upgrade blockers, availability risks and cost exposure. An engineer checks every finding and its proposed action before delivery. We ask targeted questions where application context could change the conclusion, such as whether a single-replica workload is intentionally non-critical. AI assistance requires your written consent for the intended tools and data. Recommendations state what to verify after a fix and which decisions need a separate test or implementation scope.

  4. Days 6–7 — Report & 60-minute walkthrough

    Receive the written report and join a 60-minute walkthrough with an engineer. Discuss the highest-priority findings, the evidence behind them, dependencies between actions and the questions still open. Agree proposed owners and next steps, then revoke review access. Your team can implement the plan internally or request a separate quote for upgrade assistance; the review does not create an obligation to buy implementation work.

If access or evidence arrives late, we explain the effect on coverage and timing. An unavailable restore record remains an evidence gap; it does not become a claim that recovery works merely because the review week has ended.

Understand the access model · Compare the offers

Clear boundaries, useful decisions

This engagement is read-only. It does not include production remediation, a penetration test, compliance attestation, a restore drill or ongoing incident response. Exercises that create resources, alter workloads or handle production data need their own authorization and scope. The review can identify where those exercises would answer an important question, but it cannot claim their results.

Bring a concrete decision to the kickoff: whether an upgrade is ready to schedule, why support costs have increased, or which recovery assumptions deserve testing. We use that context to make the findings actionable while preserving the agreed inspection boundaries. The result is a documented starting point for engineering work, with uncertainties visible and next steps tied to evidence.

Questions before you start

What does $1,500 include?

A one-week, read-only EKS Production Readiness Review, a written report with prioritized findings, an upgrade readiness plan and a walkthrough. Cluster count and account boundaries are agreed before booking so the scope is explicit.

Will you change production?

No. The review uses scoped read-only access. Remediation, upgrades and restore exercises that create resources require a separate scope and customer approval.

Does this certify our security or prove recovery?

No. This is a practical technical assessment, not a certification, penetration test or compliance attestation. We inspect recovery evidence and identify gaps; a successful restore requires a separately authorized drill.

Can our team implement the recommendations?

Yes. The report is designed to support your own team. EKS upgrade help can be quoted afterward, with no obligation to purchase it.

A clear next step

Know what needs attention.
Then decide what to change.

A one-week, read-only EKS Production Readiness Review. $1,500. A prioritized plan your team can use.