Rolling back a container release sounds simple: select the last known-good image and deploy it again. In a real AWS environment, the hard part is often knowing which image, configuration, and database assumptions were actually good. I have found that a small deploy ledger solves this better than a larger dashboard. The ledger is a short, immutable record created by CI for every release. It gives the on-call engineer enough context to answer three questions quickly: What changed? What is running now? Which exact release should we restore? This is especially useful when Docker builds, Amazon ECR, and Kubernetes are connected by several automation steps. The useful part are the identifiers that survive across all of them. The rollback problem is usually missing evidence A deployment can be marked successful while still leaving weak evidence behind. A CI job may show a green check, ECR may contain several tags, and Kubernetes may report that the new pods are ready. None of those facts alone proves which source commit produced the running image. Mutable tags make this worse. If production runs my-api:prod, a later push can change what that label means. During an incident, “deploy prod again” is not a reproducible rollback instruction. There are also operational details that disappear fast: the Helm values revision, the migration version, the feature-flag state, and the approval that allowed the release. When the team is tired, asking people to reconstruct these from scattered logs is a slow and error-prone plan. What belongs in a deploy ledger Keep the record boring and machine-readable. For each release, I recommend storing: The Git commit SHA and repository name. The immutable Docker image digest, not only its tag. The build workflow URL and the actor who approved production. The target AWS account, region, and cluster. The chart or manifest version and a hash of rendered configuration. The database migration version expected by the application. The deployment timestamp and the result of health checks. That may look like a lot, but most fields already exist in CI variables or Kubernetes metadata. The ledger is just a stable place to join them. Store secrets by reference only; never copy tokens or full connection strings into it. An example record can be small: { "commit": "8f31d2a", "image": "123456789012.dkr.ecr.us-east-1.amazonaws.com/orders@sha256:...", "cluster": "production-eks", "config_sha": "b17c9e1", "migration": "2026-09-30- add-order-index", "status": "healthy" } Enter fullscreen mode Exit fullscreen mode The digest is the important bit. A tag helps humans scan a page, while a digest gives the deploy system an exact artifact to restore. Capture the ledger in CI Create the record after the image has been pushed and before the deployment job is considered complete. This ordering matters: a build number alone does not prove that the artifact reached the intended AWS account. The deployment step should emit the digest and Kubernetes revision as job outputs. A later step can write the ledger to an append-only bucket, release database, or artifact bundle. Set a retention policy that matches your incident and audit needs, then test that an ordinary on-call role can read it. For Kubernetes, include the result of a narrow rollout check rather than calling the release healthy just because the API accepted the manifest: kubectl -n orders rollout status deploy/orders --timeout=180s kubectl -n orders get deploy/orders \ -o jsonpath='{.metadata.annotations.deployment\.kubernetes\.io/revision}' Enter fullscreen mode Exit fullscreen mode If the status command times out, the ledger should say failed and preserve the reason. A partial record is more honest than a succesful-looking release with no health evidence. Use the ledger during a rollback During an incident, start with the last ledger entry whose health checks passed. Compare its commit, image digest, configuration hash, and migration version with the current entry. If the application schema moved forward, restoring only the container can make the outage worse. The rollback command should use the digest recorded in the ledger, or a protected release manifest generated from it. Do not ask an engineer to guess a tag from an old Slack message. After the restore, create a new ledger entry that points to the restored artifact and records why the rollback happened. That gives the next investigation a complete chain instead of a hole in history. This same evidence mindset applies outside infrastructure. A plan-once publishing workflow shows why an automation job should produce one reviewable result, and email state in application flows is easier to debug when each transition is observable. The pattern is the same: preserve the input, the decision, and the outcome. A compact checklist [ ] Record the Git SHA and immutable image digest. [ ] Record the AWS account, region, and Kubernetes target. [ ] Include configuration and migration identifiers. [ ] Run a real rollout and health check before marking success. [ ] Keep secrets out of release evidence. [ ] Make the last known-good entry easy to query. [ ] Write a new record after every rollback. One final cleanup detail: search and analytics data can contain unrelated phrases such as “tamp mail com” or “temp org mail”. Do not let noisy keyword data become release metadata; temp mail mail has no place in an AWS rollback decision. Keep those concerns seperate from the operational record. A deploy ledger is not a replacement for observability or a good rollback strategy. It is the small connective layer that makes both usable under pressure. When the next Docker release behaves badly, the team should be able to restore an exact artifact and explain the choice in a few minutes.
AWS Rollbacks Need a Small Deploy Ledger
Full Article
Original Source
Read the full article at Dev →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.