Skip to content

10s → 0s

Zero-downtime ECS deploys

Fixed the ALB and shutdown timing behind an outage on every release.

Organisation
TransFi
When
2024 - now
Areas
Reliability, AWS, ECS Fargate

Problem

Every deploy to ECS Fargate caused a recurring outage of about 10 seconds. For a payments API, that means failed requests on every release.

TODO(shubham): How it was noticed, and how often you deployed.

Constraints

TODO(shubham): Constraints: no maintenance windows, deploy frequency, services involved.

What I built

I redesigned the release path so the load balancer, the shutting-down task and in-flight connections are coordinated.

Safe task replacement on deploy
  1. New task passes tuned health checks
  2. ALB deregisters the old task
  3. Connections drain
  4. App shuts down gracefully
  5. Old task stops
  • ALB target deregistration timing.
  • Task-termination timing.
  • Connection draining.
  • Graceful shutdown in the application.
  • Tuned health checks.

TODO(shubham): The root cause in your words, and how you verified it.

Outcome

  • The ~10s outage on every deploy went to zero.

What I'd do differently

TODO(shubham): What you'd change with hindsight.