10s → 0s
Zero-downtime ECS deploys
Fixed the ALB and shutdown timing behind an outage on every release.
- Organisation
- TransFi
- When
- 2024 - now
- Areas
- Reliability, AWS, ECS Fargate
Problem
Every deploy to ECS Fargate caused a recurring outage of about 10 seconds. For a payments API, that means failed requests on every release.
TODO(shubham): How it was noticed, and how often you deployed.
Constraints
TODO(shubham): Constraints: no maintenance windows, deploy frequency, services involved.
What I built
I redesigned the release path so the load balancer, the shutting-down task and in-flight connections are coordinated.
- New task passes tuned health checks
- ALB deregisters the old task
- Connections drain
- App shuts down gracefully
- Old task stops
- ALB target deregistration timing.
- Task-termination timing.
- Connection draining.
- Graceful shutdown in the application.
- Tuned health checks.
TODO(shubham): The root cause in your words, and how you verified it.
Outcome
- The ~10s outage on every deploy went to zero.
What I'd do differently
TODO(shubham): What you'd change with hindsight.