ECS autoscaling race condition with task protection

Autoscaling Killed Our Healthy ECS Worker. Here Is Why.

Our ECS worker was healthy, not crashing, not stuck — yet autoscaling killed it right as it picked up a new job. The fix? Inverting the ECS Task Protection lifecycle so the worker is protected whenever it’s ready to accept work, not just while processing.

June 20, 2026 · 5 min · 916 words · Ivan Yalovets
When Cost Optimization Breaks Production

When Cost Optimization Breaks Production: A Story of EFS, ECS, and Ephemeral Storage

We reduced EFS costs by 70% by moving temporary file operations to ephemeral storage. Everything looked perfect — until production started acting strange. Here’s what went wrong, why it happened, and how we built a proper decoupled architecture following AWS Well-Architected principles.

March 1, 2026 · 3 min · 611 words · Ivan Yalovets