BuildRoulette
Back to Decision Records
Decision Records / The ASG that wouldn't scal...
DECISION RECORD (EDR) · AUTOSCALING · CLOUDWATCH · POSTMORTEM By Temi Nwosu · Principal Cloud Architect

The ASG that wouldn't scale down

2 min read 142 upvotes · 31 replies
#autoscaling #cloudwatch #postmortem

1. The Incident Context

For 11 consecutive months, our primary production Auto Scaling Group (ASG) serving API traffic never scaled down below 40 EC2 instances (c6i.2xlarge) during off-peak weekend hours. Traffic routinely dropped by 82% between 01:00 UTC and 06:00 UTC, yet the fleet remained pinned at peak daytime capacity.

This resulted in an estimated $14,200/month in unnecessary EC2 and EBS compute spend.

2. The Failure Mode & Diagnosis

A custom CloudWatch step scaling policy was configured with the metric AWS/EC2 CPUUtilization. However, the alarm statistic had been set to Minimum rather than Average during an initial provisioning terraform script.

Every 5 minutes, an internal log-shipping cron executed on a single worker node for 18 seconds, spiking that one instance's CPU to 22%. Because the CloudWatch Alarm evaluated Minimum(CPUUtilization) < 15%, and the cron ensured at least one instance stayed above 15% across interleaved sampling windows, the scale-in condition was never satisfied.

3. Architectural Decision & Fix

  1. Migrated from Step Scaling to Target Tracking: Configured Target Tracking with PredefinedMetricType: ASGAverageCPUUtilization set to 60.0%.
  2. Decoupled Cron Workloads: Extracted background cron tasks into serverless AWS Fargate scheduled tasks triggered via Amazon EventBridge.
  3. Capacity Anomaly Guardrail: Deployed a CloudWatch Composite Alarm that monitors RequestCountPerTarget vs GroupInServiceInstances to alert on capacity-traffic divergence.
new-background.jpg


Record Your Takeaway & Architectural Notes

You can come back and edit this later in your profile.