Cutting AWS Costs with Right-Sizing and Spot Instances

Cloud cost problems are architecture problems. Where the savings actually are: right-sizing, spot capacity, and the observability to keep both honest.

Cutting AWS Costs with Right-Sizing and Spot Instances
Tomislav Pree
June 16, 2026
/
Articles

Costs are an architecture signal

When I get called in to look at a runaway AWS bill, I don't start with the cost explorer. I start with the architecture, because an oversized fleet is almost never a pricing problem. It's a design assumption from six months or two years ago that nobody revisited: an instance type chosen for peak load that never materialized, a database sized for a migration that got postponed, a service that scaled out instead of up because that was easier at the time and nobody went back to check whether it still made sense. Treating cost as an architecture signal rather than a finance problem changes where I look first. A consistently underutilized instance tells me something about how the workload was actually specified versus how it actually behaves.

Right-sizing with data

I never resize on intuition. Before touching an instance type, I pull real utilization data (CloudWatch metrics at minimum, Prometheus where I have finer-grained application metrics) over a window long enough to capture actual peak behavior, not just a quiet afternoon. CPU and memory utilization tell you the obvious story, but I also look at network throughput and, for anything stateful, disk I/O, since those can be the actual constraint even when CPU looks comfortably low. Right-sizing isn't a one-time exercise either. I treat it as an iterative loop (resize, observe the new utilization pattern under real load, adjust again), because workloads drift, and a fleet that was correctly sized at launch quietly stops being correctly sized as traffic patterns change.

Spot where it's safe

Spot instances are the single largest lever I have for compute savings, but only for workloads that can tolerate interruption without breaking anything. Batch processing jobs, CI runners, and stateless inference behind a load balancer are the clearest fits. If an instance gets reclaimed, the job retries or the load balancer routes around it, and nothing user-facing notices. What I don't put on spot is anything stateful or latency-critical where an interruption has a real cost: a primary database, a session-holding service, anything where "just retry" isn't actually free. The savings on spot are large enough that it's worth deliberately identifying every workload that qualifies, rather than defaulting to on-demand everywhere out of caution.

Guardrails

None of this stays correct without guardrails, so I set budgets with alerts at multiple thresholds rather than one to catch a runaway spend before it's a surprise at month-end, and I keep cost dashboards broken out by service so a spike is traceable to a specific team or workload within minutes, not a multi-day investigation. Tagging discipline is what makes that attribution possible in the first place, every resource tagged by owner and environment, enforced at provisioning time rather than hoped for after the fact.

Takeaways

  • Treat an oversized bill as an architecture question first, a pricing question second.
  • Right-size using real utilization data, and repeat the exercise as workloads change.
  • Move interruption-tolerant workloads (batch, CI, stateless inference) to spot capacity, and keep stateful and latency-critical work off it.
  • Budgets, dashboards, and enforced tagging turn cost visibility into something you can act on quickly.