Reducing AWS costs for a hyperlocal delivery platform
A hyperlocal delivery company
We worked with a leading hyperlocal delivery company running a large microservices setup across several product verticals, with millions of transactions moving through the platform every day.
We worked with the engineering teams to understand where the AWS bill was going, bring it down, and make it easier for teams to keep track of their own spend.
$161k
saved over three months
20%
lower costs accounting for projected growth
45%
lower QA infrastructure costs
Why was the AWS bill getting harder to control?
The AWS bill was already around $800,000 a month, and it was still going up by 5–8% every month. The company had more products and more engineering teams than before, which also meant more infrastructure being added on AWS. Nobody had a clean view of what each team, service or product was actually costing.
Resources hadn’t been tagged consistently, so engineering and finance could see the total bill, but tracing that spend back to the teams and services behind it was much harder.
What did we find when we started digging into the bill?
We pulled six months of billing data into Cost Explorer and ran AWS Trusted Advisor against the account. That first pass with Trusted Advisor alone pointed us to around $30,000 a month in potential savings from things like idle RDS instances, underused load balancers and unattached EBS volumes.
There was more once we went through the resource inventory. Some EC2 instances had been running at under 5% CPU utilisation for 30 days. We also found unused Elastic IPs, S3 buckets that were no longer needed and redundant NAT instances.
We dealt with the easier fixes first. Backup retention was tightened, old snapshots and storage were cleaned up, and QA and development workloads moved to Spot Instances. That brought QA infrastructure costs down by 45%.
How did teams get a clearer view of their AWS spend?
We introduced consistent tags for team, service, product and environment. The Terraform CI/CD flow checked new infrastructure for those tags, and AWS Service Control Policies stopped untagged resources from being created. We tagged existing resources separately and used weekly reports showing anything that was still missing.
Teams could then see cost by service and team through Grafana and dashboards from the AWS billing partner. Budget alerts picked up overages, and AWS Athena gave people a SQL interface into the Cost and Usage Report data when they needed to dig further.
Team leads took responsibility for their AWS costs, while the platform team continued to own shared infrastructure.
Where did the bigger savings come from?
We checked which workloads could move to Graviton instead of forcing everything across. That brought down the overall compute costs by 9% and Production Reserved Instance coverage went up to 70% too.
EKS had another nice surprise for us. A lot of pods were asking for 2–4x more resources than they were actually using. We reset requests and limits around P95 utilisation and the same workloads could run with 20% fewer nodes.
We also found 23 over-provisioned RDS instances. Kafka was running at only around 40% of provisioned capacity, and moving compatible ElastiCache workloads to Graviton cut those costs by 15%.
S3 was another big one. We moved older data into cheaper storage classes and cleared out unnecessary data. That saved $19,000 a month.
We used VPC Flow Logs and Athena and found that around 70% of transfer costs were coming from cross-zone traffic. We changed how that traffic moved and added VPC endpoints. Data-transfer costs came down by 25% and the endpoint work alone saved around $8,000 a month.
We also moved Sentry in-house and shifted GitHub Actions runners onto EKS with Spot Instances. That cut runner costs by 60%, and local caching made them 3x faster.
How did cloud cost become something every team owned?
Team leads took responsibility for their own AWS costs, while the platform team continued to own shared infrastructure. Every two weeks, teams would look at what had changed and share anything useful they’d found along the way.
Cost also started coming up earlier in the engineering process. Estimates went into technical design docs and infrastructure cost was discussed during sprint planning.
What changed after three months?
Over three months, the work saved about $161,000. Once we accounted for the growth the AWS bill was otherwise on track for, that came out to a 20% cost reduction.
The client could also see AWS spend by team and service instead of looking at one big number. And none of the savings came at the cost of reliability or performance. In a few places, performance actually improved.
Want to make your infrastructure work harder for your team?
Tell us what your engineers are wrestling with.
