Moving a TV AdTech platform from on-premise to AWS
A TV AdTech company
We worked with a TV AdTech company that was still running its infrastructure on bare metal server in an on-prem setup.
That setup had worked for a long time, but adding capacity would take weeks and the infrastructure team was spending a lot of time keeping the on-prem environment running. We worked with the infrastructure and backend teams to plan the move to AWS and take it one environment at a time.
Weeks → minutes
to provision a new environment
99.5%
uptime SLO maintained
Code
used to create infrastructure instead of manual setup
Why was the old setup difficult to run?
There was no virtualisation layer. If the team needed more compute, they had to provision more physical machines, which could take weeks.
Some of the hardware was also getting old, which was starting to show up in performance. To stay on the safe side, the team sometimes compensated by provisioning more capacity than they really needed.
There was a lot of operational work around the setup too. The team was bogged down maintaining out-of-date services and self-hosted tools, while also dealing with environments that weren't properly isolated from each other.
Why couldn’t we just copy-paste the same setup to AWS?
The existing platform had been built around physical servers, so copying it one-for-one in AWS would’ve just carried a lot of problems over.
We sat down with the infrastructure and backend teams and went through the migration plan together. Data migration needed more thought. So did rollback. We also had to be clear about what would move first and what could wait.
So instead of treating AWS like a new place to host the old setup, we used the migration to change how the infrastructure was put together.
What did we need to set up in AWS first?
We set up AWS Organizations with separate accounts for production, staging and development, along with central logging and audit accounts and a shared platform layer.
We used Terraform and Atmos to build the infrastructure. We built reusable components for things like VPCs, EKS, RDS, S3, KMS, Secrets Manager and ALBs.
We also kept AWS console access read-only. If infrastructure needed to change, the change went through code, a pull request and a Terraform plan before it was applied.
How did we get the data across?
The data the application services depended on had to be moved to AWS first.
The source data was in Ceph clusters. We worked with the data engineering team on a custom sync process that moved it into S3.
That wasn’t completely smooth. Some consistency issues in the Ceph clusters were causing the sync to fail, so we fixed those first. We also set up the database clusters in AWS and synced the data into RDS before the application services moved over.
How did we move the different environments?
We started with a lower environment because there was simply less to worry about there. There was no production data to migrate and fewer external integrations.
That gave us a chance to catch dependency changes early and sort them out before staging and production moved.
Staging added external integrations, test data and UAT requirements.
Production was a different beast. We had to account for connectivity back to the on-prem setup, data migration, API Gateway changes and external IP allowlists.
We did take on a few bits of technical debt to keep the migration going, but it was tracked so the team could come back to it afterwards.
What did we change in deployment and security?
We used GitHub Actions for CI and ArgoCD for deployments. ArgoCD also gave developers a better view of rollouts and application logs, which helped once teams started deploying into the new environment.
We also tightened up security. We removed static AWS credentials where they existed and switched those workloads over to AWS roles. We put WAF in front of public endpoints and moved secrets into Secrets Manager. For the AWS side, we used GuardDuty, CloudTrail and Security Hub, and VPC Flow Logs gave us a view into network activity.
What happened after the migration was done?
Provisioning a new environment went from weeks to minutes.
We also got to a point where infrastructure was being created through automation instead of manually. The platform maintained a 99.5% uptime SLO without breaching its monthly error budgets.
The infrastructure team didn’t have to plan weeks ahead for physical capacity anymore. They could add capacity when they needed it, and spend less time maintaining the on-prem setup and the tooling around it.
Want to make your infrastructure work harder for your team?
Tell us what your engineers are wrestling with.
