When I started planning my cloud computing capstone, I kept coming back to the cost of cloud hosting. I understood the appeal of moving workloads into the cloud, but I was interested in what a small business could do when a full migration wasn’t the right fit.
It was easy to think of cloud adoption as an all-or-nothing decision. Keep the servers on-premises or move everything to AWS. I wanted to explore the space between those choices. Could a business keep its everyday workloads where they were and use the cloud strategically to protect its ability to operate?
That became the idea behind my capstone: run the application on-premises, keep its backups offsite, and create a recovery environment in AWS when it was needed. I wanted to demonstrate a path toward the kind of resilience businesses associate with enterprise infrastructure, at a cost that made sense for a much smaller organization.
I had an idea of how the pieces should fit together. Getting them to work would require something I had studied but had never actually used... Terraform.
A business problem worth solving
To give the project a purpose, I created Bayou Outdoor Supply, a fictional Gulf South distributor whose order system depended on a single on-premises server. It was a capstone scenario, but the problem was concrete: lose the server, and the business loses access to the application and data it needs to process orders.
That could happen because of failed hardware, an extended power outage, or a disaster affecting the site. A backup stored alongside the original system would offer little comfort if the same event took out both.
The question became how to give that business a practical way back. A second physical server would still need maintenance and could share the same site risks. A permanently running cloud environment would add recurring costs even during the long stretches when it wasn’t needed.
I wanted the cloud to serve a specific purpose in the design. It would hold the offsite backups during normal operation, then provide the infrastructure to recover the application if the original environment was unavailable.
That decision gave the project its shape. The next step was to build something I could actually lose and restore.
Giving the recovery plan something to recover
I built a small order application using Flask and PostgreSQL, running in Docker containers on an Ubuntu virtual machine in my Proxmox environment. The application gave me a way to create orders and retrieve them again, which would matter later when I needed to prove that recovery had worked.
With the primary environment running, I set up a scheduled database backup. Every hour, the job exported PostgreSQL, compressed the backup, and uploaded it to Amazon S3.
Now the database had a copy outside the Proxmox host. The hourly schedule also defined a tradeoff: recovery would use the latest successful backup, so changes made after that backup could be lost. The design aimed for a one-hour recovery point objective, assuming the backups completed as intended.
That was an important part of the business decision. A company would need to decide whether that potential loss was acceptable or whether it needed a more demanding design.
For this project, hourly backups gave me a clear starting point. I could see the backup objects in S3, but there was still a large gap between having a SQL dump and having a working order system.
Closing that gap was where Terraform entered the project.
From understanding Terraform to using it
Terraform was the part I was most uncertain about, and the part I was most excited to try. I understood what infrastructure as code meant. I had studied how Terraform describes resources and provisions them through a provider. What I hadn’t done was use it to build an environment of my own.
This project gave that learning a job to do.
I needed AWS networking, an EC2 instance, and permission for that instance to retrieve the backups. Terraform defined the recovery infrastructure, including the VPC, subnet, routing, security group, IAM role, and instance profile.
The instance profile let the recovery server read the backup bucket without placing an AWS access key in its bootstrap script. The access was limited to the bucket operations needed for recovery.
Provisioning the server was only the beginning, though. Once it existed, something still had to turn it into a useful recovery environment.
I used cloud-init and a bootstrap script to install the dependencies, start the Docker Compose application stack, retrieve the newest S3 backup, and restore PostgreSQL. That connected the infrastructure definition to the application and its data.
The recovery procedure started with an operator running terraform apply. From there, the provisioning and restoration steps were automated. I wasn’t building automatic disaster detection or instant failover; I was building a recovery process that someone could deliberately initiate.
Once those pieces were connected, I could finally test the question that had started the project.
Bringing the orders back
Before the simulated outage, I confirmed that the on-premises database contained three sample orders. They gave me something specific to look for on the recovery side.
I then simulated losing the primary environment and initiated recovery in AWS. This was the part I had been looking forward to: seeing whether the Terraform configuration and bootstrap process would produce the working environment I had planned.
The controlled test took 3 minutes 55 seconds from starting terraform apply to successfully validating the application’s health endpoint.
I checked the containers and queried the orders endpoint as well. The original three records were there. The application had come back with its data.
That was the result the project needed. A running EC2 instance would have shown that provisioning worked. Seeing the original orders come back showed that the recovery path had carried the data through, too.
I also destroyed and re-provisioned the AWS infrastructure to check that I could repeat the process. The more useful the configuration became, the more I appreciated having the recovery environment described in files I could review and use again.
What the test proved, and where I go next
The successful test gave me evidence that this small application could be reconstructed in AWS from its offsite backup. The measured time belonged to that lab run, starting when recovery was initiated. A real outage would also include recognizing the failure, deciding to recover, and directing users to the recovered service.
There was more work to do before this could serve a real business. The demo needed stronger secrets handling, TLS for the application, monitoring, and a plan for DNS failover. It also needed stricter restore-failure handling and automated checks that confirmed the data had returned before declaring success.
Those improvements followed naturally from the test. Once I knew the basic path worked, I could see more clearly what would make it dependable in ordinary operation and during a crisis.
The cost logic remained what had drawn me to the project. S3 storage and requests still have ongoing costs, and recovery or testing incurs compute and associated charges. The design avoids keeping a duplicate recovery compute environment running continuously. Whether that makes financial sense for a particular business depends on its workload, recovery requirements, and existing infrastructure.
For my capstone, it let me demonstrate a more selective use of the cloud: preserve the on-premises workload, protect the data offsite, and provision recovery capacity when needed.
It also turned Terraform from a concept I understood into a tool I had used to solve a problem. That first successful deployment was exciting, and it gave me a reason to keep going. Now that I’ve graduated, Terraform Associate is the first certification I’m pursuing.
I started this project wondering how a small business could get more value from the cloud without moving everything into it. By the end, I had a tested recovery path, three restored orders, and a new skill I wanted to develop further.

