4 minute read

When we are thinking about disaster recovery, a complex picture forms in our heads. Many data centres, in many locations, so our product can reach customers even in the worst of times. Database replicas and streaming backups. File storage redundancy, and everything else under the sky. At least this should be the goal, but in most cases it’s an afterthought and usually done for compliance reasons. Being a small startup struggling to get paying customers, you have different priorities. Like doing sales and product research. But losing your app data because of a poor backup strategy is a big no-no and in almost all cases a product killer.

Disaster recovery also doesn’t have to mean having a super complicated setup. A simple database/storage backup and restore strategy is usually enough. You can define your infrastructure in code and set it up to be reproducible anywhere at any time. Imagine building your infrastructure to be antifragile from the start . Benefiting from economic benefits from it in the long run. When you think about it, even the compliant disaster recovery solution is overlooking the elephant in the room. Which is allowing for everything to spread out on the same cloud provider. Sometimes the only thing we are looking for is to check off that dreaded box. And claim we can theoretically recover if the primary data centre goes down. Being paranoid as I am, it’s hard to say I’m satisfied with what passes as resilience nowadays. In this case we can both have our cake and eat it too. We can have complete control of the whole stack, and the ability to run it anywhere we want at any time we want. Accepting some simple principles from the beginning will make your infrastructure super resilient. And it won’t be hard to maintain as well.

First, I want my applications to be able to run anywhere and everywhere. Luckily we have Docker now that actually does this instead of Java that claimed to do this. This might prevent us from using some of the faster options that cloud providers offer. But you can spin Docker containers even on AWS Lambda so it doesn’t matter that much.

With hope, a true disaster will never happen to you and your product. But if it does, you can recover wherever you want whenever you want. I guess I don’t have to mention that you have to have full control of the off-site backups. You have to take many geographic regions, countries and political blocs into account. A company based in a bloc that is not the one where you live (i.e. US cloud in Europe or any other combo) is a liability, big or small. If you are an EU company, it is in your best interest to be on a completely EU-owned cloud/hosting provider. There are multiples to choose from, depending on what your business needs are. If you are a US-based company, the same thing goes for you. We shouldn’t put all our eggs in one basket. For the ultimate peace of mind, we should be able to run on all of them with very little effort (if our DPA permits it of course). This way we avoid vendor lock-ins which might look good in the start, but they increase your risk in the long run.

This approach allows you to migrate to a different provider with near zero downtime (within whatever your RTO states). This will not only allow you to have a quick solution to recover from almost any disaster but will also produce an economic benefit to you. Imagine your cloud provider increasing their prices or neglecting the hardware for a few years or even a decade in some cases. If you are a loyal customer, you might be paying for shitty old racks at double/triple the price of new more performant ones. And there might be ways of saving a lot of money on server costs alone (depending on your infrastructure size of course). Think of it like changing mobile providers every two or three years. Who wouldn’t want to get the new customer signup bonus which drops your costs 50% for a year or two. Acquisition departments often have better budgets than retentions ones and we can use that to our advantage. Since savings are all about percentages here, the more you are spending now, the more you can save at another provider.

I’m aware that switching to another cloud provider is not the same as switching to a new mobile provider. But the downtime was longer when I switched mobile providers than with cloud migrations. Sure, the number of affected users is 1 compared to all the people using your service. You can afford 15 minutes of migration downtime every other year given it saves you thousands (or hundreds of thousands) per year. Most likely it will result in better performance for your end users as well so it’s a win-win for everyone. Except the person doing the migration on a Sunday morning of course, but I love doing these things and looking at the results afterwards.

If you want to know when I write posts like this in the future and find out more, subscribe to my newsletter below. Go here to book a free consultation call where we can go over your concerns: https://schedule.berislavbabic.com/30-minute-meeting

Comments