Forget keeping all your files and documents on your own server โ thatโs not safe, lacks resilience. What you should do is chuck it into our network of web-connected buckets, which are totally resilient because theyโre located all over the world. Plus, itโll be cheaper and easier.
Thatโs the argument that convinced the corporate world to sign up to cloud computing, an idea with older origins that really started with the launch of Amazon Web Services in 2002 followed by the arrival of Elastic Compute Cloud four years later. Google and Microsoft both jumped on the bandwagon a few years after.
A decade and a half later, the vast majority of companies are using some sort of cloud, and itโs making AWS โ which holds 30% market share โ about $30bn in revenue and $10bn in income in recent quarters.
Anyway, cloud computing truly is a modern marvel. Except, of course, when it falls over. And thatโs what happened on Monday, when one AWS region โ a big one, taking in North Virginiaโs data centres โ had a DNS issue. Amazon swiftly fixed the issue, but it took nearly a full day for services to recover.
In the meantime, pretty much every company you can think of had outages of varying degrees, from Reddit to Slack and Roblox to McDonaldsโ app. Amazonโs own services fell over โ affecting Ring cameras and Kindle downloads โ and oddly so did Microsoft Teams and Office 365, which is hilarious given Microsoftโs Azure is AWSโ main competitor. So much for eating your own dogfood, Microsoft.
The incident raises plenty of issues: how can DNS still cause so many problems, why did one region falling over cause so many outages when surely these companies have backups in other geographies, and whatโs the point of the cloud if it isnโt resilient?
Concentrated risk, shared responsibility
Forrester principal analyst Brent Ellis noted that the incident โhighlights how concentration risk โ a dangerously powerful yet routinely overlooked systemic risk โ arises when so many companies across all industries become dependent on a single cloud provider and, more pertinently, a single region covered by that vendorโ.
But heโs not surprised that has happened. First, he pins it on false assumptions that bigger is better. โThereโs great appeal to using tech giants, but assuming they are too big to fail or inherently resilient is a mistake, with the evidence being the current outage and past ones,โ Ellis says.
Itโs easier to just dump it all on AWS or Azure or Google Cloud and let the problems sort themselves out. โConvenience often overshadows navigating the complex, nested dependencies in highly concentrated environments,โ Ellis says.
โThe entrenchment of cloud, especially AWS, in modern enterprises, coupled with an interwoven ecosystem of SaaS services, outsourced software development, and virtually no visibility into dependencies, is not a bug โ itโs a feature of a highly concentrated risk where even small service outages can ripple through the global economy.โ
And then Ellis raises the issue of AWSโ shared responsibility model. โAs for resilience promises, AWS directs customers to its shared responsibility model as a way to highlight where it takes ownership of service availability and what customers are responsible for. But when core services like DNS fail, even well-architected applications can become unstable,โ he adds.
โAWS works to fix its infrastructure, but many enterprises are left to wait until that is done even though they have followed recommended design patterns. This is not exclusively an AWS problem, but it has become a recurring issue, specifically for the US-East region, with customers left holding the bag when it comes to the impact of the outage.โ
Heart of the problem: DNS
Letโs start with the tech solution. DNS, or domain name system, is often referred to as the phone book of the internet, and perhaps we need to reboot this digital yellow pages, or at least add in some protections around it.
โThis particular outage exposes core issues with cloud resilience that stem from overreliance on services such as DNS, which were not architected for cloud-era technology demands,โ noted Forresterโs Ellis.
That said, itโs not a surprise, notes Dr Soohyun Jeon, Assistant Professor/Lecturer in Operation and Information Systems Management, Brunel University of London. โThis kind of configuration-level failure, while not unusual in large-scale distributed systems, can cascade rapidly because so many online services rely on shared cloud infrastructure,โ Jeon says.
โWhat begins as a localized fault can therefore manifest globally within minutes, impacting platforms ranging from gaming and social media to banking, telecommunications, and government portals.โ
What can be done?
Ellis advises enterprise tech leaders to add these two items to their to-do lists: โBuild the tools to increase technology systemsโ reliability, and address contractual gray areas related to shared responsibility models with cloud (and SaaS) vendors.โ
Jeon added that the incident is a timely reminder to strengthen resilience and continuity planning. โReliance on a single cloud region or provider introduces concentration risk, which can be mitigated through multi-cloud or multi-region redundancy strategies,โ he says. โIt also underscores the importance of transparency and shared responsibility in cloud governance.โ
Indeed, Nicky Stewart, Senior Advisor at the Open Cloud Coalition, argues that we need a better cloud market that isnโt dominated by AWS or anyone else โ something that UK regulators have examined. โIncidents like this make clear the need for a more open, competitive and interoperable cloud market; one where no single provider can bring so much of our digital world to a standstill,โ Stewart says.
Dr Corinne Cath-Speth, the Head of Digital at human rights organisation Article 19, agreed, telling The Guardian: โWe urgently need diversification in cloud computing. The infrastructure underpinning democratic discourse, independent journalism and secure communications cannot be dependent on a handful of companies.โ
And itโs not just about democracy, but our sovereignty, argues Cori Crider, Executive Director of the Future of Technology Institute: โThe UK canโt keep leaving its critical infrastructure at the mercy of US tech giants.”
Take back your files?
Thereโs another way: returning to the old ways. Last year, AWS told the Competition and Markets Authority that it faces competition from โon-premises ITโ โ in other words, cloud repatriation, when companies take their workloads out of AWS buckets and bring them home.
One example is SaaS firm 37 Signals, the maker of Basecamp, which decided it was tired of paying more than $3 million a year on the cloud and instead invested in its own servers, saving $10 million over five years โ and that includes the cost of shelling out for hardware.
Will this outage lead to a rush of repatriation? Perhaps not, as frankly if you canโt keep the lights on with AWSโ help, it may not be worth trying to achieve that on your own, and many smaller companies may well lack the tech skills or millions to invest in servers that made 37 Signalsโ shift possible.
Professor Oli Buckley, an expert in Cyber Security at Loughborough University, said the incident was a reminder to give resilience as much consideration as security. โIn short: yes, this is a big deal; yes, we should take notice; but also: no, this is not a reason to cancel your cloud migration plans overnight. It is, however, a reason to check how resilient your systems really are,โ he says.
Cloud computing may be a marvel, but even it has limitations, and itโs clear the promise of resilience canโt always be kept.
Three articles you may want to read next