AWS falls over, makes the world reconsider cloud computing – so what lessons should we learn from the AWS outage?

Forget keeping all your files and documents on your own server – that’s not safe, lacks resilience. What you should do is chuck it into our network of web-connected buckets, which are totally resilient because they’re located all over the world. Plus, it’ll be cheaper and easier.

That’s the argument that convinced the corporate world to sign up to cloud computing, an idea with older origins that really started with the launch of Amazon Web Services in 2002 followed by the arrival of Elastic Compute Cloud four years later. Google and Microsoft both jumped on the bandwagon a few years after.

A decade and a half later, the vast majority of companies are using some sort of cloud, and it’s making AWS – which holds 30% market share – about $30bn in revenue and $10bn in income in recent quarters.

Anyway, cloud computing truly is a modern marvel. Except, of course, when it falls over. And that’s what happened on Monday, when one AWS region – a big one, taking in North Virginia’s data centres – had a DNS issue. Amazon swiftly fixed the issue, but it took nearly a full day for services to recover.

In the meantime, pretty much every company you can think of had outages of varying degrees, from Reddit to Slack and Roblox to McDonalds’ app. Amazon’s own services fell over – affecting Ring cameras and Kindle downloads – and oddly so did Microsoft Teams and Office 365, which is hilarious given Microsoft’s Azure is AWS’ main competitor. So much for eating your own dogfood, Microsoft.

The incident raises plenty of issues: how can DNS still cause so many problems, why did one region falling over cause so many outages when surely these companies have backups in other geographies, and what’s the point of the cloud if it isn’t resilient?

Concentrated risk, shared responsibility

Forrester principal analyst Brent Ellis noted that the incident “highlights how concentration risk – a dangerously powerful yet routinely overlooked systemic risk – arises when so many companies across all industries become dependent on a single cloud provider and, more pertinently, a single region covered by that vendor”.

But he’s not surprised that has happened. First, he pins it on false assumptions that bigger is better. “There’s great appeal to using tech giants, but assuming they are too big to fail or inherently resilient is a mistake, with the evidence being the current outage and past ones,” Ellis says.

It’s easier to just dump it all on AWS or Azure or Google Cloud and let the problems sort themselves out. “Convenience often overshadows navigating the complex, nested dependencies in highly concentrated environments,” Ellis says.

“The entrenchment of cloud, especially AWS, in modern enterprises, coupled with an interwoven ecosystem of SaaS services, outsourced software development, and virtually no visibility into dependencies, is not a bug – it’s a feature of a highly concentrated risk where even small service outages can ripple through the global economy.”

And then Ellis raises the issue of AWS’ shared responsibility model. “As for resilience promises, AWS directs customers to its shared responsibility model as a way to highlight where it takes ownership of service availability and what customers are responsible for. But when core services like DNS fail, even well-architected applications can become unstable,” he adds.

“AWS works to fix its infrastructure, but many enterprises are left to wait until that is done even though they have followed recommended design patterns. This is not exclusively an AWS problem, but it has become a recurring issue, specifically for the US-East region, with customers left holding the bag when it comes to the impact of the outage.”

Heart of the problem: DNS

Let’s start with the tech solution. DNS, or domain name system, is often referred to as the phone book of the internet, and perhaps we need to reboot this digital yellow pages, or at least add in some protections around it.

“This particular outage exposes core issues with cloud resilience that stem from overreliance on services such as DNS, which were not architected for cloud-era technology demands,” noted Forrester’s Ellis.

That said, it’s not a surprise, notes Dr Soohyun Jeon, Assistant Professor/Lecturer in Operation and Information Systems Management, Brunel University of London. “This kind of configuration-level failure, while not unusual in large-scale distributed systems, can cascade rapidly because so many online services rely on shared cloud infrastructure,” Jeon says.

“What begins as a localized fault can therefore manifest globally within minutes, impacting platforms ranging from gaming and social media to banking, telecommunications, and government portals.”

What can be done?

Ellis advises enterprise tech leaders to add these two items to their to-do lists: “Build the tools to increase technology systems’ reliability, and address contractual gray areas related to shared responsibility models with cloud (and SaaS) vendors.”

Jeon added that the incident is a timely reminder to strengthen resilience and continuity planning. “Reliance on a single cloud region or provider introduces concentration risk, which can be mitigated through multi-cloud or multi-region redundancy strategies,” he says. “It also underscores the importance of transparency and shared responsibility in cloud governance.”

Indeed, Nicky Stewart, Senior Advisor at the Open Cloud Coalition, argues that we need a better cloud market that isn’t dominated by AWS or anyone else – something that UK regulators have examined. “Incidents like this make clear the need for a more open, competitive and interoperable cloud market; one where no single provider can bring so much of our digital world to a standstill,” Stewart says.

Dr Corinne Cath-Speth, the Head of Digital at human rights organisation Article 19, agreed, telling The Guardian: “We urgently need diversification in cloud computing. The infrastructure underpinning democratic discourse, independent journalism and secure communications cannot be dependent on a handful of companies.”

And it’s not just about democracy, but our sovereignty, argues Cori Crider, Executive Director of the Future of Technology Institute: “The UK can’t keep leaving its critical infrastructure at the mercy of US tech giants.”

Take back your files?

There’s another way: returning to the old ways. Last year, AWS told the Competition and Markets Authority that it faces competition from “on-premises IT” – in other words, cloud repatriation, when companies take their workloads out of AWS buckets and bring them home.

One example is SaaS firm 37 Signals, the maker of Basecamp, which decided it was tired of paying more than $3 million a year on the cloud and instead invested in its own servers, saving $10 million over five years – and that includes the cost of shelling out for hardware.

Will this outage lead to a rush of repatriation? Perhaps not, as frankly if you can’t keep the lights on with AWS’ help, it may not be worth trying to achieve that on your own, and many smaller companies may well lack the tech skills or millions to invest in servers that made 37 Signals’ shift possible.

Professor Oli Buckley, an expert in Cyber Security at Loughborough University, said the incident was a reminder to give resilience as much consideration as security. “In short: yes, this is a big deal; yes, we should take notice; but also: no, this is not a reason to cancel your cloud migration plans overnight. It is, however, a reason to check how resilient your systems really are,” he says.

Cloud computing may be a marvel, but even it has limitations, and it’s clear the promise of resilience can’t always be kept.

About The Author

Nicole Kobie
Nicole Kobie

Nicole is a journalist and author who specialises in the future of technology and transport. Her first book is called Green Energy, and she's working on her second, a history of technology. At TechFinitive she frequently writes about innovation and how technology can foster better collaboration.

Read more from this author.

We take journalism seriously. To learn more on why you should trust us, head to our editorial guidelines page or meet our team.