AWS falls over, makes the world reconsider cloud computing – so what lessons should we learn from the AWS outage?

Forget keeping all your files and documents on your own server โ€“ thatโ€™s not safe, lacks resilience. What you should do is chuck it into our network of web-connected buckets, which are totally resilient because theyโ€™re located all over the world. Plus, itโ€™ll be cheaper and easier.

Thatโ€™s the argument that convinced the corporate world to sign up to cloud computing, an idea with older origins that really started with the launch of Amazon Web Services in 2002 followed by the arrival of Elastic Compute Cloud four years later. Google and Microsoft both jumped on the bandwagon a few years after.

A decade and a half later, the vast majority of companies are using some sort of cloud, and itโ€™s making AWS โ€“ which holds 30% market share โ€“ about $30bn in revenue and $10bn in income in recent quarters.

Anyway, cloud computing truly is a modern marvel. Except, of course, when it falls over. And thatโ€™s what happened on Monday, when one AWS region โ€“ a big one, taking in North Virginiaโ€™s data centres โ€“ had a DNS issue. Amazon swiftly fixed the issue, but it took nearly a full day for services to recover.

In the meantime, pretty much every company you can think of had outages of varying degrees, from Reddit to Slack and Roblox to McDonaldsโ€™ app. Amazonโ€™s own services fell over โ€“ affecting Ring cameras and Kindle downloads โ€“ and oddly so did Microsoft Teams and Office 365, which is hilarious given Microsoftโ€™s Azure is AWSโ€™ main competitor. So much for eating your own dogfood, Microsoft.

The incident raises plenty of issues: how can DNS still cause so many problems, why did one region falling over cause so many outages when surely these companies have backups in other geographies, and whatโ€™s the point of the cloud if it isnโ€™t resilient?

Concentrated risk, shared responsibility

Forrester principal analyst Brent Ellis noted that the incident โ€œhighlights how concentration risk โ€“ a dangerously powerful yet routinely overlooked systemic risk โ€“ arises when so many companies across all industries become dependent on a single cloud provider and, more pertinently, a single region covered by that vendorโ€.

But heโ€™s not surprised that has happened. First, he pins it on false assumptions that bigger is better. โ€œThereโ€™s great appeal to using tech giants, but assuming they are too big to fail or inherently resilient is a mistake, with the evidence being the current outage and past ones,โ€ Ellis says.

Itโ€™s easier to just dump it all on AWS or Azure or Google Cloud and let the problems sort themselves out. โ€œConvenience often overshadows navigating the complex, nested dependencies in highly concentrated environments,โ€ Ellis says.

โ€œThe entrenchment of cloud, especially AWS, in modern enterprises, coupled with an interwoven ecosystem of SaaS services, outsourced software development, and virtually no visibility into dependencies, is not a bug โ€“ itโ€™s a feature of a highly concentrated risk where even small service outages can ripple through the global economy.โ€

And then Ellis raises the issue of AWSโ€™ shared responsibility model. โ€œAs for resilience promises, AWS directs customers to its shared responsibility model as a way to highlight where it takes ownership of service availability and what customers are responsible for. But when core services like DNS fail, even well-architected applications can become unstable,โ€ he adds.

โ€œAWS works to fix its infrastructure, but many enterprises are left to wait until that is done even though they have followed recommended design patterns. This is not exclusively an AWS problem, but it has become a recurring issue, specifically for the US-East region, with customers left holding the bag when it comes to the impact of the outage.โ€

Heart of the problem: DNS

Letโ€™s start with the tech solution. DNS, or domain name system, is often referred to as the phone book of the internet, and perhaps we need to reboot this digital yellow pages, or at least add in some protections around it.

โ€œThis particular outage exposes core issues with cloud resilience that stem from overreliance on services such as DNS, which were not architected for cloud-era technology demands,โ€ noted Forresterโ€™s Ellis.

That said, itโ€™s not a surprise, notes Dr Soohyun Jeon, Assistant Professor/Lecturer in Operation and Information Systems Management, Brunel University of London. โ€œThis kind of configuration-level failure, while not unusual in large-scale distributed systems, can cascade rapidly because so many online services rely on shared cloud infrastructure,โ€ Jeon says.

โ€œWhat begins as a localized fault can therefore manifest globally within minutes, impacting platforms ranging from gaming and social media to banking, telecommunications, and government portals.โ€

What can be done?

Ellis advises enterprise tech leaders to add these two items to their to-do lists: โ€œBuild the tools to increase technology systemsโ€™ reliability, and address contractual gray areas related to shared responsibility models with cloud (and SaaS) vendors.โ€

Jeon added that the incident is a timely reminder to strengthen resilience and continuity planning. โ€œReliance on a single cloud region or provider introduces concentration risk, which can be mitigated through multi-cloud or multi-region redundancy strategies,โ€ he says. โ€œIt also underscores the importance of transparency and shared responsibility in cloud governance.โ€

Indeed, Nicky Stewart, Senior Advisor at the Open Cloud Coalition, argues that we need a better cloud market that isnโ€™t dominated by AWS or anyone else โ€“ something that UK regulators have examined. โ€œIncidents like this make clear the need for a more open, competitive and interoperable cloud market; one where no single provider can bring so much of our digital world to a standstill,โ€ Stewart says.

Dr Corinne Cath-Speth, the Head of Digital at human rights organisation Article 19, agreed, telling The Guardian: โ€œWe urgently need diversification in cloud computing. The infrastructure underpinning democratic discourse, independent journalism and secure communications cannot be dependent on a handful of companies.โ€

And itโ€™s not just about democracy, but our sovereignty, argues Cori Crider, Executive Director of the Future of Technology Institute: โ€œThe UK canโ€™t keep leaving its critical infrastructure at the mercy of US tech giants.”

Take back your files?

Thereโ€™s another way: returning to the old ways. Last year, AWS told the Competition and Markets Authority that it faces competition from โ€œon-premises ITโ€ โ€“ in other words, cloud repatriation, when companies take their workloads out of AWS buckets and bring them home.

One example is SaaS firm 37 Signals, the maker of Basecamp, which decided it was tired of paying more than $3 million a year on the cloud and instead invested in its own servers, saving $10 million over five years โ€“ and that includes the cost of shelling out for hardware.

Will this outage lead to a rush of repatriation? Perhaps not, as frankly if you canโ€™t keep the lights on with AWSโ€™ help, it may not be worth trying to achieve that on your own, and many smaller companies may well lack the tech skills or millions to invest in servers that made 37 Signalsโ€™ shift possible.

Professor Oli Buckley, an expert in Cyber Security at Loughborough University, said the incident was a reminder to give resilience as much consideration as security. โ€œIn short: yes, this is a big deal; yes, we should take notice; but also: no, this is not a reason to cancel your cloud migration plans overnight. It is, however, a reason to check how resilient your systems really are,โ€ he says.

Cloud computing may be a marvel, but even it has limitations, and itโ€™s clear the promise of resilience canโ€™t always be kept.

Nicole Kobie
Nicole Kobie

Nicole is a journalist and author who specialises in the future of technology and transport. Her first book is called Green Energy, and she's working on her second, a history of technology. At TechFinitive she frequently writes about innovation and how technology can foster better collaboration.