Key ingredients of a modern AI supercomputer

Supercomputers have long been the apex predators of the IT world, but just like nature, technology keeps evolving.


Sponsored by HPE New Logo

Edge-to-cloud, built to transform your business. Learn more about it here.


There are few names, if any, that carry more weight in the history of supercomputing than Cray. The Cray-1, unveiled in 1975, is widely regarded as the seminal supercomputer due to the stunning technology like liquid cooling inside and its sheer commercial success. Cray continued to push the boundaries in supercomputing, most notably with the Cray-2 in 1985, which pioneered immersion liquid cooling for high performance computing.

As the face of computing changed, and supercomputers moved towards massively parallel architectures, the dominance that Cray had enjoyed in the 1970s and 1980s diminished. While Cray evolved to compete in the new, massively parallel world, it was just one of many players.

That all changed in 2019 when Hewlett Packard Enterprise bought Cray for $1.3 billion.  It became the driving force behind HPE’s supercomputing ambitions. Today, the top three supercomputers in the world are HPE Cray Supercomputing EX systems, with six of the top ten spots held by HPE Cray machines. But HPE’s supercomputing leadership extends far beyond the top ten supercomputers, as the share numbers from the latest ranking of the world’s top 100 supercomputers demonstrate.

Today, HPE Cray SC EX systems deliver:

  • 55% of the aggregate computing power of the world’s top 100 most powerful supercomputers
  • 60% of the aggregate accelerator/GPU cores of the world’s top 100 most powerful supercomputers
  • 70% of the aggregate compute power of the world’s top 100 most energy-efficient supercomputers

It’s safe to say that with HPE, Cray is back on top, but today’s HPE Cray SC EX supercomputers are a far cry from 1975’s Cray-1.

The era of exascale

It’s impossible to comprehend the exponential increase in compute power that we’ve seen over the decades, but some milestones are more significant than others. Most recently, that milestone was exascale computing: supercomputers capable of exaflop performance.

An exaflop is one quintillion (1,000,000,000,000,000,000) floating point operations per second. To put that into perspective, the Cray-1 could perform 150 million floating point operations per second. You would need more than 6 billion Cray-1 supercomputers lined up next to one another 437 times around the earth to achieve the same level of power.

In 2022, HPE installed the Frontier supercomputer at the US Department of Energy’s Oak Ridge National Laboratory in Tennessee. Not only did it instantly becoming the fastest supercomputer on earth, it was the first verified exascale computer boasting performance of 1.1 exaflops.

The Frontier supercomputer was “a first-of-its-kind system” that would “benefit humanity” said Justin Hotard, Executive Vice President and General Manager, HPC & AI at HPE. A supercomputer “that was envisioned by technologists, scientists and researchers to unleash a new level of capability to deliver open science, AI and other breakthroughs”.

But while Frontier was the world’s first exascale supercomputer, it didn’t remain the fastest for very long. In 2024 the HPE built the El Capitan supercomputer for the National Nuclear Security Administration (NNSA) and Lawrence Livermore National Laboratory (LLNL). Despite Frontier’s performance increasing over time to 1.353 exaflops, El Capitan’s 1.742 exaflops placed it at number one by some margin. In terms of Cray-1 systems, that’s 11.6 billion.

Exascale computing isn’t just about numbers and rankings, though. These incredible levels of compute power can push the boundaries of AI in myriad fields, from drug discovery to climate change.

Of course, passing that exascale milestone wasn’t easy, and while the three fastest supercomputers in the world can all deliver performance in excess of one exaflop, it speaks volumes that all three were built by HPE.

El Capitan supercomputer
El Capitan comprises more than 11,000 compute nodes (image: Garry McLeod/LLNL)

Anatomy of a supercomputer

HPE’s success in the field of supercomputing is based on a similar principle to Cray’s success decades before – the understanding that every single part of a supercomputer needs to deliver the highest possible performance. In a world where thousands upon thousands of compute nodes are being utilised, connectivity and data storage speeds are just as important as the processors crunching the numbers.

That’s not to say that processors aren’t important – quite the contrary when AI workloads have joined classic modelling & simulation and high performance data analysis in the realm of high-performance computing. Powerful CPU cores will always form the foundation of any supercomputing platform, as they always have. But CPU cores aren’t enough: these are now joined by equally powerful GPU-based accelerators. These provide the astonishing levels of data crunching power that data scientists and AI applications demand.

Both sides of that compute equation require lightning fast, low latency memory, with enough bandwidth to ensure that no compute cycles are wasted.

For El Capitan, HPE used AMD Instinct MI300A Accelerated Processing Units (APUs), which combine CPU cores, GPU accelerators and high-bandwidth, shared memory, into one package. This tight integration of compute and memory ensures that data can be fetched and processed as fast as possible, without the latency or bandwidth hit that would accompany an “off-die” memory solution.

Download “How To Start Your AI Journey Without Getting Locked In” from HPE

The timeline for launching a successful AI pilot program will vary depending on factors such as the project’s complexity, available resources, and organizational readiness. You don’t need a complex toolset and a multitude of data scientists to get your first AI use case off the ground. It’s possible to launch an AI pilot relatively quickly – in as little as two to three months – by following a strategic plan. Here’s how.

Evolution of supercomputer storage

Any supercomputer will be processing huge amounts of data that needs to be stored somewhere with equally huge capacity. Back in Cray’s original heyday that meant hard disk drives the size of washing machines, providing a tiny fraction of the capacity seen in modern smartphones, but as with the rest of supercomputing, storage has evolved.

HPE Cray Supercomputing Storage Systems E2000 provide the performance, cost-effectiveness and flexibility required for any supercomputing application. A single storage rack can transfer the data equivalent of 500 full-length high-definition movies to the compute nodes of a supercomputer in a single second. And a single storage rack can store data equivalent to more than 1,100 printed collections of the U.S. Library of Congress.

Those record-breaking capabilities are delivered in the most cost-effective way as HPE Cray SC Storage Systems E2000 embed an open-source file system with enterprise-grade customer support, exploit the strength of different storage media (SSD and HDD) without incurring their weaknesses and connect directly into the interconnect network that interconnects the compute nodes of the supercomputer.

When it comes to supercomputing, it’s all about scale – millions of compute cores pulling petabytes of data from myriad storage drives – and ensuring that application performance scales alongside hardware infrastructure requires fast and robust interconnect technology.

HPE Slingshot interconnect 400 delivers the kind of high ban

dwidth, low latency environment that modern supercomputing applications demand. It scales application performance better and more efficiently than any other high bandwidth, low latency interconnect. HPE Slingshot interconnect technology was vital to HPE’s exascale computing ambitions – this becomes abundantly clear when you consider that El Capitan employs 44,544 AMD Instinct MI300A APUs, containing a staggering 1,051,392 CPU cores and 9,988,224 GPU cores all connected to 400 petabytes of shared data in a single file system.

It’s all cool

HPE ProLiant Compute XD server
The HPE ProLiant Compute XD server family optimizes for AI model training and tuning (image: HPE)

There’s one unavoidable byproduct with any form of high-performance computing: heat. Data centers are extremely power hungry, with cooling demanding significant amounts of power along with the processors.

Just as Cray did back in the 1980s, HPE is pioneering new 100% fanless liquid cooling technology. HPE’s 100% fanless direct liquid cooling architecture delivers two key advantages. First, by avoiding the airflow needed for traditional cooling it allows for higher density systems – essentially more processing power can be packed into each rack leading to performance and cost-effectiveness benefits. Second, direct liquid cooling significantly reduces the power consumption and associated costs compared to air cooling.

“This direct liquid cooling architecture yields 90% reduction in cooling power consumption as compared to traditional air-cooled systems,” said Antonio Neri, President & CEO, HPE. He adds, “as organizations embrace the possibilities created by generative AI, they also must advance sustainability goals, combat escalating power requirements, and lower operational costs.”

Supercomputers have undoubtedly changed over the past 50 years, but the underlying ethos remains the same – every component must operate at optimum efficiency to deliver maximum performance on a system level. Where we go from this era of exascale remains to be seen, but as the needs and application of AI grows, so will the performance of the hardware required to deliver it.

Avatar photo
Riyad Emeran

Riyad is a highly experienced writer and editor who has spent over 30 years writing about the technology industry. He co-founded and edited the website Trusted Reviews in 2003, having previously been Editor-in-Chief of Personal Computer World magazine, affectionately known as PCW.