AMD has announced that Zyphra has developed ZAYA1, a large-scale mixture-of-experts foundation model trained entirely on AMD Instinct MI300X GPUs, AMD Pensando networking and ROCm open software, powered by IBM Cloud.
AMD has labelled it as a “major milestone” in large-scale AI model training.
“AMD leadership in accelerated computing is empowering innovators like Zyphra to push the boundaries of what’s possible in AI,” AMD artificial intelligence group’s AI and engineering corporate VP Emad Barsoum said.
“This milestone showcases the power and flexibility of AMD Instinct GPUs and Pensando networking for training complex, large-scale models.”
The results, detailed in a Zyphra technical report, showed ZAYA1-base achieved performance comparable to leading models such as Alibaba’s Qwen3-4B and Google’s Gemma3-12B. It outperformed models including Meta’s Llama-3-8B across reasoning, mathematics, and coding benchmarks.
Zyphra also reported that the memory capacity of AMD Instinct MI300X helped the company simplify its training capabilities, while achieving speeds that were more than 10 times faster than other models.
“Efficiency has always been a core guiding principle at Zyphra,” said Zyphra CEO Krithik Puthalath. “It shapes how we design model architectures, develop algorithms for training and inference, and choose the hardware with the best price-performance to deliver frontier intelligence to our customers.”
He added: “ZAYA1 reflects this philosophy and we are thrilled to be the first company to demonstrate large-scale training on an AMD platform. Our results highlight the power of co-designing model architectures with silicon and systems, and we’re excited to deepen our collaboration with AMD and IBM as we build the next generation of advanced multimodal foundation models.”
Zyphra partnership with AMD and IBM
Zyphra said it co-designed ZAYA1 around AMD silicon, resulting in the introduction of advanced routing architecture, compressed convolutional attention, and lightweight residual scaling, so it could achieve higher training throughput.
The results build on prior work between AMD, Zyphra and IBM to design and deploy a large-scale training cluster powered by AMD Instinct GPUs with AMD Pensando networking. The jointly engineered AMD and IBM system, announced earlier this quarter, combines AMD Instinct MI300X GPUs with IBM Cloud’s high-performance fabric and storage architecture. This provides the foundation for ZAYA1’s large-scale pretraining.
“AMD leadership in accelerated computing is empowering innovators like Zyphra to push the boundaries of what’s possible in AI,” said Emad Barsoum, Corporate Vice President of AI and engineering, Artificial Intelligence Group, AMD. “This milestone showcases the power and flexibility of AMD Instinct GPUs and Pensando networking for training complex, large-scale models.”
“As AI creates opportunities for enterprises to innovate, foundation models are key to unlocking accelerated development, efficiency and productivity,” added IBM Cloud GM Alan Peacock.
“We are proud to deliver IBM’s scalable AI infrastructure as the foundation for ZAYA1’s large-scale model and are excited to continue collaborating with AMD on AI model development across our mutual clients.”
You might also be interested