Technology & Science

AMD Unveils Helios: A Fully Integrated Rack-Scale AI System Designed to Redefine High-Performance AI Computing

Advanced Micro Devices (AMD) has officially launched Helios, its groundbreaking, fully integrated rack-scale AI system, marking a pivotal moment in the company’s ambitious push into the high-performance artificial intelligence market. Unveiled at the "Advancing AI 2026" event, Helios represents AMD’s first complete, in-house designed AI system, meticulously integrating 72 of its Instinct MI455X GPUs, Epyc Venice processors, Pensando networking technology, and the comprehensive ROCm software stack into a singular, cohesive unit. This launch signals AMD’s bold ambition to construct the most powerful AI rack globally, underpinned by a steadfast commitment to open standards and an integrated architecture. The system is now officially in production, poised to address the escalating demands of hyperscale AI workloads.

The Genesis of Helios: A Strategic Vision for Integrated AI

The development of Helios began over a year ago, driven by a clear and challenging mandate. Andrew Dieckmann, Corporate Vice President and General Manager of Data Center GPU at AMD, encapsulated the project’s core mission: "Our North Star with Helios was to build a system that utilizes the full capacity of our compute, in an open rack architecture with open standards, and that raises the bar in terms of efficiency, maintainability, and reliability. That was the team’s mission." This statement underscores AMD’s strategic pivot towards offering not just components, but a holistic, optimized AI solution designed to overcome the limitations of traditional, disaggregated systems.

For years, the AI hardware landscape has been dominated by a component-centric approach, where data centers assemble AI clusters from individual GPU servers interconnected via complex, multi-hop networks. This traditional model often leads to increased latency, network congestion, and inefficiencies that impede the performance of large-scale AI models. Helios directly challenges this paradigm. As Mark Chubb, Corporate Vice President of Platform Architecture, who joined AMD from rack-builder ZT Systems, succinctly puts it, "The rack becomes the new system boundary." This shift signifies a fundamental rethinking of how AI infrastructure is designed, moving towards a unified, system-level approach where the entire rack functions as a single, coherent computing entity. This integration is crucial for maximizing the performance and efficiency of massive AI models, which require seamless, high-bandwidth communication between hundreds or thousands of accelerators.

AMD Helios doorgelicht: het rack is de nieuwe server - ITdaily

Architectural Marvel: Unpacking Helios’s Design Prowess

The Helios system is engineered for unprecedented scale and performance. At its heart, all 72 Instinct MI455X GPUs share a colossal 31 TB of High Bandwidth Memory (HBM4), accessible via a single switch-hop with an impressive 260 TB/s of scale-up bandwidth. This architectural choice ensures that every GPU can communicate with any other GPU within the rack at identical speeds, regardless of its physical location. This uniform access simplifies software development, as developers no longer need to account for network topology, allowing them to focus entirely on optimizing AI workloads. The sheer computational power is staggering: each Helios rack delivers 2.9 exaflops of FP4 compute and 1.4 exaflops of FP8 compute, alongside 1.7 petabytes per second of memory bandwidth and 43 TB/s of scale-out bandwidth to the broader data center network. Such figures position Helios as a formidable contender in the race for extreme-scale AI processing.

Physically, Helios conforms to the Open Rack Wide (ORW) standard, a specification developed in collaboration with Meta through the Open Compute Project (OCP). This double-wide rack measures 1.2 by 1.3 meters and stands 44 rack units (OU) tall. Within this robust chassis, 18 compute trays and six switch trays slide into place, interconnected at the rear by four interchangeable cassettes utilizing full copper cabling. This modular design emphasizes both performance and serviceability.

Each compute tray, occupying a mere 10U, houses four liquid-cooled Instinct MI455X GPUs and a single Epyc Venice SP7 processor. These components are tightly coupled via AMD’s coherent Infinity Fabric connections, providing 256 GB/s of direct bandwidth per GPU and allowing the CPU to actively participate in the shared memory domain. The density and power of these trays are considerable; each weighs approximately 77 kilograms and features 576 differential connections, requiring 120 kilograms of force to fully insert into the rack.

The switch trays are even more impressive in their engineering. Each tray contains two Broadcom Tomahawk 6 switch ASICs, designed to handle immense data traffic. With 1,728 connections and demanding 310 kilograms of insertion force, these trays are equipped with noticeably long levers to provide the necessary leverage. A total of twelve switch ASICs within the rack create a single-hop, multi-plane fabric based on UALink over Ethernet (UALoE), ensuring unparalleled communication efficiency across all GPUs. A critical design philosophy for Helios is in-rack maintainability: trays can be fully extended to replace DIMMs or other components without the need for server lifts or crash carts, significantly reducing downtime and operational complexity.

AMD Helios doorgelicht: het rack is de nieuwe server - ITdaily

Networking at the Core: Pensando’s Pivotal Role

In an AI system of this magnitude, the network is not an auxiliary component but a foundational element. This is where AMD’s Pensando division, acquired in 2022, plays a crucial role. Krishna Doddapaneni, Corporate Vice President at Pensando, emphasizes this integration: "AI scales at the speed of data transport." Helios employs a sophisticated, multi-tiered networking architecture designed to optimize data flow for diverse AI workloads.

The front-end network is managed by the Pensando Salina 400-DPU (Data Processing Unit). This programmable card offloads critical network virtualization, firewalling, encryption, and NVMe storage services from the main CPUs. Crucially, it also handles key-value (KV) cache traffic, which is becoming increasingly vital for the efficient operation of large language models and other AI applications.

The scale-up network is responsible for interconnecting the 72 GPUs, enabling them to function as a single, logical memory domain via UALoE. This ensures high-speed, low-latency communication essential for collective operations in distributed AI training.

For scale-out capabilities, connecting to thousands of GPUs across the broader data center, each GPU within Helios is equipped with three Pensando Vulcano 800G AI-NICs (Network Interface Cards). This configuration delivers an astounding 2.4 terabits per second of bandwidth per GPU. The programmable Vulcano NICs support standard RoCE (RDMA over Converged Ethernet) and also incorporate new transport protocols like MRC (Multi-Rail Communication), a technology developed by AMD in collaboration with OpenAI to further optimize large-scale AI communication. This advanced networking suite ensures that Helios can seamlessly integrate into and scale within existing and future hyperscale data center environments.

AMD Helios doorgelicht: het rack is de nieuwe server - ITdaily

Resilience and Manageability: Engineered for Enterprise AI

Given the immense scale and complexity of Helios, AMD designed the system with an explicit focus on resilience and fault tolerance. Recognizing that component failures are not just a possibility but an inevitability in systems of this size, Helios is "designed to fail gracefully." Should a switch crash or require replacement, traffic is automatically rerouted, allowing ongoing training jobs to continue with a temporary, minor reduction in bandwidth, rather than catastrophic failure. This "always-on" capability is critical for enterprise-grade AI operations, where downtime can translate into significant financial losses and project delays.

Management and control are also built with redundancy in mind. The AMD Fabric Manager, responsible for orchestrating the entire rack’s operations, runs in a triplicated configuration, leveraging the compute resources of the switch trays themselves. This ensures that even the failure of a single management instance will not disrupt the fabric’s integrity. Furthermore, Helios supports virtual pods, enabling operators to logically partition the rack into isolated sub-clusters for different tenants or workloads. If one virtual pod experiences an issue, it remains contained, preventing cascading failures across the entire system. The network control plane operates on SONiC (Software for Open Networking in the Cloud), an open-source network operating system, and AMD is committed to upstreaming any of its contributions back to the open-source community, reinforcing its dedication to open standards.

A Growing Customer Base: Hyperscalers Embrace Helios

The true testament to Helios’s potential lies in its early adoption by some of the world’s largest AI players. The string of multi-gigawatt deals announced even before the official launch underscores the industry’s confidence in AMD’s integrated AI solution. Last autumn, OpenAI inked a multi-year agreement, committing to a substantial 6 gigawatts of AMD infrastructure. This was followed in February by Meta, which signed a similarly large deal focusing on custom AMD GPUs. Most recently, Microsoft announced large-scale Helios deployments on its Azure cloud platform.

AMD Helios doorgelicht: het rack is de nieuwe server - ITdaily

At the "Advancing AI 2026" event itself, AMD further solidified its market position by announcing a multi-gigawatt agreement with Anthropic, a leading AI safety and research company, which also includes a long-term technical collaboration. Beyond these hyperscale giants, AMD anticipates rolling out "multiple gigawatts" of Helios infrastructure over the coming year, with Oracle and a range of emerging "neoclouds" also standing in line. These significant commitments from industry leaders highlight a strategic imperative: while performance is paramount, diversification of AI hardware suppliers is becoming increasingly critical to mitigate supply chain risks and foster competition in a market historically dominated by a single vendor.

Challenging the Incumbent: AMD vs. Nvidia in the AI Arena

AMD is positioning Helios in direct competition with Nvidia’s forthcoming Vera Rubin NVL72 system, setting the stage for an intense rivalry in the high-end AI infrastructure market. AMD boldly claims that Helios offers a compelling performance advantage: 15 percent more FP4 compute, 50 percent more HBM memory, 50 percent more scale-out bandwidth, and up to 30 percent more "tokens per euro" compared to Nvidia’s announced specifications. The "tokens per euro" metric is particularly significant for hyperscalers, as it directly translates to the cost-efficiency of running large-scale AI models.

However, a degree of nuance is warranted in these comparisons. Both Helios and Nvidia’s Vera Rubin are slated for release in the second half of 2024. AMD’s benchmarks are based on its internal measurements and models, comparing them against Nvidia’s publicly disclosed specifications for a system that is also yet to be deployed at scale in customer environments. Nvidia, for its part, has made equally bold claims for Vera Rubin, promising five times the inference performance of its current Blackwell platform and ten times lower token costs.

Moreover, the software ecosystem remains a crucial battleground. Despite AMD’s significant strides with its ROCm software stack, the vast majority of AI tooling and developers still operate within the CUDA ecosystem, which has benefited from years of NVIDIA’s investment and market dominance. While Andrew Dieckmann dismisses CUDA as a "non-event" today, asserting that "everyone programs at higher abstraction levels, and AI agents help customers to optimize for our platform," the inertia of a deeply entrenched ecosystem is not to be underestimated. Nvidia also benefits from a massive installed base of Blackwell customers, providing a natural upgrade path to Vera Rubin. In this highly competitive landscape, the burden of proof undeniably rests with the challenger to demonstrate real-world performance and seamless integration.

AMD Helios doorgelicht: het rack is de nieuwe server - ITdaily

The Open Ecosystem Advantage: ROCm and Beyond

AMD’s commitment to open standards, exemplified by the Open Rack Wide design and the open-source SONiC integration, extends deeply into its software strategy with ROCm. This open-source platform for GPU computing is AMD’s direct answer to CUDA, providing developers with the tools and libraries needed to accelerate AI, high-performance computing (HPC), and graphics workloads on AMD hardware. The strategy to embrace open standards and foster an open ecosystem is a deliberate move to reduce vendor lock-in, encourage broader adoption, and accelerate innovation. By collaborating with partners like OpenAI on new protocols such as MRC and contributing to open-source projects, AMD aims to make its platform more accessible and appealing to a wider developer community. The success of this strategy is critical for long-term market penetration, as the hardware’s capabilities must be fully unleashed by a robust and accessible software environment.

Supply Chain and Market Implications

In an era of complex global supply chains and geopolitical uncertainties, the availability of components is as crucial as performance. Andrew Dieckmann addressed this directly, stating that AMD is taking direct control of the supply chain for Helios: "We don’t sell Helios as an AMD system, but we work with the entire ecosystem to ensure that all components are available in the right quantities, at the right time, and with the right quality. We have been working on this for many, many months." This hands-on approach reflects the strategic importance of Helios and the need to guarantee consistent delivery to hyperscale customers who operate on immense timelines and scales.

The launch of Helios has profound implications for the AI hardware market. It introduces a powerful, integrated alternative that could significantly challenge Nvidia’s long-standing dominance. Increased competition is likely to spur further innovation, potentially leading to faster performance improvements and more cost-effective solutions for AI development and deployment. For hyperscalers, the availability of a viable alternative like Helios offers not only competitive pricing leverage but also the critical benefit of supply chain diversification, reducing their reliance on a single vendor.

AMD Helios doorgelicht: het rack is de nieuwe server - ITdaily

Conclusion: The Road Ahead for AMD’s AI Ambition

With the Instinct MI455X GPUs as its engine, Epyc Venice processors as its conductor, and an open ecosystem as its compelling sales argument, AMD’s Helios system presents a complete, integrated rack that, on paper, stands toe-to-toe with the industry’s best. The multi-gigawatt endorsements from AI titans like OpenAI, Meta, Anthropic, and Microsoft unequivocally signal that major buyers are not only believing AMD’s narrative but are also actively seeking robust alternatives to diversify their critical AI infrastructure.

The coming months will be crucial. As Helios moves from factory production lines into real-world data centers during the second half of the year, its ultimate success will hinge on its ability to translate impressive paper specifications into consistently high "running tokens" for complex AI workloads. Only then will the industry definitively know if AMD’s rack is a truly formidable alternative that can reshape the AI landscape, or if it will primarily serve as an excellent negotiation tool for customers seeking better terms from established players. Regardless of the final outcome, Helios undeniably represents a significant leap forward for AMD and a powerful catalyst for competition and innovation in the rapidly evolving world of artificial intelligence.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button