Meta-backed architecture could make an entire data center operate like one computer

Meta-backed architecture could make an entire data center operate like one computer

AI datacenters could soon be designed to behave less like collections of separate machines and more like a single giant computer. That is the idea behind a new architecture proposed by semiconductor company Panmnesia and Meta. The design uses Compute Express Link (CXL) to connect CPUs, AI accelerators, and memory across racks while keeping their operation tightly coordinated. The problem becomes more serious as AI models grow. Training models with trillions of parameters can require hundreds or thousands of accelerators to exchange terabytes of data. Even if most of those devices finish their tasks quickly, one slow component can hold up the entire operation. That makes latency predictability increasingly important. Inside a rack, accelerators can already communicate through dedicated high-speed connections. Across racks, however, data typically travels through conventional networking equipment and software layers, introducing more variation in how long individual requests take. One fabric replaces rack boundaries Panmnesia and Meta’s approach extends a CXL domain beyond individual racks and into the wider datacenter. Instead of treating each rack as a separate computing island, the architecture aims to make resources across the facility operate as a coordinated system. Three pieces of hardware are central to the design: a high-fan-out non-blocking CXL switch, a link acceleration unit, and a fabric controller. Together, they are intended to reduce unpredictable delays as data moves between devices. The architecture also uses optical connections to overcome the physical distance limitations of electrical signaling. CXL-over-optics could therefore extend the fabric across larger portions of a datacenter without abandoning the underlying CXL model. The proposed system changes the scale at which computing resources can work together. In the reference configuration, a CPU is connected to two accelerators. Under the new architecture, that number rises to 16, while a single coherence domain could encompass as many as 960 accelerators. Latency becomes the new battleground The researchers say accesses that would normally leave a rack and travel through a conventional network could instead follow a more predictable path. Round-trip latency could fall from the microsecond range to several hundred nanoseconds, representing a reduction of up to an order of magnitude. That predictability could matter as much as raw computing power. Large AI workloads often operate collectively, meaning the slowest participant can determine when the entire operation progresses. Reducing those delays could allow more accelerators to function as one execution environment. The architecture could also make failures less disruptive. Rather than replacing or isolating an entire server when something goes wrong, the proposed system narrows the replacement unit to an individual device. “As AI systems continue to scale, the ability to connect large numbers of accelerators and memory devices quickly and efficiently is becoming just as important as the performance of individual accelerators. This research outlines a direction for next-generation AI infrastructure, where CXL enables the entire datacenter to operate as a single computing system,” Myoungsoo Jung, CEO of Panmnesia, said. Panmnesia says it has already implemented the core components in silicon, completed validation, and is preparing them for commercial supply. The work was published in Nature Reviews Electrical Engineering. Get the latest in engineering, tech, space & science - delivered daily to your inbox.With over a decade-long career in journalism, Neetika Walter has worked with The Economic Times, ANI, and Hindustan Times, covering politics, business, technology, and the clean energy sector. Passionate about contemporary culture, books, poetry, and storytelling, she brings depth and insight to her writing. When she isn’t chasing stories, she’s likely lost in a book or enjoying the company of her dogs.

Original Source

Read the full article at Interestingengineering →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.