Robotics and edge AI put new pressure on computing infrastructure Physical AI is forcing the technology industry to rethink the entire computing stack. Robots, autonomous systems and intelligent devices need economical inference, secure data access and infrastructure that works beyond conventional clouds. Rafay Systems is addressing those demands through orchestration that lets providers offer open models without dedicating entire GPU systems to individual customers, potentially lowering enterprise AI costs while protecting sensitive information, according to Haseeb Budhani (pictured, right), co-founder and chief executive officer of Rafay Systems Inc. “The serverless part at the infrastructure level, we’ve solved that … we’ve done for a while,” Budhani said. “Now we have a slice of essentially a confidential GPU. I can now have better security on top of the GPUs that I already have, which means as a provider, I get to monetize every second of my infrastructure. But, for the enterprise, it’s actually a better deal. They like that because they look at the total cost of ownership, they look at the capex and they don’t have the big money. It’s too expensive, otherwise. This allows them to have a secure solution, which actually delivers a better price point. That’s what everybody wants.” Budhani, Eiman Ebrahimi (left), CEO of Protopia AI Inc., and other industry leaders spoke with theCUBE’s John Furrier and guest host Howie Xu, chief AI & innovation officer of Gen Digital Inc., as part of theCUBE + NYSE Wired: Robotics & AI Infra Leaders event, during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. They examined the infrastructure, economics and technology required to move AI into the physical world. The event builds on theCUBE’s Machina AI Summit coverage of robotics moving from demonstrations into production. (* Disclosure below.) Physical AI changes the infrastructure equation The economics behind physical AI begin with tokens. Agentic applications consume far more tokens than simple chat interactions because every new action may require previous context to be processed again. That makes token volume, model selection and workload design increasingly important measures of infrastructure demand, explained Max Kan, tokenomics technical lead at SemiAnalysis LLC. “In general, agentic workloads consume exponentially more tokens than the chat workloads everyone was using a year ago because it’s all inherently multi-turn,” he said. “And every time you ask your agent a follow-up or it uses some tool like web search or maybe it searches your code base, it has to reprocess all previous tokens in your conversation, which really causes total tokens produced/processed to go up exponentially.” The inference market is splitting into specialized workloads, with prefill demanding compute, while decode depends on memory capacity and bandwidth. Positron AI Inc. is targeting those requirements through systems optimized for tokens per dollar and tokens per watt, including air-cooled designs suited to enterprise data centers that cannot support liquid-cooled racks, according to Darren Chien, managing director of APAC at Positron AI. “We believe that inference is fundamentally an economics problem,” he said. “We need to make sure that these tokens are generated as cheaply and as widely available to users as possible. The way we compete is on optimizing the two metrics of tokens per dollar and tokens per watt.” As physical AI moves into factories, vehicles and distributed environments, infrastructure providers need more than faster accelerators. Axiado Corp. addresses the resulting demands for power, security and system control through silicon-based platform management that adjusts cooling, frequency and voltage by workload while reducing operational pressure on CPUs and GPUs at scale, noted Gopi Sirineni, founder, president and CEO of Axiado. “We’re bringing that intelligence into a silicon level, and it will be offloaded off of your GPUs, CPUs,” he said. “We’re monitoring the system itself, including the fans, liquid cooling and all that. Number one, we manage the efficiency by managing the fans and fan controls and liquid cooling … and then, two, we do dynamic frequency and voltage scaling for the platform.” Intelligence moves closer to the physical world Physical AI increasingly depends on compact models that can run locally despite limits on power, memory and connectivity. Liquid AI Inc. is developing customizable foundation models for laptops, vehicles and other devices, allowing specialized intelligence to reduce inference costs, improve privacy and deliver faster responses without relying entirely on centralized data centers, pointed out Ramin Hasani, co-founder and CEO of Liquid AI Inc. “One small model is not going to be generally intelligent,” he told theCUBE. “What they can do, they can be fine-tuned. You can customize them. The nice thing about this thing is that customizing a small model, the cost of it is going to be shockingly low.” As intelligence moves onto devices, networking becomes part of the physical AI application architecture. Aria Networks Inc. applies specialized AI to high-resolution telemetry, helping operators detect failures and performance issues before they disrupt costly workloads. The goal is to strengthen human decision-making with timely context, not remove people from network operations, according to Mansour Karam, founder and CEO of Aria Networks. “It starts with collecting telemetry. There are two aspects to collecting telemetry,” he said. “You need to be collecting telemetry at the microsecond resolution. That’s number one. You need it at the right resolution. Number two, the network, and we can talk about all the different networks in an AI factory, but it spans very different domains.” Enterprise AI is moving beyond chatbots toward computer-use agents that can operate existing software, capture workflows and execute routine tasks without lengthy integrations. H Company is targeting that shift by placing a human-centered intelligence layer over legacy systems, helping organizations automate repetitive work while keeping consequential actions under human control, emphasized Gautier Cloix, CEO of H Company. “What’s hard is to do it with the safety part, without the risk,” he said. “That’s the same thing for enterprise. The way we built our agents is that they cannot execute irreversible tasks – sending an email, deleting something, ordering something – without a human validation at first … that’s where we focus a lot of energy, making sure that deterministically our agents cannot do the wrong things.” New markets emerge around compute capacity AI’s rapid growth is creating demand for financial tools that can manage volatile GPU pricing and availability. Silicon Data and The Compute Exchange Inc. are developing compute benchmarks and instruments that could help infrastructure operators hedge falling prices while protecting major AI consumers against rising costs as demand shifts over time, explained Carmen Li, founder and CEO of Silicon Data and CEO of The Compute Exchange. “I’m helping probably a dozen market participants set up their compute desks. People are actually helping their clients to hedge,” she said. “Think about if your client, if you’re a bank, your client can be the hyperscaler or neoclouds, they have a long exposure, they try to help them manage the futures and volatility for their revenues. Or … your client can be the massive AI company, they’re going to consume a lot of GPU tokens. They help them manage the short exposure.” Compute markets are becoming more distributed as organizations combine owned infrastructure with capacity from neoclouds and regional providers. San Francisco Compute Co. is building a marketplace for GPU capacity that supports this shift, using secure multi-tenancy to match infrastructure with workload, location and availability needs without exposing customer data, emphasized Alan Butler, chief business officer of San Francisco Compute Co. “That was the genesis of having a compute layer where you can rent out on short-term provisioning that became then the basis of, in essence, a marketplace,” he said. “There’s a lot of marketplace companies that are flipping GPUs over a fence with no SLAs behind it. And everything we do has an SLA standing behind it with a data center provider.” Edge deployments require infrastructure that runs models near where data is created without forcing enterprises to rebuild data centers. Axelera AI B.V. is extending its edge-efficient architecture into servers and cloud systems, using familiar software frameworks so developers can adopt new hardware while avoiding proprietary application redesigns and reducing deployment barriers. “They’re designed together very intentionally,” said Alexis Crowell, chief marketing officer and general manager of the Americas at Axelera AI. “We have as many folks working on the software side of our business as we do on the hardware side. It’s not that I’m trying to monetize the software. I’m trying to … make sure that developers have a frictionless experience.” To watch more of theCUBE’s coverage of theCUBE + NYSE Wired: Robotics & AI Infra Leaders event, here’s our complete video playlist: (* Disclosure: Neither ScaleFlux Inc., the presenting sponsor of theCUBE + NYSE Wired: Robotics & AI Infra Leaders event, nor other sponsors have editorial control over content on theCUBE or SiliconANGLE.) Photo: SiliconANGLE A message from John Furrier, co-founder of SiliconANGLE: Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities. 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network. Are you AWS customer? Support SiliconANGLE Financially by buying your AWS services from our Marketplace portal page and links. About SiliconANGLE Media SiliconANGLE Media is a recognized leader in digital media innovation, uniting breakthrough technology, strategic insights and real-time audience engagement. As the parent company of SiliconANGLE, theCUBE Network, theCUBE Research, CUBE365, theCUBE AI and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI. Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.
Robotics and edge AI put new pressure on computing infrastructure
Full Article
Original Source
Read the full article at Siliconangle →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.