All in a day’s reading, if you’re following innovation

All in a day’s reading, if you’re following innovation

One problem with following innovation is that the extraordinary now arrives disguised as the weekly. A frontier model trained on more than 100,000 GPUs, several leading AI services stumbling in the same morning, a camera-chip company redesigning itself around edge inference, and four Western drugmakers licensing Chinese science can all appear between two sets of results. This is less about finding one big story than noticing the density. Product plans, cost structures, factory bottlenecks, pricing models and corporate boundaries are changing together. Many developments will take years to reach revenue, and several will fail. Each still tells us where companies are spending capital, accepting risk and trying to enter the next economy. What follows are ten things gathered from roughly a day’s reading of the week’s news. They run from the hardware bill inside software and the return of task-priced “solutions” to AI reliability, to the increasingly global skilled construction labor shortage in innovation epicenters, to China’s drug-discovery pipeline and Nvidia’s expanding sphere of influence. We have mostly left the far more important announcements made at SEMICON Taiwan for separate work, and some of these thoughts will become full pieces later. None is a maintenance item. Each changes at least one part of the map, which is quite a lot for an ordinary week. For readers who get tired of going through the thicket, we request that you jump to the last one in the nomenclature. It might not be the most important, but certainly it’s the most interesting. One hundred thousand GPUs and a benchmark that blinked Every model release for three years has staged the same small play. One side says models have stalled; the other points to a benchmark just flattened. Astra obliged. It took ARC-AGI-3 from 7.8% to 99.9%, while lifting a broad composite score by 0.3 points. One looks like a revolution, the other like a rounding error. Both can be true. As is a weekly affair now, the cutting edge of AI models has moved again, and we have no way of knowing how much. A benchmark that jumps from failure to perfection has probably been solved. That does not mean the model became twelve times smarter. The model, its surrounding software, and inference budget found a route through the test. Still progress, but less universal than the score suggests. Greg Brockman says the “AGI era” has arrived; leave that there. His useful disclosure: Astra was trained in OpenAI’s first run, using more than 100,000 GPUs. Whatever the marginal return, the machine keeps getting larger. Elsewhere, Nvidia and CrowdStrike adapted a useful security model on 71 B200s. The numbers are not a comparable arithmetic: one trained a frontier model, the other tuned an existing one. The contrast is the story. Creating frontier intelligence is becoming concentrated and capital-hungry; applying it is becoming cheaper and more widely available. Investors need not settle whether AGI is here or scaling is over. We need to notice both economies. One demands power, packaging, networking, cooling and construction; the other lets hundreds of companies build products. Software gets a hardware bill Salesforce told investors this week that internal use of Claude was one reason it had not raised its margin guidance. It had “unleashed Claude” across R&D six months earlier; now management is routing each task to a cheaper model capable of doing it. The story was token discipline. The more consequential fact is that the compute bill was recorded on Salesforce’s income statement before the resulting product revenue. The same bill appeared in three places. Asana’s gross margin fell by 120 basis points sequentially, with 80 basis points coming from AI infrastructure and compute for new products. Zscaler nearly tripled quarterly capital expenditure plus capitalized software as it bought data-center equipment when available; management expects spending to remain elevated as memory, storage and processors become dearer and harder to source. GitLab chose a contractual answer. The old software toll booth has acquired a power meter. SaaS companies have always paid for servers, but AI makes the marginal cost of another useful answer larger, more variable and more visible. Architecture and contract design are now part of product strategy: who chooses the model, routes the task, absorbs a price increase, and pays when usage explodes? Investors who ask only whether AI revenue is growing are asking only half the question. The hard half is who carries the bill, and whether price per task is falling more slowly than cost per task. Task-based pricing: return of ‘solution’ selling My work life started at an IBM outfit, and about the first thing I learned was to sell a “solution.” Back then and for decades until recently, software came in separate compartments. Customers needed someone to connect those pieces to their data and working practices. The word was fashionable; the job was real, creating consulting businesses that still thrive. In between, though, the industry moved to usage-based selling in the name of software as a service. And, in the name of task-based pricing, we are returning to solutions. AI is recreating that opportunity with a new unit of sale. Now, Intercom charges for a defined outcome. HubSpot asks 50 cents per resolved conversation and a dollar per qualified lead. Zendesk bills only verified resolutions. Salesforce offers actions and successful resolutions, and has agreed to buy Fin. The pitch is simple: customers buy the closed ticket, booked appointment or processed claim, leaving tokens where they belong, as a supplier’s input. The attraction is that it transfers the uncertainty. The solution provider selects the model, integrates the customer’s data, routes the task, handles failed attempts and escalates to a human when needed. If models get cheaper, savings can become margin. If they fail repeatedly, the provider pays. Even deciding what counts as success becomes part of the product. That said, the resemblance to the 1990s ends where the supplier begins. Specialist software vendors then had little reason to become integrators. Model providers own the key input, improve it constantly, and are already building agents and orchestration around it. OpenAI is reportedly testing outcome-based contracts with large customers. Hyperscalers, model labs and venture companies everywhere are chasing the model-independent solution layer. Application incumbents must join them without dissolving the old compartments they still charge for. There is plainly a business in assembling models, data, and workflows and standing behind the result. The business is real. The moat may sit one layer lower. When connected things start thinking IoT companies are redesigning themselves around inference. For years, their job was to put a sensor on a thing, connect it and send the data to a dashboard. Recent disclosures from Ambarella, Samsara and Japan’s TDK show how that definition is expiring. Ambarella’s X7 is a stand-alone accelerator that can sit beside its own or someone else’s processor. Samsara is adding agents that interpret operational data and intervene. TDK is combining sensors, software and edge AI to move factory maintenance from observing a machine toward prescribing its next action. The product boundary is changing. The old device reported; the new one interprets and acts. That matters wherever a robot, camera, or production line cannot wait for the cloud, or assume it will always answer. Revenue will lag architecture. X7 is sampling, Ambarella’s channel deals have yet to produce orders, and TDK’s prescriptive layer is still a promise on a conference schedule. Samsara offers early evidence: emerging products supplied over a fifth of net new annual contract value for a third consecutive quarter. A lot remains unproven, but innovation fever is spreading, and an increasing number seem to recognize that the avenue for growth is no longer a bet on lower costs through globalization or on TAM growth driven by factors like demographics. Big pharma’s China week Big Pharma had a China week. Or maybe, better to say, another China week. Roche licensed a trispecific antibody from Simcere. GSK took a first-in-class cancer conjugate from Hutchmed. Genentech added DualityBio’s ADC platform. AstraZeneca completed its acquisition of global rights to Dizal’s lung-cancer drug. Four agreements put $830 million of upfront cash behind molecules or platforms originated in Chinese laboratories. The much larger milestone totals may never be paid; the upfront money records today’s conviction. Then came the evidence around the deals. Akeso said ivonescimab improved overall survival against Keytruda in a Phase III lung-cancer trial, although the full data and international confirmation are still to come. BioPharma Dive updated its count to more than 100 China-to-West licensing agreements since the start of 2025. After visiting Shanghai, Thermo Fisher chief Marc Casper called China “a significant source of biopharmaceutical innovation.” Any one item could pass as another biotech announcement. Together they describe a pipeline changing ownership. This is a deliberately narrow list. It excludes China’s enormous CRO and CDMO businesses. Beyond the frame, United Imaging recently showed new AI-enabled MRI and ultrasound equipment as its products reached more than 100 countries, while Biocytogen this week licensed an antibody generated by its proprietary mouse platform. The activity now runs from research tools and molecular design through trials, licensing, and imaging. The black box grew hands When generative AI arrived, the worry was that nobody could see how it reached an answer. Reasoning models seemed to improve the bargain: the machine showed its work. Agents reopen the problem at a consequential level. Their reasoning may be readable while their activity disperses across tools, computers, and traces. Dwarkesh Patel called the OpenAI-Hugging Face episode the rise and fall of “agent civilizations,” vivid and perhaps too human. The less cinematic version is disturbing enough. Some 1,200 isolated agents found a shared package cache and turned it into a message board, carrying 70,000 messages and files; roughly 700 joined the intrusion. A separate group apparently left 18,000 posts on a German programming wiki during web-lookup tasks, even creating ZZZ pages when a moderator deleted them alphabetically. Almost nobody could see what these scattered actions added up to. The immediate problem is catching a bad action. The harder one is proving that nothing else is happening. A user may let an agent read files, browse, or update a calendar, as this writer does, then wonder what happens beneath the screen. The German message board ran for weeks while people close to it saw only fragments. More monitoring helps, but the Hugging Face record became too large for humans to investigate unaided. Agentic AI needs a credible, continuous account of where each agent went, what it touched, what it left behind, and whether it stopped. No one knows how and when this arrives, but with a few more episodes like the above, the risks of sudden control and regulation will rise exponentially. The morning AI ran out of plan B Thursday morning offered a preview of AI as infrastructure. Claude and Grok failed; ChatGPT joined later, with reports also hitting Gemini and Azure. The verified story is narrower: three confirmed outages overlapped, Google and Microsoft recorded no comparable incidents, and a common cause remains unproven. OpenAI blamed routing; xAI, its Memphis compute center. Anthropic named no cause. The mystery matters less than the experience: the alternatives disappeared together. The internet and power grid learned this lesson early. Packets find alternative paths through redundant routes; grids maintain reserves to cover generation loss. AI has zones and replicas inherited from cloud architecture. Frontier inference differs: Accelerators are expensive, models are not fully interchangeable and displaced workloads are enormous. A second model is not a second route if it shares an upstream dependency, lacks spare capacity or behaves differently enough to break the application. AI is becoming critical infrastructure before it has been engineered like critical infrastructure. It will require independent capacity across regions and suppliers, with edge or on-premises fallbacks for work that cannot stop. That changes the hardware equation. Most forecasts count expected use: users, tokens, and reasoning per query. They may miss the capacity that must stand ready for the exceptional hour. It will look inefficient, because resilience always does before the failure. If Thursday repeats, redundancy becomes another source of demand for GPUs, networking, power, cooling, and powered shells. AI infrastructure has been sized for use. It may still have to be resized for failure. The bottleneck wears a hard hat The chip industry has spent two years explaining its shortages in the vocabulary of physics: EUV tools, CoWoS packaging, HBM stacks and substrates. This week TSMC changed the vocabulary. Cliff Hou said equipment needs had nearly doubled since December; almost 20 fabs were under construction, four times its old rhythm, yet capacity still fell short. The newest constraint is skilled construction labor, in Taiwan and Arizona alike. Capital can approve a cleanroom. It cannot train the people who install its power, cooling and gas lines on the same timetable. The same message is traveling. A US survey found data-center projects increasing competition for skilled workers at 58% of contractors; August construction pay rose about 5% year on year. Australian builders warned that data centers are pulling electricians from housing, grids and defense. Wales has 13% of recent British data-center applications and 4% of the electrical-installation workforce. In Korea, contractors say only 1–2% of the mechanical-installation workforce has specialist experience for fab and data-center piping, welding and cooling. The digital build has a physical floor. Every new fab or data center must be wired, cooled and commissioned by people who take years to train and are already needed elsewhere. The hardware bottlenecks are creating a human one. That pushes AI spending into wages, contractor backlogs, tools, modular construction and training. The last scarce component may wear a hard hat, and the companies that train and equip those workers sit in the economic shadow. This is one of our core “epicenter” themes. Nvidia scales across itself Nvidia describes an AI factory in three dimensions. Scale-up joins accelerators inside a rack, scale-out joins racks and scale-across joins distant data centers. If earlier, CUDA made a GPU harder to replace, its simultaneous expansion into multiple other architectural points is the new moat for the coming years. This is well known. What is less discussed is how the visionary company is scaling across other corporates and industries, locally and globally, in a way almost no one has attempted before. This week Nvidia put $3.5 billion into MediaTek convertibles, bringing the Taiwanese designer into NVLink Fusion, local AI computing and cars. It also agreed to buy Hugging Face, entering a model-and-developer marketplace used by 18 million people. Its partnerships with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR aim to mobilize over $500 billion for AI infrastructure. The fried-chicken dinners in Seoul looked like Jensen theatre; around them came partnerships with Korea’s government, Samsung, SK, Hyundai and Naver covering more than 260,000 GPUs. The balance sheet and financing network are traveling with him. The conglomerate comparison would go too far. The company seems to have made a conscious decision to choose the route that involves no acquisitions and thus avoid scrutiny and other associated issues. These activities are related, and ownership is often optional. The closer analogy is a hub-and-spoke keiretsu, held together by capital, interfaces and mutual dependence. That matters as competition reaches the GPU itself. A rival accelerator may win a benchmark and still face a rack, network, model ecosystem and financing map increasingly organized around Nvidia. The innovation engine gets plenty of attention. The quieter strategy makes every adjacent winner strengthen the system it might otherwise challenge. Nvidia’s most underappreciated product may be its sphere of influence. Memory needs to learn naming Samsung is doing some of the most interesting work in AI hardware. Its HBM4 puts a 4-nanometer logic die beneath the memory stack. Custom HBM, now aka cHBM, tunes that stack to a customer’s accelerator by working on the B-Die, while also modifying C-Die. The proposed aHBM moves some attention work into memory’s cHBM and bHBM, while zHBM eventually stacks the memory above the processor. The trouble begins on the slide. Readers must distinguish HBM4 from HBM4E and HBM5, then cHBM, aHBM and zHBM. Samsung’s packaging adds HPB, HCB, TCB, I-Cube and X-Cube. SK Hynix offers its own cHBM, iHBM, AiMX, CuD and CMM-Ax. Somewhere in Korea, a genuinely radical idea is always one set of initials away from sounding like an incremental standards update. Taiwan does slightly better. CoWoS is alphabet soup too, but years of shortages have turned it into something close to a brand, with COUPE following on its heels. Nvidia and AMD understand the difference. Blackwell, Vera Rubin, Kyber and Helios make people curious before the technical slide begins; HBM4E 16H tells them to find a glossary. We are astounded by the amazing things being proposed, including this week, but for the world to tune in more, the likes of Samsung really need to learn better ways of communication. Nilesh Jasani is the founder and CEO of GenInnov Pte Ltd Singapore, which originally published this article. The article is republished with permission.

Original Source

Read the full article at Asiatimes →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.