Key Takeaways: Bloomberg Intelligence projects AI inference will reach $1.3 trillion by 2032, doubling the training market and redrawing the chip competitive map.
Key Takeaways: Bloomberg Intelligence projects AI inference will reach $1.3 trillion by 2032, doubling the training market and redrawing the chip competitive map.

AI inference is the fastest-growing slice of the AI infrastructure market, projected by Bloomberg Intelligence to reach $1.3 trillion by 2032 — double training — as Nvidia, Cerebras, and AMD race for share.
Bloomberg Intelligence projects inference will double the size of the AI training market by 2032, reaching $1.3 trillion, according to the research firm's forecast. The projection has pushed chipmakers to re-engineer around memory access and latency rather than raw compute.
Nvidia (NVDA), the training-market leader with a $5.3 trillion market cap, acquired Groq and its language processing units (LPUs), which embed SRAM (static random-access memory) directly on chips. Its systems pair GPUs for the pre-fill phase — reading the prompt — with LPUs for the decode phase, when models answer queries. Cerebras (CBRS), valued at $66 billion, builds wafer-sized chips five to six times faster than LPUs, though they require specialized cooling and sell only as part of its CS-3 systems. AMD (AMD), at a $774 billion market cap, uses a chiplet design that packages GPUs with more high-bandwidth memory (HBM) to cut latency.
The stakes are large because inference is where AI models generate responses — the most compute-intensive and costly part of deployment. Nvidia, trading at 34.3 times trailing earnings, has the scale advantage; Cerebras brings speed; AMD brings cost efficiency through its Helios rack-scale systems and recent acquisitions of memory-optimization firm MEXT and chip start-up Taalas.
Nvidia's LPU Bet Targets the Decode Bottleneck
Nvidia's move to fold Groq's LPUs into its lineup addresses the decode phase of inference, when large language models (LLMs) assemble answers token by token. Because SRAM sits on the chip, the LPUs avoid slower trips to external memory that throttle conventional GPUs. The result is a complete system where Nvidia's GPUs read the prompt and its LPUs generate the response, cutting latency at the point users feel it most.
The approach keeps Nvidia anchored in AI infrastructure as the market shifts from training to inference. Its data center segment generated $75.2 billion in revenue in the fiscal first quarter, up 92 percent from a year earlier, and the company is guiding to about $91 billion in total revenue for the fiscal second quarter, a 95 percent jump. Nvidia is also preparing its Vera Rubin systems, which the company says can cut inference token costs by 90 percent versus the Blackwell platform.
AMD and Cerebras Team Up to Cut Inference Costs
Cerebras has signed deals with OpenAI and Amazon's AWS, but its wafer-scale engines carry a premium price and demand specialized cooling and power management. A new partnership with AMD pairs AMD's Helios rack-scale solution — which handles the pre-fill phase more cheaply — with Cerebras' Wafer-Scale Engine for the faster decode phase, a combination aimed squarely at Nvidia's integrated systems.
AMD is also attacking inference from other angles. Its acquisition of MEXT targets the memory bottleneck by offloading seldom-accessed data from DRAM to unused flash, then using predictive AI to pull it back before an application requests it — reducing the need for expensive DRAM. The Taalas acquisition brings chips with AI models hardwired directly onto silicon, faster and cheaper for specific models, though less flexible. AMD plans to pair its GPUs for pre-fill with Taalas chips for decode.
For investors, the three approaches carry different risk profiles. Nvidia trades at 34.3 times trailing earnings, a discount to its 10-year average of 61.6, giving it room if inference demand accelerates. Cerebras, at a $66 billion market cap, is the purest play on speed but carries execution risk as it scales beyond niche deployments. AMD, at $774 billion, offers the broadest portfolio but must prove its inference stack can close the gap with Nvidia's integrated systems.
This article is for informational purposes only and does not constitute investment advice.