Morgan Stanley's three frameworks show hyperscaler AI inference investments can generate returns that far exceed market expectations, challenging the narrative that $1.4 trillion in cumulative capital expenditure will go unrewarded.
Morgan Stanley analysts Brian Nowak, Stephen C Byrd and Adam Wood published three bottom-up return-on-invested-capital frameworks on July 27 covering the generative AI inference value chain, and each points to incremental ROIC well above the single-digit figures some investors have feared. The most profitable model — first-party model API services run on owned infrastructure — delivers an estimated 46% ROIC, while GPU leasing through hyperscaler IaaS platforms returns roughly 31% and third-party model APIs running on rented compute yield about 25%.
"The market has underestimated the long-term return potential of the inference phase," the analysts wrote, noting that the three frameworks all show "compelling incremental unit economics" even under conservative pricing assumptions.
The report arrives as the four largest hyperscalers — Amazon.com Inc., Alphabet Inc., Microsoft Corp. and Meta Platforms Inc. — face mounting investor scrutiny over a capital expenditure trajectory that Morgan Stanley estimates will exceed $1.4 trillion in aggregate, with compute capacity quadrupling to roughly 120 gigawatts between 2025 and 2028. A significant portion of that spending funds model training across frontier, mid-size and small models, with monetization expected to come primarily through inference services.
Three frameworks, three return profiles
The first framework models a hyperscaler's GPU-as-a-service IaaS business using Nvidia Corp.'s GB300 GPU as the reference chip. A single 1-gigawatt data center housing roughly 41,000 GB300 GPUs at 75% utilization and an $8.50 hourly rental rate generates about $229 billion in annual revenue per gigawatt. After $76 billion in operating costs — including $50 billion in IT equipment depreciation, $10 billion in non-IT facility depreciation and $20 billion in energy and other expenses — the business yields a roughly 67% incremental EBIT margin and a 31% ROIC. In a sensitivity analysis with rental prices ranging from $7 to $10 per hour, ROIC spans 23% to 39%.
The second framework examines model developers — such as Google's Gemini, Meta's recently released API and various private model labs — that offer API access through their own data centers. With 65% of compute allocated to inference, each GPU processing 2,750 tokens per second and a blended token price of $1.75 per million tokens, annual revenue reaches roughly $304 billion per gigawatt. The same $76 billion cost structure produces a 75% incremental EBIT margin and a 46% ROIC. When token pricing varies from $1 to $2.50 per million tokens and inference compute share ranges from 50% to 80%, ROIC expands from 19% to 63%.
The third framework covers model providers that rent compute from hyperscalers or other third parties. At a GPU lease rate of $7.75 per hour, the model generates about $405 billion in annual revenue per gigawatt but pays $279 billion in compute rental costs — the "middleman margin" that compresses returns. Incremental EBIT margin falls to roughly 31%, with NOPAT margin at 25% and ROIC at about 25%. At a token price of $1 per million, the business turns negative with a NOPAT margin of negative 9%.
Who wins, who competes
Morgan Stanley maintained bullish ratings on Amazon, Google, Microsoft and Meta, arguing that the returns justify continued investment and that data center capacity is becoming a strategically scarce asset. The analysts also noted that healthy returns will attract new entrants and open-source alternatives, which in turn will intensify pressure on product innovation and compute efficiency — two variables the report identifies as the primary levers for token pricing and token throughput.
For Nvidia, the GB300 serves as the reference architecture across all three frameworks, meaning any shift in GPU pricing or utilization directly affects the ROIC calculations. The report's base case assumes GPU rental pricing holds near $8.50 per hour, a level that supports attractive returns for hyperscalers while leaving room for negotiation with third-party tenants.
Amazon shares trade at roughly 22 times forward earnings, while Microsoft and Alphabet trade at similar multiples. Meta, which has signaled aggressive AI infrastructure spending, trades at approximately 24 times forward earnings. The Morgan Stanley analysis suggests that if inference ROIC materializes in the 25% to 50% range, current valuations may not fully reflect the long-term earnings power of these AI investments.
This article is for informational purposes only and does not constitute investment advice.