Citi's latest AI Tracker shows model intelligence accelerating faster than the infrastructure needed to run it, driving costs higher across the stack.
Citi's latest AI Tracker shows model intelligence accelerating faster than the infrastructure needed to run it, driving costs higher across the stack.

Citi's latest AI Tracker shows model intelligence accelerating faster than the infrastructure needed to run it, driving costs higher across the stack.
AI model intelligence is accelerating faster than the infrastructure can support it: the median top-20 score rose 9 points in six weeks, while Blackwell GPU rentals climbed 27% year to date, Citi Research reported July 24.
"The AI industry's return on investment is accelerating toward the infrastructure layer, and the next competitive moat will shift from compute acquisition to efficient output and proprietary data," the Citi analysts said.
Kimi K3, the largest publicly released open-source model, scored 57 points — third globally behind Anthropic's Claude Fable 5 at 60 and OpenAI's GPT-5.6 Sol at 59. The prior open-source high was 51 points. DeepSeek V4 Pro scored 44 at $0.04 per million tokens, while Xiaomi's MiMo-V2.5-Pro scored 42 at $0.03 — near-frontier intelligence at two orders of magnitude lower cost. Google released Gemini 3.6 Flash on July 21 while Gemini 3.5 Pro remains in testing and Gemini 4 pre-training has already begun, the report noted. The "Flash first, Pro delayed" release cadence signals that frontier model development is becoming harder, consistent with infrastructure bottlenecks across the industry.
The findings point to sustained demand for GPU manufacturers, power infrastructure companies, and data center operators. But the report also warns that pure compute accumulation is losing its edge as a competitive differentiator. Labs that control proprietary data and achieve higher inference efficiency will build more durable moats, the analysts said.
The Bottleneck Has Moved
Trillion-parameter models now spend more time moving weights and KV-cache data across HBM memory and GPU interconnect networks than performing matrix calculations, Citi found. The bottleneck has shifted from FLOPs — how fast chips compute — to memory bandwidth, GPU interconnect, and power supply.
Some labs are bypassing the grid entirely. SpaceX purchased $1 billion worth of 1-gigawatt mobile turbine generators on July 15, and Georgia Power signed a service agreement on July 22 — moves that would have been unthinkable a year ago, the report said.
Chinese frontier model pricing jumped 45% week over week and month over month, from $0.60 to $0.87 per million tokens — the first significant move in two months. Global frontier model pricing rose 6.8% weekly and 11.3% monthly. US and European pricing was relatively stable, down 0.5% weekly but up 4.1% monthly. Citi expects pricing pressure to ease as incremental infrastructure comes online, making proprietary data and task-specific performance the more durable competitive advantages.
Agents Get Smarter, and Riskier
The report described a sharp escalation in autonomous AI agent risks. In April, the concern was agents "interrupting lunch." By July, the scenario had escalated to agents "infiltrating Hugging Face infrastructure," Citi said. Independent testing also showed Kimi K3 underperforming US frontier models in planning cyberattacks, highlighting the uneven safety profile of rapidly advancing open-source models.
METR data from July 21 showed the economic gap between AI agents and human workers narrowing, which will accelerate token consumption and further strain infrastructure. Meanwhile, model providers still cannot fully explain why their models behave as they do — the interpretability problem remains unsolved, the report noted.
Nvidia commented on AI regulation on July 24, and US Treasury Secretary Scott Bessent weighed in on July 22, but no regulatory framework is imminent. Even without formal rules, enterprises running proprietary models face rising compliance costs.
For investors, the report reinforces a rotation toward infrastructure plays — Nvidia, AMD, power and cooling providers, and data center REITs — while raising questions about the sustainability of pricing power among model providers. Nvidia shares, trading at elevated multiples on data center revenue expectations, face a paradox: demand is surging, but the bottleneck is no longer just about buying more GPUs. The companies that solve the memory bandwidth and power constraints — or own the proprietary data that makes models valuable — will capture the next phase of value creation.
This article is for informational purposes only and does not constitute investment advice.