A 27-billion-parameter AI model now runs entirely on an iPhone, compressing a data-center-scale workload into 3.9 gigabytes of local memory.
A 27-billion-parameter AI model now runs entirely on an iPhone, compressing a data-center-scale workload into 3.9 gigabytes of local memory.

A 27-billion-parameter AI model now runs entirely on an iPhone, compressing a data-center-scale workload into 3.9 gigabytes of local memory.
On-device AI crossed a threshold July 14 when Caltech-founded PrismML released Bonsai 27B, a 3.9-gigabyte model that runs on Apple's iPhone 17 Pro while keeping 90 percent of the original's performance.
"Apple is in very early discussions about the technology," PrismML's CEO told CNBC, declining to specify a timeline for any integration. The startup's advance is not compression itself — shrinking open models is routine — but doing so without the performance collapse that has plagued prior attempts.
Bonsai is built on Alibaba's open-weight Qwen3.6, a 27-billion-parameter architecture that would normally exceed a phone's usable memory. At 3.9 gigabytes, it runs on the iPhone 17 Pro, iPad, Mac, and PCs. Apple's A19 and A19 Pro chips in the latest iPhones feature dedicated neural accelerators, which the company says deliver a significant boost to AI performance. On its second-quarter earnings call in April, Apple management described the Mac as "the best platform for AI," with silicon capable of running advanced models on-device. The compression breakthrough matters because it removes the last barrier to local inference: a model that is both small enough to fit in device memory and smart enough to be useful.
The stakes extend beyond consumer convenience. OpenAI surpassed 900 million weekly active users in February with roughly 50 million paying subscribers — a conversion rate near 5.5 percent. If free, open-weight models running locally on devices continue to improve, the frontier labs' consumer subscription model faces structural pressure. Apple's position differs: its business sells hardware, so a local-AI wave means more memory and capacity per device, raising cost of goods sold but strengthening the upgrade cycle.
The A19 and A19 Pro neural accelerators were designed for on-device inference, and Apple's own AI execution has been uneven. The rebuilt Siri fell short in internal testing in February, though management sounded optimistic on its July 30 earnings call. Siri AI has been in public beta for several weeks with positive early feedback. In July, Apple sued OpenAI in federal court, alleging trade-secret theft tied to former engineers who joined the lab, including claims that secrets were taken to help OpenAI build its own devices.
The lawsuit puts the competitive tension between Apple's device-first approach and OpenAI's cloud-subscription model into sharp relief. Apple's silicon advantage — the neural accelerators in A19 and A19 Pro — was designed precisely for the kind of local inference that Bonsai enables. If third-party open-weight models can run effectively on Apple hardware, the company's internal foundation model work becomes less critical to its AI story, and the hardware itself becomes the differentiator.
Most consumer AI usage is free, and open-weight alternatives are eroding the paid tier. Apple's lawsuit against OpenAI puts the consumer AI fight front and center, though the company remained silent about the litigation during its earnings call. For Apple, the opportunity is not subscription revenue but device differentiation: capable local AI that runs offline, keeps data on-device, and requires no monthly fee. Apple shares closed at $308.91 on Friday, down 7.35 percent, with a market capitalization of $4.9 trillion, as investors weighed surging memory costs against the on-device AI opportunity. With breakthroughs such as PrismML's, consumers may not have to wait much longer for capable AI that runs offline and keeps data on the device without a subscription.
This article is for informational purposes only and does not constitute investment advice.