Employees can now ask a business question in plain English and get an investigated answer, a shareable dashboard and a suggested next step, after OpenAI shipped a Data agent inside ChatGPT Work on Sept. 10. What the company did not ship is a number telling buyers how often the agent gets the answer right.
The agent reaches into Amazon Redshift, Google BigQuery, ClickHouse, Databricks, MongoDB, Snowflake and Datadog, pulls files from Google Drive and SharePoint, and writes back into Omni, Oracle BI, Power BI, Sigma, Tableau and ThoughtSpot. It investigates why a metric moved rather than returning a query, and can execute follow-up actions once a user approves them.
"It combines all the various contexts in their state, as opposed to generating a net new context layer itself," Arpan Shah, general manager of the enterprise technology vertical at OpenAI, told VentureBeat. The agent stitches context together on the fly rather than maintaining a persistent layer across a customer's tools.
The internal version it was generalized from served more than 3,500 users across roughly 70,000 datasets and over 600 petabytes of data, OpenAI disclosed in a January 2026 engineering post. Nearly all of OpenAI's product team and more than two-thirds of its go-to-market organization now use it, the company said. That scale is the sales argument: the agent was stress-tested on OpenAI's own warehouse before any customer touched it.
The benchmark gap Databricks is already filling
For a product whose entire value is answering questions about enterprise data, OpenAI has not published a retrieval accuracy or correctness figure that external customers can compare against rivals. Shah said the internal benchmark checks the agent's results against OpenAI's own data tools to confirm parity with what the company expects externally. That is an internal comparison, not an independently verified number.
The gap matters because enterprise data questions are harder than general chat. An answer may sit across several tables, a BI dashboard, a document and a Slack thread, with different teams defining the same metric differently. Without a published score, a buyer cannot tell whether the agent resolves 60 percent of those questions or 95 percent.
Databricks moved in the opposite direction this week, publishing research claiming its Adaptive Instructed-Retriever matches the answer quality of Claude Sonnet 5, GPT-5.6 Luna and DeepSeek-V4-Flash while responding in an average of 5.8 seconds — more than twice as fast. Those results come from Databricks' own testing and have not been independently verified either, but they give procurement teams a number to argue over.
OpenAI's substitute is a claim about process. Shah described it as hill climbing: iterative internal refinement that made the company's models better at data work before the product shipped. "We've gone through the journey internally of hill climbing, and that in turn has made our models really good at this stuff already," he said.
NTT Data and ServiceTitan test the self-serve pitch
Alpha customers are running the agent on sales, spending and reporting work. NTT Data, Thermo Fisher, ServiceTitan, Zipline, Empower, CookUnity and Turing are among the named participants.
Yuji Shono, head of NTT DATA Group's Global AI Office, said licensing costs and technical expertise had limited how far dashboards could spread inside his organization. With the agent, he said, "many non-engineers, particularly in sales and corporate functions, have been able to build and update their own dashboards using plain language."
ServiceTitan's Ankur Bhatt said his team used it to show that users of Atlas, the company's AI sidekick, launched campaigns at roughly three times the rate of nonusers — a finding he said is simplifying onboarding. Zipline co-founder Ryan Oksenhorn said early testing surfaced findings that would otherwise have taken one of the company's best people hours to dig up.
Governance is the constraint that decides how far this spreads. Administrators choose which connections and roles are available, and queries inherit existing table-, row- and column-level permissions. Snowflake's head of developer experiences, Umesh Unnikrishnan, said employees can draw on Snowflake data under their existing access controls.
OpenAI is not positioning the agent as a replacement for Databricks, Snowflake or Tableau. Shah framed it as orchestration across tools rather than depth in any one, and the partner statements from AWS, ClickHouse, Databricks, Snowflake, MongoDB, Tableau, Microsoft and ThoughtSpot read as endorsements rather than defenses. The bet is that as the number of tools a company runs grows, spanning all of them beats mastering one.
The investor question is whether that layer captures value or compresses it. If a natural-language agent becomes the default entry point to enterprise data, the BI seat becomes a commodity and pricing power shifts toward whoever owns the question — a direct read-through for Tableau parent Salesforce, Microsoft's Power BI franchise and Snowflake's consumption model. Databricks, privately held and last valued in a 2025 funding round, is the loudest voice arguing that retrieval quality, not interface, is the defensible asset.
OpenAI has scheduled a webinar for Sept. 10 at 9:30 a.m. PT on how its own data team uses ChatGPT Work. The next thing to watch is whether the company publishes an external benchmark before enterprise renewal cycles force the comparison anyway.
This article is for informational purposes only and does not constitute investment advice.