OpenAI's Data Agent Has No Benchmark: Write Your Own

You have a question like "which customers churned last quarter, by plan?" and nobody who can answer it without a day of SQL. Now every vendor is selling an agent that claims to fix that. OpenAI's newest one, the Data agent in ChatGPT Work, launched without a published accuracy number, so the only way to know if any data agent works on your data is to measure it yourself.
This post covers what OpenAI shipped, what the missing benchmark does and doesn't tell you, and how to build a small evaluation harness that works for any data agent, whether you buy one or build one.
What OpenAI actually launched
OpenAI announced the Data agent in ChatGPT Work on September 10, 2026. It is a plugin that connects to approved company data sources and builds shareable, interactive dashboards.
The details from OpenAI's announcement:
- Data sources listed: Amazon Redshift, Datadog, Google BigQuery, ClickHouse, Databricks, MongoDB and Snowflake, "and more."
- Dashboard tools it can build and interact with: Omni, Oracle BI, Power BI, Sigma, Tableau and ThoughtSpot.
- Permissions: per an explainx.ai write-up of the announcement, it enforces the connected account's existing permissions, including table, row and column restrictions. Admins choose which connections exist and which roles can use them.
- Access: it appears as "Data" in the ChatGPT Work Plugins directory. Admins install it through Workspace settings, and users start a conversation with
@Data, according to Unite.AI. - Early customers in the Alpha program: NTT Data, Thermo Fisher and ServiceTitan.
I'm not stating pricing or plan tiers because I couldn't confirm which plans include the Data agent. Check OpenAI's official help documentation for the current plan requirements before you plan around it.
I'm also not naming an underlying model. I couldn't confirm one in OpenAI's announcement, so check the announcement directly for the latest details.
The missing benchmark: what it means and what it doesn't
No published accuracy figure means you can't compare this agent to others on paper. It does not mean the agent is inaccurate. VentureBeat reported that OpenAI's Data agent has no published benchmark, while Databricks released one the same week.
OpenAI does have an internal benchmark that compares the agent's results against OpenAI's own data tools, but it hasn't published a figure. That makes it an internal check, not an independently verified benchmark. Instead of an accuracy claim, OpenAI's Arpan Shah describes a process called "hill climbing": iterative internal refinement that shaped the tool before any customer used it.
That's a legitimate engineering practice. It's also not evidence you can verify.
Don't over-read the Databricks number either
The same week, on September 9, 2026, Databricks introduced Adaptive Instructed-Retriever. Per ITBrief, it reported an average end-to-end latency of 5.8 seconds, which it said was more than twice as fast as the comparison models. SiliconANGLE reports the comparison set was Claude Sonnet 5, GPT-5.6 Luna and DeepSeek-V4-Flash, across seven held-out internal and external benchmarks. Databricks' own blog says the model matches leading third-party models at 2x lower latency.
Three caveats before treating that as a scoreboard:
- These are Databricks' own results, not independently verified.
- It's a retrieval model benchmark. It is not a head-to-head with OpenAI's Data agent.
- Latency and "quality parity" on Databricks' benchmark sets says little about whether the agent gets your revenue definition right.
| Claim | Source of claim | Independently verified? | Comparable to OpenAI's Data agent? |
|---|---|---|---|
| Data agent internal benchmark vs OpenAI's own tools | OpenAI (via VentureBeat) | No, no figure published | n/a |
| 5.8 s average latency, 2x faster | Databricks | No | No, it's a retrieval model |
| Quality parity with frontier models | Databricks | No | No |
The benchmark you should worry about is a different one
Databricks' earlier OfficeQA benchmark used 246 questions over roughly 89,000 pages of Treasury Bulletins. VentureBeat reported that even the best agents scored below 45% accuracy on those enterprise-style document tasks.
That result is about documents, not SQL, so it doesn't transfer directly. But it's a useful reminder that agents which look strong on abstract tests can stall on messy real-world material. Your warehouse is messy real-world material.
Why "built it for ourselves first" matters more than the benchmark
The most useful fact in this story isn't a number. It's the origin. OpenAI's own employees couldn't get fast answers to basic data questions, so the company built a fix for itself and then decided to sell it.
The evidence is in OpenAI's January 2026 post on its in-house data agent. It says the platform serves more than 3.5k internal users across over 600 petabytes of data and 70k datasets. It calls the tool a custom internal-only tool, "not an external offering."
Attribution matters here: those figures describe the internal data platform and its agent, not the externally released Data agent. Don't read them as adoption numbers for the product.
In the September announcement, OpenAI says nearly all of its product team and over two-thirds of its go-to-market organization use data agents in ChatGPT Work. It credits an internal data team that set shared business definitions, access rules and safeguards for sensitive data.
Read that last part again. The credit doesn't go to a model. It goes to definitions, access rules and safeguards. Those are the parts you can build yourself, at any company size.
I want to be careful about the generalization. OpenAI is one documented case of "build internally, then externalize." I haven't found a source showing it's a universal pattern, so treat it as one credible data point, not a law. What I can say from building these systems is that the unglamorous layer under the model is where most of the accuracy comes from.
Build your own benchmark in an afternoon
A useful data-agent eval is a spreadsheet of real questions with known-correct answers, run repeatedly, scored automatically. You don't need a research team. You need 30-50 questions and a read-only database connection.
Step 1: collect real questions
Pull them from Slack, email, and your own head. Take the last month of "can you pull..." requests. Include the ugly ones: ambiguous time ranges, refunds, test accounts, timezone edges. Those are where agents fail.
Step 2: write each question with several phrasings
Databricks' Genie Agent documentation, in its guide to testing and monitoring a Genie Agent, lets teams add benchmark questions and recommends two to four phrasings of the same question. That's good advice regardless of vendor, because people don't phrase things the way you'd like.
# eval/questions.yaml
- id: churn_by_plan_q2
phrasings:
- "How many customers churned last quarter by plan?"
- "Churn by plan, previous quarter"
- "Which plans lost the most customers in Q2?"
gold_sql: |
SELECT plan, COUNT(*) AS churned
FROM subscriptions
WHERE canceled_at >= '2026-04-01' AND canceled_at < '2026-07-01'
AND is_test = false
GROUP BY plan
ORDER BY plan
notes: "Excludes test accounts. Quarter is calendar, not fiscal."
Note the notes field. When the agent gets this wrong, you'll want to know why the gold answer is what it is. Half the value of this exercise is discovering your team disagrees about what "churned" means.
Step 3: score results, not SQL text
Two different queries can return the same correct answer. Compare result sets, not strings.
# eval/run_eval.py
import yaml, json, sqlite3 # swap for your read-only warehouse driver
from pathlib import Path
def run_sql(conn, sql):
cur = conn.execute(sql)
rows = cur.fetchall()
return sorted(map(tuple, rows)) # order-insensitive compare
def ask_agent(question: str) -> str:
"""Call your agent (vendor API, MCP tool, custom) and return SQL it ran."""
raise NotImplementedError
def main():
conn = sqlite3.connect("file:warehouse.db?mode=ro", uri=True) # read-only
cases = yaml.safe_load(Path("eval/questions.yaml").read_text())
results = []
for case in cases:
gold = run_sql(conn, case["gold_sql"])
for phrasing in case["phrasings"]:
try:
got = run_sql(conn, ask_agent(phrasing))
ok = got == gold
err = None
except Exception as e:
ok, err = False, str(e)
results.append({"id": case["id"], "q": phrasing, "pass": ok, "err": err})
Path("eval/last_run.json").write_text(json.dumps(results, indent=2))
passed = sum(r["pass"] for r in results)
print(f"{passed}/{len(results)} passed")
if __name__ == "__main__":
main()
Illustrative math with round numbers: 40 questions with 3 phrasings each is 120 runs per pass. If 84 pass, that's 70%. The number itself matters less than the failing 36 rows, which tell you what to fix.
A caveat: some agents return dashboards or prose, not SQL. In that case, extract the underlying numbers and compare those. You lose some automation but keep the principle. If a vendor tool gives you no way to see what it computed, that's a finding in itself.
Step 4: track the same set over time
Keep last_run.json in version control. When a vendor updates their model, when you change a metric definition, when you add a table, rerun. This is "hill climbing" in the boring, honest sense: change one thing, measure, keep or revert.
Fix definitions before you touch the model
When a data agent gets an answer wrong, the cause is usually a missing definition or a wrong table, not the language model. OpenAI credits its internal data team's shared business definitions, access rules and safeguards for the tool's uptake. Copy that priority.
Write down your metrics once, in a place the agent can read:
# definitions/metrics.yaml
metrics:
active_customer:
definition: "Paying subscription with no cancellation date and is_test = false"
table: subscriptions
owner: finance
mrr:
definition: "Sum of monthly-normalized plan price; annual plans divided by 12"
table: subscriptions
exclude: ["trial", "refunded"]
owner: finance
churned_customer:
definition: "Subscription with canceled_at set in the period"
table: subscriptions
notes: "Downgrades are NOT churn"
Then sort your eval failures into buckets. This tells you where the effort goes:
| Failure type | Typical cause | Fix |
|---|---|---|
| Wrong number, plausible query | Metric definition ambiguity | Add to definitions file |
| Query errors | Agent doesn't know schema | Add table/column descriptions |
| Right table, wrong filter | Test accounts, soft deletes, timezones | Add explicit exclusion rules |
| Different answers per phrasing | Underspecified question | Add a clarification rule |
| Sees data it shouldn't | Over-broad connection | Tighten permissions |
Permissions are not optional
Connect the agent with a read-only role scoped to the tables it needs. OpenAI's Data agent inherits the connected account's table, row and column restrictions, per the explainx.ai summary. That's the right model: the agent should never have more access than the person asking. If you build your own, enforce this at the database, not in the prompt. A prompt saying "don't show salary data" is a suggestion. A column-level grant is a rule.
Buy, build, or mix: a decision guide
Choose based on where your data lives and how much of the definitions work you've already done, not on whichever vendor has the best launch week.
| Situation | Reasonable path |
|---|---|
| Data already in Snowflake/BigQuery/Databricks etc., you already use ChatGPT Work | Try the Data agent on a scoped, read-only connection. Run your eval set against it before rolling out. |
| Heavy Databricks shop | Look at Genie Agents. Their docs support adding benchmark questions directly. |
| Data in Postgres plus spreadsheets plus a SaaS tool or two | A small custom agent with a definitions file and eval harness is often simpler than forcing everything into a warehouse. |
| Sensitive data with strict access rules | Start with permissions design first. Any tool is only as safe as the account it connects with. |
| No one has written down what "active customer" means | Stop. Do that first. No agent fixes this. |
Whichever route you pick, the eval set is the constant. It costs nothing to keep and lets you compare vendors on your own data instead of on their slides.
Here's a sensible order of operations:
- Write 30-50 real questions and gold answers (a day).
- Write the metric definitions file (half a day, plus an argument or two).
- Create a read-only, scoped database role.
- Run the eval on each candidate: vendor agent, custom agent, or both.
- Fix the failures that are definition problems. Only then compare tools.
- Roll out to a small group of users. Log every question and answer, and add the wrong ones to the eval set.
How BizFlowAI approaches this
I build AI automations for solopreneurs and small teams, and I don't trust any output I haven't measured.
I'm not going to claim benchmark numbers for a system I haven't measured on your data. The process above (pick the questions, write the gold answers, set the permissions, rerun after every change) is the part you can start on today. If you want help with automation, get in touch.
Work with BizFlowAI
If you'd rather have this built for you, that's what we do: production AI automation for solo founders and small teams — agents, integrations, and document pipelines that actually ship.
Book a free discovery call — 30 minutes, we map the highest-ROI automation in your workflow. No pitch deck, just engineering.
More guides like this on the BizFlowAI blog.
Frequently asked questions
Does OpenAI's Data agent in ChatGPT Work have a published accuracy benchmark?
No. OpenAI launched the Data agent in ChatGPT Work on September 10, 2026 without a published accuracy figure. It has an internal benchmark comparing the agent's results against OpenAI's own data tools, but no number has been released and nothing is independently verified. Lack of a benchmark doesn't mean the agent is inaccurate, only that you can't compare it on paper.
Which data sources and dashboard tools does OpenAI's Data agent support?
OpenAI lists Amazon Redshift, Datadog, Google BigQuery, ClickHouse, Databricks, MongoDB and Snowflake as data sources, plus 'and more.' It can build and interact with dashboards in Omni, Oracle BI, Power BI, Sigma, Tableau and ThoughtSpot. Reportedly it enforces the connected account's existing table, row and column permissions. Admins choose which connections exist and which roles can use them.
How do I test whether a data agent works on my own data?
Build a small evaluation harness: collect 30-50 real questions from Slack, email and past requests, including messy edge cases like refunds, test accounts and timezones. Write a gold SQL query for each, plus two to four phrasings of the question. Run the agent against a read-only database connection and compare result sets rather than SQL text. Repeat runs regularly to track accuracy.
Can I compare Databricks' Adaptive Instructed-Retriever results to OpenAI's Data agent?
No. Databricks reported 5.8 seconds average end-to-end latency, more than twice as fast as comparison models, and quality parity with leading third-party models. These are Databricks' own unverified results on a retrieval model, not a head-to-head with OpenAI's Data agent. They say little about whether an agent gets your business definitions right.
Why score data agents on result sets instead of SQL text?
Two differently written SQL queries can return the same correct answer, so comparing strings produces false failures. Comparing sorted result sets checks whether the agent produced the right data regardless of how it got there. Keep a notes field with each gold query explaining definitions, such as whether a quarter is calendar or fiscal, so you can diagnose why a failure happened.