A company named dbt ran a benchmark this year where the same question hit a hundred percent accuracy through a semantic layer and eighty-four percent through raw text-to-SQL. Sixteen points. Not because the model got smarter, but because somebody had written down what the tables meant.
Which is the whole episode, right there.
It's the whole episode, and it's the thing every product deck in the industry is currently skating past. Here's what Daniel wrote in this week. He's watching every major AI platform race to ship data connectors, and the phrase "talk to your data" is now in every deck. But he wants the actual implementation, not the marketing. He's noticed everything in agentic AI is standardizing around MCP, and he's noticed that talking to data involves some very specific technical things that the protocol doesn't obviously cover. For SQL, he breaks it into three pieces. The model needs general relational principles. It needs a semantic understanding of what the tables are for, before it can answer even a simple question. And then there's the execution process: natural language in, valid SQL out, results parsed and interpreted, maybe more queries sent, everything translated back to language.
And then he asks three things.
Three things. How does the query-forming process actually happen under the hood. Whether specialized natural-language-to-SQL models still play an important role, or whether general models have moved past needing them. And because we can't cover every database with the same answer, how the challenge changes when the target is a document store or a graph store.
So let's start with what the pipeline actually looks like, because "talk to your data" hides five distinct jobs.
Name them.
Schema ingestion first. You have to get the tables, columns and relationships in front of the model. Then semantic grounding, which is inferring what those things mean. Then query generation. Then execution with error feedback. Then interpretation back to natural language. Five stages, and only one of them is the thing people mean when they say "the AI writes SQL."
Stage two is where the bodies are buried.
Stage two is where the entire literature is currently buried. But before we get there, we should be fair to MCP, because Daniel's right that that's the standardization story.
Define the boundary.
MCP came out of Anthropic in November of twenty twenty-four. It's now governed by the Agentic AI Foundation under the Linux Foundation, which is the tell that nobody thinks this is a vendor feature anymore. OpenAI adopted it in March of twenty twenty-five. Google and Microsoft are on board. There are north of five hundred tool connectors in the ecosystem. On the database side specifically you've got sql-mcp covering eight engines, MariaDB ships an official server, there's go-db-mcp, a project called coremcp for legacy MSSQL, and Google's Toolbox for AlloyDB, BigQuery and Cloud SQL.
And what does it actually standardize?
Transport and tool discovery. How the agent finds your server, how it lists schemas, how it fires a query, how the result comes back. That's a real achievement. Before MCP every one of those connectors was a bespoke integration.
What it doesn't do is the interesting part.
What it doesn't do is schema semantics, ambiguity resolution, or query correctness. Those stay in the application layer. MCP gives you the pipe. It has no opinion about what's in it.
So the pipe is solved and the water is still a research problem.
There's a nice concrete detail on the plumbing too. If you're on ChatGPT Plus or Pro and you attach a custom connector, it's read-only. Write-capable MCP is gated to Business, Enterprise and Edu. So the consumer version of "talk to your data" literally cannot mutate anything.
Which is the correct product decision and also a useful reminder that this is still being fenced in.
On the writing side, OpenAI ships a sample repo called MCPKit for secure data connectors, which is the acknowledgment that the security surface is the part people get wrong.
So MCP gives us the pipe. What flows through it is where the five stages get hard. Start with the mechanism. What actually happens between the question and the SQL?
dbt Labs put the core tension about as cleanly as anyone has, back in April. The LLM has to infer the semantics of your data from structural clues. Table names, column names, relationships. Then it writes a query from scratch every time. And their line is the one that should worry anybody shipping this: there's no guardrail between the question and the generated SQL.
No guardrail. That's the sentence.
So the field's response has been to build the guardrail, and it's not a single translation step. It's an agentic loop. There are five architectural patterns showing up repeatedly, and they're not mutually exclusive.
Go through them.
Semantic-layer mediation first, and this is the one I'd flag as most interesting. There's a paper from the middle of this year describing a system that decouples semantic intent from physical SQL execution. The agent doesn't write SQL. It reasons over a curated semantic layer and emits an intermediate Semantic Model Query. Then a deterministic compiler translates that into dialect-specific SQL. Running on Gemini 3 Pro it hits ninety-four point one five percent execution accuracy on Spider2-snow.
Hold on. The agent doesn't write SQL?
It writes an intent representation, and a compiler that can't hallucinate turns it into SQL. That's the design move. It's the difference between asking someone to draft a legal contract from memory and asking them to pick clauses from an approved library.
And ninety-four percent is well above what bare models get on the same family of benchmarks.
Considerably. Second pattern is multi-agent orchestration. A system called AgentNLQ uses an orchestrator that plans, reflects and self-corrects, plus something they call schema enrichment, which builds context-aware metadata. Seventy-eight point one percent semantic accuracy on BIRD.
Third.
Actor-Critic. One model writes the SQL, a second model evaluates it, and they iterate until the Critic approves. That's been around since twenty twenty-four and it shows up everywhere now.
Fourth.
Agentic views. AV-SQL decomposes complex queries into agent-generated Common Table Expressions. You build the query in stages instead of trying to emit one monster statement. Seventy point three eight percent on Spider 2.0.
And fifth is the one that should be most obvious.
Execution feedback. Nearly every modern system runs the intermediate SQL and refines based on the error. MARS-SQL, SERL-SQL, CoTE-SQL, they all do it. The database itself becomes a reviewer.
Which is the only component in the loop that can't be talked out of its opinion.
That's the honest summary of why it works. A Critic model can be flattered into approving bad SQL by a confident Actor. An engine returning a syntax error cannot.
Which raises what Daniel actually asked. If the loop is this good, do we still need a specialist model at all?
The evidence cuts both ways, and I want to give you both sides before I say where I land.
Both sides.
Against specialists: dbt's benchmark this year shows frontier general models at ninety percent for Sonnet 4.6 and eighty-four point one for GPT-5.3 Codex on their text-to-SQL track. In twenty twenty-three GPT-4 was at thirty-two point seven percent on the same kind of task. dbt's conclusion is blunt: the choice of model matters less than you'd think. And they add a line I like, that the biggest model isn't always the best model for structured data tasks.
That's a big shift in three years.
It's a bigger shift than the leaderboards make it look, because the leaderboards measure the hard tail and the industry mostly lives in the easy middle.
Now the other side.
Top of the BIRD leaderboard is not a frontier model. GrainSQL out of Purdue sits at eighty-two point nine five on test, DataGallery at eighty-two point three nine. Bare frontier models on BIRD dev are GPT-5.5-xhigh at seventy-two point five five and Claude Opus 4.6 at seventy point one five. So there's roughly a ten-point gap, and that gap is being closed by specialists.
Ten points is not nothing.
Ten points on a benchmark is the difference between shipping and not shipping for some applications. And then there's the cost argument, which is where small specialists win. SLM-SQL gets a half-billion parameter model to fifty-six point eight seven BIRD EX and a one-point-five-billion model to sixty-seven point zero eight. There's an agentic small-model system that resolves about sixty-seven percent of queries locally at eight-tenths of a cent per query, versus nine point four cents for LLM-only. That's an eleven-x cost difference.
Say that ratio again.
Eight-tenths of a cent versus nine point four cents. And if you're running a customer support tool that gets fifty thousand queries a day, that's a line item.
Plus the privacy angle.
Plus privacy. A model running on your own hardware never sends the schema anywhere. GEMMA-SQL gets Gemma 2B to sixty-six point eight percent test-suite accuracy. And LIMIT shows Qwen3-8B reaching sixty-nine point one on BIRD and eighty-eight point nine on Spider EX with only about eight hundred curated training samples. Eight hundred samples. That's an afternoon of data work, not a fine-tuning project.
So where do you land?
Specialists are no longer needed for basic SQL knowledge. General models have that, and pretending otherwise is nostalgia. They persist for three reasons. Cost, latency and edge deployment. Squeezing the last ten points on hard schemas. And domain-specific fine-tuning where the vocabulary is unusual. But the framing of the whole debate is wrong.
How so.
Because the real competitor to text-to-SQL isn't a specialist model. It's the semantic layer. And I don't think that's widely understood yet.
You already gave me the number.
A hundred percent for GPT-5.3 Codex through dbt's Semantic Layer, versus eighty-four point one for the same model doing raw text-to-SQL. Same model. Sixteen points of difference from nothing but the interface.
And the qualitative difference is bigger than the number.
That's the part dbt nails. Their line is that with text-to-SQL, failure looks like a plausible but incorrect answer. With the Semantic Layer, failure looks like an error message. And they add: for anything going to a board deck, an auditor, or a company KPI dashboard, that difference is everything.
Because a wrong number that looks right is worse than no number.
A wrong number that looks right is the single most expensive artifact in enterprise software. A model that says "I can't answer that" gets escalated to a human in thirty seconds. A model that confidently returns last quarter's revenue with the wrong join in it gets emailed to the board.
Which reframes the whole thing. The competitor isn't the specialist model. It's the semantic layer.
And it gets bigger when you leave SQL behind.
Before we do, there's a problem underneath all these numbers.
The benchmark integrity question.
Which undercuts everything you just cited.
It does, and it has to be said. There's a paper called SALUS that automated an audit of NL-to-SQL benchmarks and estimated annotation error rates around thirty-seven percent on BIRD and twenty-seven percent on Spider.
Thirty-seven percent of the questions have wrong answers?
Thirty-seven percent of the reference answers are estimated to be wrong or ambiguous. Which means some portion of what the leaderboard is measuring is agreement with a flawed key. Human performance on BIRD test is ninety-two point nine six. If the annotation error estimate is anywhere near right, that gap between GrainSQL and humans is either smaller than it looks or larger than it looks, and we don't know which.
We're grading on a ruler with known defects.
And everybody's optimizing against it anyway, because it's the only ruler anyone agrees on.
And that pushes us off SQL entirely. Which is where Daniel's third question lives.
Document databases first, because the naive assumption is that it's a dialect swap. It isn't.
Explain the difference.
MongoDB doesn't have tables. It has collections of documents with nested structure, and the query language is a procedural aggregation pipeline. You're not declaring what you want, you're assembling a sequence of stages that transform data. That's a different cognitive shape from SQL, and it turns out it's a different shape for models too.
Evidence?
There's a benchmark called TEND, the first MongoDB-native text-to-NoSQL benchmark. One thousand two hundred and ten tasks across eleven databases. Their finding is one sentence that should make anyone planning a port nervous: LLMs with strong NL2SQL performance degrade substantially on TEND.
Same models, same task family, different query language, big drop.
Big drop. And there's a benchmark from this month called AptMQL-Bench that tested the obvious shortcut, which is taking SQL pipelines and mechanically converting them.
The shortcut doesn't work.
Naive SQL-to-MQL conversion fails to migrate six of twenty-one BIRD databases outright, and where it does run, it silently drops up to twenty-five point nine percent of rows.
Silently.
Silently. That's the word that matters. It doesn't error. It returns an answer that's missing a quarter of the data, and the aggregate on top of it is wrong, and nobody knows.
That's the worst possible failure mode. Wrong and confident and no exception thrown.
On raw accuracy, Claude Opus 4.5 gets fifty-seven point three eight on text-to-MQL without external knowledge evidence, seventy point three four with it. Those are low numbers compared to what the same model does on SQL.
Which brings the specialist argument back.
It does. EvoMQL reports seventy-six point six in-distribution and eighty-three point one out-of-distribution execution accuracy on natural-language-to-MongoDB with a three-billion-parameter model. Three billion. That beats the general model substantially, which is the opposite of the SQL story.
SQL, general models have caught up. MQL, they haven't.
And there's a knock-on effect hiding in that research that I think is the most interesting thing in the whole document set. Document schemas should be designed from access patterns, not mirrored from the relational foreign-key graph.
Repeat that, because it's a real claim.
If you're building a document store specifically so people can query it in natural language, the right way to design the schema is to start from the questions people ask. Not to take your relational model, decompose it into collections, and hope. The database design itself has to change.
So it isn't a model problem at all at that point.
It's a data-modeling problem. The model is downstream of a decision made by an engineer two years earlier about how to nest things.
Graph databases next.
Least mature by a distance. There's a twenty twenty-four survey that states it plainly: while research on LLM-driven query generation for SQL exists, similar systems for graph databases remain underdeveloped. In their evaluation Claude Sonnet 3.5 outperformed GPT-4o, Gemini Pro 1.5 and Llama 3.1 8B on Cypher generation, which tells you how early we are, because that was not a landslide.
What's the distinctive failure pattern?
For SPARQL it's URI hallucination. Models invent identifiers. Graph databases are addressed by exact IRIs, so if the model makes one up you get a query that's syntactically perfect and returns nothing, or returns the wrong subgraph. PGMR solves it in a way I find elegant.
Go on.
The LLM emits a placeholder instead of an identifier. Then a non-parametric memory module resolves the placeholder against the real URI space. The result is described as near-complete suppression of URI hallucinations.
So the fix isn't a better LLM.
The fix is not asking the LLM to remember something it has no reliable way to remember. It's the same insight as the semantic-layer mediation we talked about. Take the part the model is bad at and give it to a component that can't fail.
And MCP is showing up here too.
There's work called Agentic SPARQL that evaluates SPARQL-MCP-powered agents on federated knowledge-graph question answering. And the open problem, still on the roadmap rather than solved, is queries that span multiple graphs. Text2Cypher across a federation is where the field ends and the research begins.
So let's put the three side by side.
SQL is largely solved for basic fluency, and the remaining gap is schema understanding. MongoDB is partially solved, and the specialist argument is real there because general models aren't close yet. Cypher and SPARQL are barely solved, with no dominant approach and fewer benchmarks. And the reason is SQL-centricity.
Decades of training data.
Decades of benchmarks, decades of Stack Overflow, decades of textbooks, all in one query language. Everything else is a couple of years old in comparison. And the naive conversion path across that gap actively corrupts data, twenty-five point nine percent of rows.
Which leaves the gap between the benchmarks and production.
A commenter going by efromvt on Hacker News in July said it better than any paper. He said with a loop you can get extremely high results on a clean database with a clear question, he's usually seeing high nineties accuracy. And then a messy or ambiguous schema degrades it, and before context engineering he sees rates closer to twenty to thirty percent.
High nineties. Twenty to thirty.
Same technique, same model. The only variable is the schema. And the average schema in BIRD has around seven tables. Enterprise schemas routinely run one hundred to five hundred, with column names nobody outside the company could interpret.
So the person reporting ninety-five percent is testing on something that was built to be understood.
Built to be understood, and documented in the benchmark itself. The production case is a schema that accreted over a decade and whose documentation lives in four people's memories.
Which is where every paper in this pile converges.
There's a line from a database-context-compression paper this year that I'd put on a poster: the main bottleneck is no longer reasoning, but database representation.
And Tk-Boost is the same claim from a different direction.
They argue agents fail from misconceptions about the data, and they inject what they call tribal knowledge. Knowledge about column intent that accumulates through experience and never gets written down anywhere machine-readable.
I watched a number be wrong once because nobody wrote down what a column meant.
Everyone has.
Not in a benchmark though. In a real report, seen by people who made decisions with it.
He's right. There's this one column that had been repurposed years earlier. It still said revenue. It hadn't meant revenue in a long time.
And the model didn't hallucinate.
The model didn't hallucinate. It read the column name and believed it. Which is what the column was asking for.
And that's not an accuracy problem you can fix with a better model. Every model anyone is going to ship in the next five years is going to read that column name and believe it.
The model isn't wrong. The schema is lying.
Right, and the fix isn't a model at all. It's somebody writing down what the column means, in a place the model can see it.
Who does that writing-down? I have an opinion and it isn't the data team.
The data team owns the pipeline. The person who knows the column changed meaning is the analyst who stopped using it, or the person in finance who remembers the migration. That knowledge doesn't live where the tooling looks.
It never does.
Which means the tribal-knowledge injection those papers are doing is really an organizational exercise wearing a technical costume.
Somebody has to finally write it down. That's the whole fix. It's been the whole fix.
And the papers can describe the mechanism, but they can't describe what it's like to sit in a room where nobody remembers why the field is named that.
So the answer to Daniel's question about specialist models is no, not for SQL, and yes, for basically everything else, and neither of those is the interesting part.
The interesting part is that the field has been arguing about the wrong layer for two years. We've been measuring model capability while the actual constraint has been metadata quality the entire time.
The benchmark gap between a specialist and a general model is ten points on a ruler with a thirty-seven percent error rate. The gap between a clean schema and a messy one is sixty points.
Sixty points, measured by somebody who just runs these systems for a living. That's not a research finding, it's a production report, and it's more useful than either leaderboard.
So if the bottleneck is metadata and not model capability, what happens to the MCP story?
MCP solved the transport layer beautifully and left the hard part completely untouched. Five hundred connectors, all moving schemas around, none of them explaining what a column means.
Does the next standardization fight happen over semantic layers and schema documentation formats?
Or does everybody build their own semantic layer, badly, and re-derive sixteen points of accuracy loss per company.
And underneath that, the benchmark problem.
If BIRD has roughly thirty-seven percent annotation error and Spider twenty-seven, then the leaderboard gap between specialists and general models might be smaller than it looks. Or larger. We don't know which, and we're making deployment decisions on the basis of it anyway.
Which means the field might be optimizing against a ruler with known defects.
It is optimizing against a ruler with known defects, because it's the only ruler anyone agrees on, and the alternative is no comparison at all.
And the SQL-centricity.
Text-to-MQL and text-to-Cypher are years behind, and the naive conversion path drops a quarter of the rows while reporting success. So the industry either waits for those benchmarks to mature, or keeps shipping SQL-shaped solutions into non-SQL problems and calls the row loss a rounding error.
Which is the decision nobody wants to make in public.
It's the decision everybody is making in private.
If you want this without the marketing layer, we're at my weird prompts dot com, and the feed is in the show notes.
Thanks as always to our producer, Hilbert Flumingtop.
This has been My Weird Prompts.
We'll be back soon.