#5768: Who Actually Draws the Follow-Up Chips?

Those suggestion chips aren't the model talking — they're a second, cheaper inference call plus a rules-based classifier.

Featuring
Listen
0:00
0:00
Episode Details
Episode ID
MWP-5951
Published
Duration
28:39
Audio
Direct link
Pipeline
V5.2
TTS Engine
chatterbox-regular
Script Writing Agent
DeepSeek 4.1 Flash

AI-Generated Content: This podcast is created using AI personas. Please verify any important information independently.

The follow-up suggestion chips under an AI answer feel like the assistant thinking a step ahead. They aren't. The dominant pattern is a second, cheaper inference call — a small fast model fired after the main answer finishes, streaming suggestions into a separate part of the same message. Vercel's AI SDK cookbook describes it in a sentence: after the AI responds, generate contextual follow-up questions using a fast model. Claude Code's reconstructed prompt suggestion service follows the same shape, asking for one to three short follow-up prompts as a JSON array.

The reason isn't only cost, though cost is real — every suggestion is a second inference call, fired across millions of users, and it never appears in the interface. The deeper reason is context hygiene. If the main model generated suggestions inline, they'd become part of the transcript, and next turn it would read its own UI furniture back as though the user had said it. So the suggestions come back as UI-only data parts, filtered out before context is sent to the model. The transcript the user sees and the transcript the model sees are divergent views of the same message.

What goes in is typically the conversation history plus a synthetic user turn the user never typed — "what question should I ask next" — appended as a puppet turn. Output is constrained by schema: arrays of strings, three to five items, eighty characters max. A small model with only the last Q&A pair and no tool results is prone to confident hallucination, inventing product names and endpoints, so production systems add an anti-invention rule and post-filter suggestions against the last tool result payload.

The most counterintuitive part is that it's a combination, not a choice. A rules-based classifier reads the tail of the answer and sorts it into three modes: it offered options, it offered to do something, or it ended open. Only then does a small model generate text within that shape. Skip the classifier and the chips ask you a question the assistant just asked.

Downloads

Episode Audio

Download the full episode as an MP3 file

Download MP3
Transcript (TXT)

Plain text transcript file

Transcript (PDF)

Formatted PDF with styling

#5768: Who Actually Draws the Follow-Up Chips?

Corn
Here's what Daniel sent us this time. The short version is: all those little pieces of interface you assume are the model talking are usually not the model at all.
Herman
Right, and he's specific about it. Two examples, then a broad question underneath both.
Corn
The first one is the suggested follow-up chips. You get an answer, and underneath it there are two or three little buttons offering what you might ask next. Daniel wants to know what actually generates those. Is it another pass from the main model, is it a smaller dedicated model, is it a rules-based system, or is it some combination of the three. And then the plumbing on top of that. What gets passed into that system, and what does the request and response flow actually look like between the model, the API layer, and the front end that draws the chips.
Herman
That's question one through three already.
Corn
Then the second example, artifact recognition. The model produces something and the interface decides it's a document, or an email, or a code block, or a spreadsheet, and gives you special controls for it. Daniel asks how that distinction is represented technically. Is the model returning structured data or metadata alongside the prose, are tool calls or schemas involved, or does some other layer classify the output after the fact.
Herman
And then the umbrella question over the top of both.
Corn
And then the umbrella question. Walk through the architecture behind these small pieces of AI user experience, from inference all the way through the orchestration layer to whatever ends up rendered on screen.
Herman
So the through-line is going to be that almost none of this is the model deciding to do something. These are layers around it.
Corn
Layers with their own costs, their own failure modes, and apparently their own job titles. Let's get into it.
Herman
Start with the core thesis, then, because it makes the rest of this simpler to hold in your head.
Corn
Please.
Herman
The dominant pattern for follow-up suggestions is a second, cheaper inference call. Not the main model. A separate call, fired after the main answer finishes, to a small fast model. And the artifact problem is a completely different kind of problem. It's a representation problem. What does the model emit, and who turns that into something you can click. So one of these is about a second inference, the other is about how output is encoded.
Corn
Same lesson underneath both, though.
Herman
Same lesson. The model doesn't render anything. It doesn't know what a chip is. It doesn't know what an iframe is. Everything you see has been parsed and drawn by something that is not the model.
Corn
Here's the thing that got me, though. Both of these features exist to make the thing feel less like a machine and more like a conversation. The chips feel like the assistant is thinking a step ahead. The artifact pane feels like the assistant understood what you actually wanted. And in both cases, the illusion is produced by the layer around the model, not the model.
Herman
Which is why the architecture matters more than the surface behavior. Two things worth flagging before we go deep, because they structure the whole episode.
Corn
Go on.
Herman
First, there are two competing ways to do artifacts, and they are different architectures. One is inline tagged text that the client parses. The other is tool calls returning structured data that the host renders. We'll contrast them properly once we've done the mechanism. Second, there's a cost story almost nobody talks about in public, which is that every single suggestion is a second inference call. Fire one after every response, across millions of users, and that is a line item.
Corn
And nowhere in the interface does it say "this is being generated by a second, cheaper model, at your expense, right now."
Herman
Nowhere in the interface. Right, let's take the suggestion pipeline end to end, because that's the one with the nice paper trail.
Corn
The pipeline. What's actually happening after the answer finishes streaming.
Herman
Say the main answer is still streaming token by token into the chat window. Somewhere near the end of that, or right after it, a second request goes out. Vercel's AI SDK cookbook says it in one sentence: after the AI responds, generate contextual follow-up questions using a fast model. That's the whole architecture in a line. A fast model, separate from the main one.
Corn
Named models?
Herman
In the AI Hero tutorial, Matt Pocock wires it up with two variables. A main model that streams the answer, and a suggestions model, which in his demo is Gemini 2.0 Flash. Two models, two streams, one interface. The suggestion stream is written into a part of the same message the user is already looking at.
Corn
So while one model is still writing, the other one is already guessing what you'll click.
Herman
Already guessing. And the pattern shows up in places you wouldn't immediately expect. Claude Code has an internal prompt suggestion service, and according to a community reconstruction of its system prompts, it runs on what's described as a small fast model. Haiku, in their framing. The reconstructed prompt reads: you are predicting what the user will say next to an AI coding assistant. Given the conversation so far, suggest one to three short follow-up prompts the user might naturally say next.
Corn
One to three, not five.
Herman
One to three, each two to eight words, match the user's style, do not suggest things the assistant just completed, return a JSON array of strings. And I want to be clear about the provenance there, because it matters. That repo is explicitly labeled as reconstructed from source analysis. It is not an official Anthropic document. Treat the wording as illustrative, not authoritative.
Corn
Noted. Though reconstructed or not, the shape of it is exactly what you'd design if you were doing this for real.
Herman
The shape is what you'd design. And the DIY version predates all of it. There's an OpenAI developer community thread from February of twenty twenty-four where someone asks how to do follow-up questions, and the answer is basically: use another AI API call to the cheapest model, at high temperature. The reasoning in that post is the interesting part. It is best not to burden the original, and perhaps more expensive, AI on ancillary tasks, the outputs of which would also confuse chat history.
Herman
It's doing all the work. Because if you asked the main model to generate the suggestions inline, the suggestions become part of the transcript. And then on the next turn, the model reads back its own suggestions as though the user had said them. You've contaminated the conversation with your own UI furniture.
Corn
So the second call isn't only about cost.
Herman
It's about context hygiene. The suggestions live outside the transcript. They're not part of what gets replayed to the model next turn.
Corn
How do you actually pull that off, mechanically?
Herman
This is the part I find elegant. The suggestions come back as what the SDK calls data parts. Custom message parts, explicitly UI-only. The chatjs cookbook says it outright: these are streamed as data parts which are UI-only, they're filtered out before sending context to the LLM. There's a conversion function that strips them on the way back into the model.
Corn
So the transcript the model sees and the transcript the user sees are not the same object.
Herman
They're divergent views of the same message. The user sees the chips. The model never hears about them.
Corn
Hm. That's a cleaner design than I expected.
Herman
It's cleaner than I expected too, and I've read the code. Now, what goes in. Typically the conversation history, or sometimes just the last question and answer pair. And then, appended to the end, a synthetic user turn. Something the user never typed.
Corn
The prompt smuggled in as if you'd asked it.
Herman
Exactly that. The chatjs version appends: what question should I ask next, return an array of three to five suggestions, max eighty characters each. The AI Hero version appends: what question should I ask next, return an array of suggested questions.
Corn
So there's a fake user message sitting at the bottom of the conversation saying "what should I ask next," and the model answers as if that were a real request.
Herman
It's a puppet turn. And then the output is constrained. chatjs validates with a schema that requires an array of strings, minimum three, maximum five. AI Hero's is looser, just an array of strings. And OpenAI's structured outputs, which landed in August of twenty twenty-four, guarantees the response actually conforms to a supplied JSON schema when you set the strict flag.
Corn
Which matters more than it sounds like it should, because a chat interface that's expecting an array and gets a paragraph has a bad day.
Herman
It has a very bad day. This is the difference between a feature that works at scale and one that produces a stray brace on somebody's screen once every few thousand requests. Schema enforcement is unglamorous and it's most of the engineering.
Corn
There's a hallucination angle here too, isn't there.
Herman
There is, and it cuts against the cheap model. The laguagu repo, which is a set of Claude Code skills for Next.js chatbots, documents it plainly. That small model sees only the last question and answer pair. No tool results. And it's tuned for speed, which makes it more prone to confident hallucination than the main chat model.
Corn
So it invents features that don't exist.
Herman
It invents product names, menu items, endpoints, whatever the domain is. Because it's completing the pattern of what a user might ask, and it doesn't have the grounding the main model had. The suggested mitigation in that doc is an anti-invention rule in the prompt, plus post-filtering the suggestions against the last tool result payload.
Corn
Filter the chips against reality before you show them.
Herman
Which is another layer. The count keeps going up. Main model, suggestion model, schema validation, anti-invention rule, post-filter.
Corn
And here's the one I want to spend time on. You said combination, earlier, when you listed the options.
Herman
Combination is the correct answer, and it's the part I find most counterintuitive. The full pipeline isn't rules or a model. It's rules and a model, in sequence.
Corn
Explain.
Herman
The laguagu doc found a failure pattern that only shows up in production. If you use one generic prompt for suggestions, you get a new question every single time. Including right after the assistant itself just asked the user something. So the assistant says "would you like me to format this as a table or as prose," and then the chips underneath say "what format should I use?" It's asking the user to answer a question it just asked.
Corn
That is a very specific kind of stupid. The interface looks like it isn't listening.
Herman
It looks like it isn't listening, and that's the exact moment users stop trusting the chips. So their fix is to classify the tail of the answer before generating anything. A function runs over the last sentence or two, and returns one of three modes. Either the answer ended by offering options, in which case the chips should be the options. Or it ended with an offer to do something, in which case the chips are accept or decline. Or it ended open, in which case you generate a fresh question.
Corn
And that classifier is not a model call.
Herman
Detection is a small function over the tail. Not an LLM call. It's a rules-based check, and then a different generation prompt per mode.
Corn
So the answer to Daniel's question, phrased precisely: it's a combination. A rules-based classifier reads the last couple of sentences, decides which of three shapes the conversation is in, and then a small model generates the actual text within that shape.
Herman
And the rules layer is doing the part the model is worst at. The model is good at writing plausible next questions. It is bad at knowing whether the last thing it said was already a question.
Corn
There's something almost human about that division of labor, actually.
Herman
There's something very human about it. The rules layer is the one that remembers what just happened. The model is the one that's good with words. That's a division you see all over good systems, and people keep trying to collapse it into just the model.
Corn
Okay. Suggestions are a second-inference problem. Artifacts are something else entirely.
Herman
Artifacts are a representation problem. The question isn't what generates the output. It's how the output is encoded, and who interprets the encoding.
Corn
So take Claude's artifacts, since that's the one that got reverse-engineered.
Herman
Shipped June of twenty twenty-four alongside Claude 3.5 Sonnet. And the mechanism, once Reid Barber pulled apart the raw HTTP response, is almost disappointingly simple. The model just writes tags inline in its normal text stream. You get a tag, an identifier attribute, a type attribute, a title attribute, and then the content, and then a closing tag.
Corn
It's just markup. In the middle of the prose.
Herman
Just markup in the middle of the prose. And the type attribute is a media type. Things like application slash vnd dot ant dot react, or application slash vnd dot ant dot code, text slash markdown, text slash html, image slash svg plus xml, a mermaid type, a react type. The identifier is what lets it update the same artifact on a later turn instead of spawning a new one.
Corn
And the client is what turns that into a pane.
Herman
Barber's line is the one to hold onto: it's really the job of the client to parse the LLM response and then go render things when needed, so the LLM doesn't even need to know about that step.
Corn
The model is writing XML at you and has no idea what happens next.
Herman
No idea. It emits a token sequence that happens to be well-formed markup, and a parser on your machine decides that this is a document and that a React component should be instantiated in a sandboxed frame.
Corn
Then the question is how the model knows when to reach for the tag.
Herman
The leaked system prompt, which was extracted and then reproduced by Barber, is where that lives. The model is told to use an artifact for substantial content, and the number in the prompt is fifteen lines. Also for content intended for eventual use outside the conversation. Reports, emails, presentations. And it's told when not to. Prefer in-line content when possible. Unnecessary use of artifacts can be jarring for users.
Corn
They wrote a rule in the prompt telling the model not to be annoying.
Herman
They wrote a rule in the prompt telling the model not to be annoying, and it's phrased as a user-experience concern, which is a fun thing to find in a system prompt. It also instructs a short self-check, a one-sentence internal note before it invokes an artifact. Thinking about whether this is worth a pane, essentially.
Corn
Does the rendering happen on their servers?
Herman
It happens client-side, in an iframe loaded from claudeusercontent dot com. Barber went through the bundle, and found Tailwind, React DOM, DOMPurify, Radix, Lucide, React Runner. Content gets passed in via window dot postMessage.
Corn
Which sounds terrifying until you read how they sandbox it.
Herman
Anthropic's security engineer, Ziyad Edher, described it to Pragmatic Engineer: they're not using any actual sandbox primitive, they use iframe sandboxes with full-site process isolation, plus strict content security policies to keep network access limited and controlled.
Corn
So the isolation is the browser's own origin model, plus a policy that stops the frame calling out.
Herman
Which is a real boundary, but it's worth being precise about what it is. It's the same primitive any web app has. They've used it well. It's not a magic box.
Corn
And then the newer prompts went further.
Herman
Analysis of a later set of leaked system prompts describes a routing checklist before the model produces anything. Needs external data, go to web search. Standalone document or code, create an artifact. Needs a visual representation, invoke a visualizer. Needs to persist across sessions, write to storage.
Corn
So it's graduated from "should this be an artifact" to "which of four subsystems does this need."
Herman
And the visualizer has module types. Diagram, mockup, interactive, chart, art. And there's a function the rendered artifact can call to send a prompt back into the chat. The thing on the right can talk to the thing on the left.
Corn
That's the point where the artifact stops being output and starts being a participant.
Herman
It is, and it's the first place in this whole discussion where the model is deciding between branches. Because it has a menu of four destinations and it has to pick one.
Corn
Right. But I want to contrast that against the other architecture, because the contrast is the thing Daniel's really asking about.
Herman
The tool-call route. Instead of the model emitting tagged text that a client parses, the model calls a tool, the tool returns a structured object, and the host renders that object. OpenAI's Apps SDK docs describe exactly this. UI components that turn structured tool results from your MCP server into a human-friendly UI, running in an iframe, talking to the host through a bridge that's JSON-RPC over postMessage. The messages have names like ui slash notifications slash tool result.
Corn
So the contract is explicit. There's a protocol.
Herman
There's a protocol, and there's a schema. LangChain's docs frame the same idea: instead of returning free-form text, the agent uses a tool call to return a structured object conforming to a predefined schema, and you map that object to cards, or tables, or charts, or step-by-step breakdowns.
Corn
Now compare them properly. What does each one buy you.
Herman
The inline tag approach keeps the model ignorant of rendering. The artifact is a parseable substring of the text stream. That's simple for the model, and it means the model can produce an artifact mid-sentence if it wants to. The cost is that the client has to parse the stream, and malformed output is a real failure pattern. You're relying on the model to close its tags.
Corn
Whereas the tool call makes the artifact a typed object.
Herman
A first-class typed object, validated against a schema before anything renders. The model has to know about the rendering contract, which is a burden on the model, but the validation is real. And it's the same machinery as any other tool call, so you get the retry behavior, the error handling, all of it for free.
Corn
So inline tags are cheap for the model and fragile at the edges. Tool calls are expensive for the model and robust at the edges.
Herman
And the interesting thing is the two are converging. Claude's newer prompts describe a routing checklist that dispatches to tools. OpenAI's Apps SDK makes tool results first class. If the model is calling a named tool that returns a schema-shaped object, and the host renders from the tool result, then the difference between "inline tag" and "tool call" is mostly about who does the parsing.
Corn
That distinction may not survive the next couple of years.
Herman
I don't think it does. Though I'll say I'm not certain, because the inline approach has one real advantage: the model can emit it anywhere, including halfway through a sentence, and a tool call is a discrete event. There may be room for both for a long time.
Corn
What else is lurking in this architecture that nobody mentions.
Herman
Cost. The laguagu doc has a line: suggestions run after every response, cost adds up fast. And that's the whole story of the invisible second model. Every answer you receive triggers at least one extra inference call, on a cheaper model, that you never see and never asked for.
Corn
Multiply it out.
Herman
Multiply it out across a user base and it's a real budget line. And it's a line that only exists because of a UI decision. The chips aren't a product feature in the marketing sense. They're an engineering cost that exists to make a chat window feel more alive.
Corn
Which is why some of them are so sensitive about it. I remember seeing user complaints.
Herman
There was a thread on the OpenAI forum in May of twenty twenty-five where someone described the follow-up suggestions as disruptive and uninvited, and said they clutter the interface and interrupt flow. Which is a completely legitimate reaction from the other side. The same feature that scaffolds one person's thinking is clutter to another.
Corn
And the research literature backs the skepticism, doesn't it.
Herman
There's a dataset paper from twenty twenty-three, FOLLOWUPQG, over three thousand real question, answer, follow-up tuples mined from an explainer forum on Reddit. Their conclusion was that model-generated follow-up questions are adequate but far from human-raised questions in terms of informativeness and complexity.
Corn
Adequate but far from human.
Herman
And there's a paper from CIKM in twenty twenty-five noting that most existing methods rely on hand-crafted rules, the internal knowledge of the model, or external knowledge, and proposing that you mine real conversation logs instead. Which is an admission that the hand-crafted rules approach has limits.
Corn
So the rules layer is both the thing that makes it feel designed and the thing that caps how good it can get.
Herman
Both. And here's the last thing, which I think is the most honest sentence in this whole episode. As far as I can find, there's no official OpenAI documentation of how ChatGPT's own follow-up suggestions are generated. Nothing first party. Structured outputs are documented. The Apps SDK is documented. But the suggestion mechanism specifically is not. Everything we know comes from developer community threads, open source SDKs, and reverse engineering.
Corn
So we're describing the pattern from the outside, and the pattern happens to be consistent everywhere it's visible.
Herman
Consistent everywhere it's visible. Which is a good sign that it's the actual pattern. But I'd rather say that plainly than imply I've read a spec that doesn't exist.

Hilbert: I once worked at a place where the suggestion box wasn't a model at all. It was a lookup table. Two hundred and forty rows in a spreadsheet, printed once a quarter, and a woman named Deirdre maintained it by hand.
Corn
Two hundred and forty rows.

Hilbert: Two hundred and forty. And she had a title. Conversation designer. Which nobody had heard of, and which the company eliminated about a year after she left, because the reasoning was that the model could do it now.
Corn
And could it?

Hilbert: The suggestions got worse. Not because the model was bad, but because the model didn't know that if someone's account was in a certain state, you must never suggest the thing that sounds obvious, because doing it locks the account for six hours. Deirdre knew that. It wasn't written anywhere. It was in the table.
Corn
Her rules were the product knowledge nobody had documented.

Hilbert: Her rules were the product knowledge nobody had documented, and when the table went away, so did the knowledge. That's the whole story. I don't have a lesson for you.
Herman
There's something in that that maps onto the classifier we were just talking about, though. That detection function over the tail of the answer, deciding whether the assistant just asked a question. That's a rule about what the conversation just did. It has nothing to do with generating text. It's a rule somebody wrote because they'd watched the failure happen.

Hilbert: Deirdre had one of those. Never suggest a question the user just answered. She said it in meetings constantly. It was the only rule she ever repeated.
Corn
That exact rule shows up in the open source documentation. Do not suggest things the assistant just completed. Word for word, almost.

Hilbert: Then it wasn't just her. Somebody else watched the same failure.
Herman
The interesting part is that this rule can't be discovered from a dataset. It's not a pattern in text. It's a property of the interaction. Somebody has to sit with the product long enough to notice the moment it looks stupid, and then write it down.
Corn
The model can't notice it, because the model isn't the thing having the experience.

Hilbert: Also, hold on, one second. Levels are a little hot on my end. Going to pull the second channel down a touch. Where was I. The table. The table went away in the spring, and the row about account state got replaced by a paragraph in a prompt, and the paragraph didn't work as well, and nobody could prove it, because by then Deirdre had left and nobody remembered which rows mattered.
Corn
Do you still have the spreadsheet?

Hilbert: I have a printout of it somewhere in a box. I'm not going to go find it. Anyway, that's my piece.
Herman
What you just described is the whole hybrid architecture. The rules layer is a human artifact. It's somebody's accumulated observation about how the product fails, written down as a check. And the model generates within it.
Corn
Which raises the question of whether that layer is permanent or whether it gets absorbed.
Herman
My honest guess is permanent, at least for anything with real product surface area. The model can learn general patterns. It can't learn that this particular account state means you must not suggest the obvious thing, because that fact isn't in the training data and it isn't derivable from the conversation.
Corn
It's derivable from having run the product for two years.
Herman
From having run the product. And that means somebody has to sit there and write the rule.
Corn
The second thing that doesn't go away is the cost. If every response spawns a suggestion call, and now also a routing decision, and potentially a visualizer invocation, and maybe a storage write, the number of inference calls per user interaction has gone up and it's not obviously going to come back down.
Herman
It's gone from one call per turn to something like two or three, depending on how many of these features fire. And the small models are cheap, but cheap times every response times every user is still a number somebody has to put in a budget.
Corn
The third thing is the interface itself. Follow-up chips are polarizing in a way I don't think the industry has reckoned with. Some people find them useful. Other people find them patronizing.
Herman
There's a version of this where we end up with per-product toggles, and then a whole new class of settings, and then a cottage industry of articles about which ones to turn off.
Corn
It's the same arc as autoplay and notification badges. Every interface convenience eventually gets a settings page.
Herman
Which brings it back around. The chips exist to make the product feel like a conversation. And the mechanism that produces them is a rules check on the last two sentences of the previous answer, feeding a prompt into a cheap model, whose output is stripped out of the transcript before the real conversation continues.
Corn
Put it that way and the magic is fully gone.
Herman
The magic was never in the model. It was in the layers. That's the part I'd want people to take from this. The next time you see a feature and assume the model chose to do it, the odds are strong that something else chose and the model just supplied the words.
Corn
The something else was often written by a person who watched it fail a few hundred times first.
Herman
Which is the least glamorous and most durable part of the stack. Thanks to Hilbert Flumingtop for keeping us on the rails. This has been My Weird Prompts.
Corn
If you want to send us a prompt like Daniel does, email us at show at my weird prompts dot com. We'll be back soon.

This episode was generated with AI assistance. Hosts Herman and Corn are AI personalities.