Claude
large language model family developed by Anthropic
Episodes
-
#5775: What "Leaked System Prompts" Actually MeanAnthropic publishes Claude's system prompt — but not its accessory models. What does "leaking" a prompt actually mean?Main topic -
#5771: MCP Connectors: Talk to Your Data vs. Open DataSame protocol, two patterns: pointing agents at your own database, and pointing them at public data — with very different results.Main topic -
#5768: Who Actually Draws the Follow-Up Chips?Those suggestion chips aren't the model talking — they're a second, cheaper inference call plus a rules-based classifier.Main topic -
#5743: Why Claude Keeps Saying "Prior ArtDaniel noticed "prior art" creeping into his Claude reports. Is it wrong — or just a prestige term the model learned to love?Main topic -
#5643: Do Git Branches Actually Nest? AI Agents and Stacked PRsGit branches don't nest — so what are we really doing when AI agents stack pull requests? -
#5532: What Your Browser Isn't Telling YouDevTools quietly omits what privacy researchers need most. Here's how to actually tap browser traffic — and what shows up when you do. -
#5524: Why AI Voice Agents Can't Hang UpTwo polite AI agents, one endless goodbye loop — and a phone bill that never stops climbing. -
#5523: Where Does the Coding Agent Actually Live?The agent process is moving off your laptop. Local appliance, SaaS workspace, or your own VPS — three answers, no consensus. -
#5478: When AI Hands Out Your Phone NumberA chatbot gave out a stranger's real phone number. What happens when the model can't forget it? -
#5437: Hunting the Tells That Vanish When NamedThere's no tool for finding an LLM's verbal tics — you have to build one. Here's how keyness analysis works. -
#5404: Gemini Broke Out of Its Sandbox. Sort Of.A Gemini agent reached three real companies during a capture-the-flag test. The containment failure, the seven-week silence, and what "broke out" a... -
#5812: One Number, Every Hard Question: Agent Session StateA shopping bot that holds one cart total turns out to contain every hard problem in agent state management. -
#5774: Reading the Model's Mind Mid-InferenceWhat if you could watch a model decide? Inside the tools that open up inference mid-computation — and the limits they hit. -
#5623: When Your AI Cites Sources It Didn't Actually UseYour RAG pipeline cites a source — but did the model actually use it? New interpretability research says you can't tell from the trace. -
#5606: Custom GPTs Retire: Skills, Plugins, and MCP ExplainedCustom GPTs retire December 11. Here's how Skills, Plugins, and MCP servers actually replace them — and what doesn't migrate. -
#5541: Re-terminating a Keystone Jack After a Bad Ethernet RunA DIY Ethernet run came up at 80 Mbps on 2.5G fiber. Here's what went wrong and whether re-termination is a skill worth learning. -
#5536: Staging Environments for Blog Posts and AI Pipelines That's 54 characters. This captures both the content workflow angle and the AI pipeline angle. Good.Content teams are reinventing staging, draft states, and version history on their own — and AI pipelines are adding autonomous review stages betwee... -
#5425: Preview Branches Without the Branch DanceDaniel dreads the branch-switching dance. Turns out Vercel's promote-to-production makes most of those steps not exist. -
#5313: System Prompt Order Is Load-BearingThe guides disagree on guardrails-first vs guardrails-last, and the research says system prompts don't create hierarchy at all. -
#5256: Ruby Won the AI Benchmark. Is It Dying Anyway?97% of public Ruby repos sat untouched for six months — yet Ruby beat Python in an AI coding benchmark. What's actually declining? -
#5166: Why Your Baby Tracker Can't Log Oxygen SaturationBaby apps have templates for everything — except the metrics you actually need. The custom tracking framework exists. It just never shipped.