#rag
91 episodes · Page 3 of 4
#1804: The Fork in the Road: Why AI Agents Check Old Receipts First
Stop your AI agent from overthinking. Learn why it checks old memories instead of booking flights—and how to fix the "eagerness" problem.
#1794: RAG Is Cheaper Than You Think (Until It’s Not)
From a $1 embedding bill to a $10k/month vector database bill, here’s the real math behind RAG in 2026.
#1792: Google's Native Multimodal Embedding Kills the Fusion Layer
Google’s new embedding model maps text, images, audio, and video into a single vector space—cutting latency by 70%.
#1784: Context1: The Retrieval Coprocessor
Chroma's new 20B model acts as a specialized "scout" for your LLM, replacing slow, static RAG with multi-step, agentic search.
#1778: Audio Is the New "Read Later" Graveyard
Why listening to AI conversations beats reading dense PDFs, and how serverless GPUs make it cheap.
#1765: The Agentic Internet: A Clean Web for Machines
We explore the tools building a parallel, machine-readable web—from SearXNG to Tavily.
#1764: Your Repo as a Knowledge Base
How to give AI agents instant memory of your entire project—without cloud costs or complex infrastructure.
#1754: From Ollama to Agentic CLIs: The Rise of the AI Harness
Explore the evolution from local LLMs to modern agentic CLIs, focusing on the "harness" that gives models context, tools, and autonomy.
#1737: Nous Research: The Decentralized AI Lab Beating Giants
Meet Nous Research, the decentralized collective outperforming billion-dollar labs with open-source AI and the self-improving Hermes-Agent framework.
#1731: Why Deep Research Agents Are Being Forgotten
Specialized research agents outperform general orchestrators by 40-60% on verification tasks, yet developer hype is fading. Here's why.
#1728: The AI Carpool: Emergent Collaboration Through Role-Playing
CAMEL AI lets two agents role-play to solve tasks autonomously. No complex code—just emergent teamwork.
#1727: The Great Architectural Heist: LSP as AI's Universal Plumbing
Explore how the Language Server Protocol is being repurposed to integrate AI directly into code editors, unifying development workflows.
#1725: The Death of the Lonely Chatbot
Forget chatbots: AI orchestration is now the key to scaling intelligent agents in the enterprise.
#1713: Why Native AI Search Grounding Still Fails
Native search grounding is expensive and flaky. Here’s why bolt-on tools still win for accurate, real-time AI answers.
#1708: Why Your AI Agent Forgets Everything (And How to Fix It)
Learn how Letta's memory-first architecture solves the AI context bottleneck for long-term agents.
#1700: Can LLMs Learn Continuously Without Forgetting?
We explore a new approach: micro-training updates every few days to keep AI knowledge fresh without constant web searches.
#1666: The Agent Mesh: Shared Context That Changes Everything
Grok 4.20’s native multi-agent architecture cuts token costs by 75% and enables real-time cross-agent reasoning.
#1629: From DAGs to Loops: Why Agents Need Stateful Cycles
Stop building linear chains and start building cycles to create agents that can reason, self-correct, and maintain complex state.
#1601: Cohere: The Switzerland of Enterprise AI
While others chase viral memes, Cohere is quietly building the secure, cloud-agnostic infrastructure powering the global enterprise.
#1592: The Vector Debt Trap: Choosing Embeddings That Last
Stop treating embedding models like plumbing. Learn how to navigate vector debt, multimodal retrieval, and database configuration for RAG.
#1565: Machine-Readable Safety: Markdown for AI Agents
Transform bloated government data into clean Markdown to power life-saving AI agents during emergencies.
#1482: The Hidden Cost of Choosing an Embedding Model
From Matryoshka models to multimodal search, discover how the fundamental units of AI memory are being optimized for efficiency and scale.
#1212: The Postgres Vector Revolution: Killing the Sprawl
Is your tech stack a sprawling suburb of microservices? Discover why a 40-year-old database is winning the AI infrastructure war.
#1123: When One Database Isn't Enough
Can Postgres 18 finally replace the data warehouse? We dive into data gravity, columnar storage, and the physics of scaling in the AI age.