#ai-inference
38 episodes · Page 2 of 2
#2040: The Rebellion Against Big Tech's AI Lock-In
Why run LLMs locally? We break down Ollama, llama.cpp, vLLM, and llamafile—and when to use each.
#2022: When AI Becomes Your IT Department
We dug into a repo of 47 real-world projects showing how OpenClaw powers everything from self-healing servers to overnight app builders.
#1831: The 79% AI Coder: Reasoning vs. Memorization
AI models now score 79% on coding benchmarks, but a 40-point drop on harder tests reveals the truth.
#1782: Jenkins, GitHub, or Tekton? Picking Your 2025 CI/CD Engine
Jenkins is still the COBOL of DevOps, but the "one size fits all" model is dead. Here’s how to pick your pipeline.
#1756: The Ferrari in the Mud: Prestige Flops
We count down the five worst serious movies of the last five years, starting with a sci-fi disaster that wasted $80 million.
#1620: Why VRAM Is the Wrong Way to Measure Your AI PC
Forget VRAM—bandwidth is the new king. Discover why your local AI feels slow and how to build a true "agent computer" for professional coding.
#1556: The War Against Latency: Engineering Real-Time AI
From KV cache monsters to sub-100ms response times, explore the hardware and software innovations making real-time AI a reality.
#1479: The Speed of Thought: Inside the New Era of Inference
The war for model size is over. Explore the engineering breakthroughs making massive AI models faster than human thought.
#1084: Why AI Models Can’t Read and Your Bill Is Rising
Why does the same prompt cost more on different models? Discover the "invisible wall" of tokenization and how it shapes AI perception.
#1056: The Vocabulary Myth: Do More Words Equal Better Thinking?
Does a massive vocabulary lead to deeper thoughts? Explore the hidden mechanics of English, Hebrew, and the famous "Inuit snow" myth.
#671: Keys to the Kingdom: Securing AI Model Weights
How do AI labs share their models without losing the secret sauce? Explore the tech keeping Claude secure in the Pentagon’s hands.
#484: The Silicon Sharing Economy: Inside Serverless GPUs
How do small teams run massive AI models without $50,000 chips? Corn and Herman dive into the hidden plumbing of serverless GPU providers.
#48: Renting vs. Building: The Hidden Choices in AI Inference
Ever wonder how AI magic happens? We demystify AI inference, exploring where and how models truly operate.
#38: Why Local AI Inference Is Beating the Cloud
AI supercomputers are landing on your desk! Discover why local AI is indispensable for enterprises facing API costs, latency, and privacy.