DeepSeek
Chinese artificial intelligence company
Episodes
-
#5673: Why Lowering Temperature Broke Daniel's ScriptsLowering temperature should make output safer, not incoherent. Daniel's DeepSeek scripts say otherwise — and the reason may be stranger than he thi...Main topic -
#5188: DeepSeek's Point Release That Isn'tDeepSeek shipped a whole new architecture and called it a point release. Here's what actually changed inside the model.Main topic -
#5182: DeepSeek V4.1 Flash: 1M Context, 437x Smaller KV CacheDeepSeek V4.1 Flash landed with a 1M-token window and a KV cache 437x smaller than V1. Here's what actually changed — and why the middle of your co...Main topic -
#5154: DeepSeek's Two-Endpoint PhilosophyDeepSeek quietly routed Pro traffic to Flash — and that routing change says everything about its two-endpoint strategy.Main topic -
#5830: Serving Your Own Fine-Tuned Model in the CloudYou fine-tuned an open-weight model. Now how does anyone actually talk to it? Dedicated GPUs vs serverless inference, and the math that decides it. -
#5520: Keeping a Fine-Tune Alive Across Model ReleasesDaniel wants to fine-tune DeepSeek Flash 4.1 on edited podcast scripts. The hard part isn't training — it's surviving the next release. -
#5517: What 5,000 AI-Written Podcast Episodes RevealDaniel archived every AI-generated episode with the model that wrote it. Now he wants to know which models loop — and how to catch it. -
#5393: Fine-Tuning a Model on 100 Hand-Edited AnswersYou don't need 10,000 examples to make a model sound like you. The real number is closer to 100 — if the edits are opinionated.