The technology of voice and sound. From text-to-speech systems and voice cloning to speech recognition and audio engineering, this channel covers the cutting edge of how machines learn to speak, listen, and sound convincingly human.
#4668: Why Podcast AI Voices Sound Too Perfect
We dig into why AI podcast voices sound too clean—and how TTS is learning to stumble, overlap, and interrupt convincingly.
#4666: Why 30 Seconds of Audio Beats 3 Minutes for Voice Cloning
Why adding more audio made Daniel's voice clones worse — and what it reveals about how voice embeddings actually work.
#4456: Inside the Podcast Pipeline: How 15 Weekly Episodes Get Made
From prompt to published episode — a full walkthrough of the automated production system running 15 shows weekly.
#4376: Which Mic Actually Lowers Word Error Rate?
Laptop mics hit 18% WER. A $70 mic drops it to 4%. Here's what actually works for voice-first dictation.
#4312: Why Speech-to-Text Still Fails at Its Own Name
When OpenAI's Whisper misheard its own name as "Wispr," it revealed why 95% word accuracy still isn't good enough.
#3446: Where to Clip a Speaker for the Best Sound
Tiny placement changes dramatically alter sound. Learn the physics of where to clip your speaker for the best audio.
#3189: Drawing the Melody: SSML's Hidden Power
How SSML gives developers narrative control over AI voices — and why ElevenLabs became its center of gravity.
#3020: How Chatterbox Locks Your Voice Clone Across Thousands of Generations
Why most single-shot TTS models drift over time—and how Chatterbox's cached embedding approach solves it.
#2982: Why Your TTS Model Nails "Shabbat" but Not "Keren Hishtalmut
Why multilingual TTS models handle loanwords but fail at niche vocabulary — and what you can do about it.
#2914: Can AI Read the Room? TTS Prosody Explained
Can TTS models truly infer emotion from text, or just mimic patterns? We break down the science of prosody.
#2886: How Acoustic Cameras Catch Honking Drivers
Can an acoustic camera pinpoint one honk in a traffic jam? The tech is real, and fines are being issued.
#2781: When Voice AI Features Enable Fraud
Voice AI platforms now let you simulate background noise, hesitation, and natural conversation — and that's a problem.
#2754: Why Your Dictation Setup Might Be Wrong
Modern ASR is shockingly robust. The biggest predictor of accuracy? How well your audio matches its training data.
#2707: Foot Pedals vs USB Buttons: The Ergonomics of Dictation
Foot pedals, USB buttons, and under-desk macro pads for voice dictation — a deep dive into the hardware that makes AI dictation work.
#2618: Text Normalization's Hidden Complexity
How to handle acronyms in text-to-speech pipelines using BERT models, lexicons, and layered preprocessing.
#2602: Letting Non-Experts Direct Audio Tools Through Conversation
How to use AI for podcast mastering — and why agentic AI works better for small tasks than big promises.
#2591: Decoupling Script from Voice
How dynamic voice replacement could let listeners choose who narrates each host's lines.
#2590: The Uncanny Valley of Clean Speech
How transformer models distinguish "um" from meaningful speech — and why removing too much makes you sound like a robot.
#2582: What Your Browser Does to Mic Audio Before It Reaches Your Server
getUserMedia returns audio, but not raw audio. Here's what browsers actually do to your mic feed before it hits your server.
#2563: How Audio Fingerprinting Actually Works
Spectrogram peaks, constellation maps, and hash matching — the elegant mechanics behind identifying any song in seconds.