The technology of voice and sound. From text-to-speech systems and voice cloning to speech recognition and audio engineering, this channel covers the cutting edge of how machines learn to speak, listen, and sound convincingly human.
#5218: What Lip Reading Can and Can't Do
Only 30–40% of speech sounds are visible on the lips. Here's what speechreading actually is — and where the spy-movie version falls apart.
#5128: The 120 Hz Problem: Why Your Window Can't Stop That Jackhammer
A bedroom-shattering 120 Hz tone reveals why low-frequency noise ignores windows, mocks active noise cancellation, and demands real physics.
#4799: How to Build Your Own TTS Audiobooks
From M4B containers to chapter timing drift — the technical pipeline for creating audiobooks with synthetic voice.
#4668: Why Podcast AI Voices Sound Too Perfect
We dig into why AI podcast voices sound too clean—and how TTS is learning to stumble, overlap, and interrupt convincingly.
#4666: Why 30 Seconds of Audio Beats 3 Minutes for Voice Cloning
Why adding more audio made Daniel's voice clones worse — and what it reveals about how voice embeddings actually work.
#4456: Inside the Podcast Pipeline: How 15 Weekly Episodes Get Made
From prompt to published episode — a full walkthrough of the automated production system running 15 shows weekly.
#4376: Which Mic Actually Lowers Word Error Rate?
Laptop mics hit 18% WER. A $70 mic drops it to 4%. Here's what actually works for voice-first dictation.
#4312: Why Speech-to-Text Still Fails at Its Own Name
When OpenAI's Whisper misheard its own name as "Wispr," it revealed why 95% word accuracy still isn't good enough.
#3446: Where to Clip a Speaker for the Best Sound
Tiny placement changes dramatically alter sound. Learn the physics of where to clip your speaker for the best audio.
#3189: Drawing the Melody: SSML's Hidden Power
How SSML gives developers narrative control over AI voices — and why ElevenLabs became its center of gravity.
#3020: How Chatterbox Locks Your Voice Clone Across Thousands of Generations
Why most single-shot TTS models drift over time—and how Chatterbox's cached embedding approach solves it.
#2982: Why Your TTS Model Nails "Shabbat" but Not "Keren Hishtalmut
Why multilingual TTS models handle loanwords but fail at niche vocabulary — and what you can do about it.
#2914: Can AI Read the Room? TTS Prosody Explained
Can TTS models truly infer emotion from text, or just mimic patterns? We break down the science of prosody.
#2886: How Acoustic Cameras Catch Honking Drivers
Can an acoustic camera pinpoint one honk in a traffic jam? The tech is real, and fines are being issued.
#2781: When Voice AI Features Enable Fraud
Voice AI platforms now let you simulate background noise, hesitation, and natural conversation — and that's a problem.
#2754: Why Your Dictation Setup Might Be Wrong
Modern ASR is shockingly robust. The biggest predictor of accuracy? How well your audio matches its training data.
#2707: Foot Pedals vs USB Buttons: The Ergonomics of Dictation
Foot pedals, USB buttons, and under-desk macro pads for voice dictation — a deep dive into the hardware that makes AI dictation work.
#2618: Text Normalization's Hidden Complexity
How to handle acronyms in text-to-speech pipelines using BERT models, lexicons, and layered preprocessing.
#2602: Letting Non-Experts Direct Audio Tools Through Conversation
How to use AI for podcast mastering — and why agentic AI works better for small tasks than big promises.
#2591: Decoupling Script from Voice
How dynamic voice replacement could let listeners choose who narrates each host's lines.