Audio & Speech

Speech recognition, TTS, voice cloning, audio engineering

77 episodes RSS Feed

The technology of voice and sound. From text-to-speech systems and voice cloning to speech recognition and audio engineering, this channel covers the cutting edge of how machines learn to speak, listen, and sound convincingly human.

#5218: What Lip Reading Can and Can't Do

Only 30–40% of speech sounds are visible on the lips. Here's what speechreading actually is — and where the spy-movie version falls apart.

speech-recognitionsensory-processinglinguistics

#5128: The 120 Hz Problem: Why Your Window Can't Stop That Jackhammer

A bedroom-shattering 120 Hz tone reveals why low-frequency noise ignores windows, mocks active noise cancellation, and demands real physics.

audio-engineeringsignal-processingsound-transmission

#4799: How to Build Your Own TTS Audiobooks

From M4B containers to chapter timing drift — the technical pipeline for creating audiobooks with synthetic voice.

text-to-speechaudio-engineeringandroid

#4668: Why Podcast AI Voices Sound Too Perfect

We dig into why AI podcast voices sound too clean—and how TTS is learning to stumble, overlap, and interrupt convincingly.

text-to-speechmultimodal-aiconversational-ai

#4666: Why 30 Seconds of Audio Beats 3 Minutes for Voice Cloning

Why adding more audio made Daniel's voice clones worse — and what it reveals about how voice embeddings actually work.

voice-cloningaudio-processingspeech-recognition

#4456: Inside the Podcast Pipeline: How 15 Weekly Episodes Get Made

From prompt to published episode — a full walkthrough of the automated production system running 15 shows weekly.

audio-engineeringgpu-accelerationspeech-recognition

#4376: Which Mic Actually Lowers Word Error Rate?

Laptop mics hit 18% WER. A $70 mic drops it to 4%. Here's what actually works for voice-first dictation.

audio-engineeringspeech-recognitionaudio-quality

#4312: Why Speech-to-Text Still Fails at Its Own Name

When OpenAI's Whisper misheard its own name as "Wispr," it revealed why 95% word accuracy still isn't good enough.

speech-recognitionautomatic-speech-recognitionhallucinations

#3446: Where to Clip a Speaker for the Best Sound

Tiny placement changes dramatically alter sound. Learn the physics of where to clip your speaker for the best audio.

audio-engineeringacousticsspeaker-placement

#3189: Drawing the Melody: SSML's Hidden Power

How SSML gives developers narrative control over AI voices — and why ElevenLabs became its center of gravity.

text-to-speechaudio-engineeringconversational-ai

#3020: How Chatterbox Locks Your Voice Clone Across Thousands of Generations

Why most single-shot TTS models drift over time—and how Chatterbox's cached embedding approach solves it.

voice-cloningopen-source-aispeech-to-speech

#2982: Why Your TTS Model Nails "Shabbat" but Not "Keren Hishtalmut

Why multilingual TTS models handle loanwords but fail at niche vocabulary — and what you can do about it.

text-to-speechtokenizationfine-tuning

#2914: Can AI Read the Room? TTS Prosody Explained

Can TTS models truly infer emotion from text, or just mimic patterns? We break down the science of prosody.

text-to-speechspeech-to-speechaudio-processing

#2886: How Acoustic Cameras Catch Honking Drivers

Can an acoustic camera pinpoint one honk in a traffic jam? The tech is real, and fines are being issued.

audio-processingsignal-processingurban-planning

#2781: When Voice AI Features Enable Fraud

Voice AI platforms now let you simulate background noise, hesitation, and natural conversation — and that's a problem.

voice-cloningai-ethicsfinancial-fraud

#2754: Why Your Dictation Setup Might Be Wrong

Modern ASR is shockingly robust. The biggest predictor of accuracy? How well your audio matches its training data.

automatic-speech-recognitionspeech-recognitionaudio-processing

#2707: Foot Pedals vs USB Buttons: The Ergonomics of Dictation

Foot pedals, USB buttons, and under-desk macro pads for voice dictation — a deep dive into the hardware that makes AI dictation work.

ergonomicsaudio-engineeringhardware-engineering

#2618: Text Normalization's Hidden Complexity

How to handle acronyms in text-to-speech pipelines using BERT models, lexicons, and layered preprocessing.

text-to-speechspeech-recognitionaudio-processing

#2602: Letting Non-Experts Direct Audio Tools Through Conversation

How to use AI for podcast mastering — and why agentic AI works better for small tasks than big promises.

audio-engineeringconversational-aiai-agents

#2591: Decoupling Script from Voice

How dynamic voice replacement could let listeners choose who narrates each host's lines.

voice-cloningtext-to-speechaudio-processing