Audio & Speech
Speech recognition, TTS, voice cloning, audio engineering
#5386: Morse Code Between Two Phones: Does It Actually Work?
Two phones, no data, one room. Morse over sound and flashlight — what the research says about echoes, timing, and whether anyone will hear you.
#5384: Android ASR Runtimes: LiteRT, ExecuTorch, and Why Your Phone Has No VRAM
Why does your phone have no VRAM number? A tour of Android's runtime layer and what it takes to run ASR locally.
#5378: SIP Speakers, Multicast, and the One-Way Page
How SIP one-way paging actually works — and why multicast turns 500 speakers into a single call.
#5328: A Shabbat Radio That Can't Be Paused
A soundproof box that plays podcasts continuously — no pause, no skip, no volume. Just open the lid and listen.
#5319: Focus Stacking Without the Phone App Guesswork
Your phone app wants three shots. The optics want three hundred. Here's how to actually stack macro focus.
#5259: When Actors Watch Themselves: Tarantino's Laugh Track
Tarantino laughs at his own films in Tel Aviv theaters. Most actors can't even watch theirs once. What's behind the gap?
#5128: The 120 Hz Problem: Why Your Window Can't Stop That Jackhammer
A bedroom-shattering 120 Hz tone reveals why low-frequency noise ignores windows, mocks active noise cancellation, and demands real physics.
#4799: How to Build Your Own TTS Audiobooks
From M4B containers to chapter timing drift — the technical pipeline for creating audiobooks with synthetic voice.
#4668: Why Podcast AI Voices Sound Too Perfect
We dig into why AI podcast voices sound too clean—and how TTS is learning to stumble, overlap, and interrupt convincingly.
#4666: Why 30 Seconds of Audio Beats 3 Minutes for Voice Cloning
Why adding more audio made Daniel's voice clones worse — and what it reveals about how voice embeddings actually work.
#4456: Inside the Podcast Pipeline: How 15 Weekly Episodes Get Made
From prompt to published episode — a full walkthrough of the automated production system running 15 shows weekly.
#4376: Which Mic Actually Lowers Word Error Rate?
Laptop mics hit 18% WER. A $70 mic drops it to 4%. Here's what actually works for voice-first dictation.
#4312: Why Speech-to-Text Still Fails at Its Own Name
When OpenAI's Whisper misheard its own name as "Wispr," it revealed why 95% word accuracy still isn't good enough.
#3446: Where to Clip a Speaker for the Best Sound
Tiny placement changes dramatically alter sound. Learn the physics of where to clip your speaker for the best audio.
#3189: Drawing the Melody: SSML's Hidden Power
How SSML gives developers narrative control over AI voices — and why ElevenLabs became its center of gravity.
#3020: How Chatterbox Locks Your Voice Clone Across Thousands of Generations
Why most single-shot TTS models drift over time—and how Chatterbox's cached embedding approach solves it.
#2982: Why Your TTS Model Nails "Shabbat" but Not "Keren Hishtalmut
Why multilingual TTS models handle loanwords but fail at niche vocabulary — and what you can do about it.
#2914: Can AI Read the Room? TTS Prosody Explained
Can TTS models truly infer emotion from text, or just mimic patterns? We break down the science of prosody.
#2886: How Acoustic Cameras Catch Honking Drivers
Can an acoustic camera pinpoint one honk in a traffic jam? The tech is real, and fines are being issued.
#2781: When Voice AI Features Enable Fraud
Voice AI platforms now let you simulate background noise, hesitation, and natural conversation — and that's a problem.