← All Tags

#speech-recognition

49 episodes

#4667: How Transformers Killed the Robot Voice

From espeak's robotic squawk to neural voices with added "ums" — how transformers made speech synthesis human.

text-to-speechtransformersspeech-recognition

#4666: Why 30 Seconds of Audio Beats 3 Minutes for Voice Cloning

Why adding more audio made Daniel's voice clones worse — and what it reveals about how voice embeddings actually work.

voice-cloningaudio-processingspeech-recognition

#4641: Mozilla's Hidden Projects: Beyond the Browser

Mozilla is more than just Firefox. Discover Common Voice, Monitor, Thunderbird, and Rust — the public infrastructure projects shaping the open web.

open-sourcespeech-recognitiondigital-privacy

#4497: Voice-First Stack Consolidation

Two transcription apps, two text expanders, no sync — and snippets about to fight each other. The path forward.

voice-firstspeech-recognitiontext-to-speech

#4456: Inside the Podcast Pipeline: How 15 Weekly Episodes Get Made

From prompt to published episode — a full walkthrough of the automated production system running 15 shows weekly.

audio-engineeringgpu-accelerationspeech-recognition

#4376: Which Mic Actually Lowers Word Error Rate?

Laptop mics hit 18% WER. A $70 mic drops it to 4%. Here's what actually works for voice-first dictation.

audio-engineeringspeech-recognitionaudio-quality

#4312: Why Speech-to-Text Still Fails at Its Own Name

When OpenAI's Whisper misheard its own name as "Wispr," it revealed why 95% word accuracy still isn't good enough.

speech-recognitionautomatic-speech-recognitionhallucinations

#3854: From Coos to Conversation: Baby's Hidden On-Ramp

How do babies go from babbling to real back-and-forth dialogue? The hidden architecture of early conversation.

child-developmentspeech-recognitionneurodivergence

#3493: Murmuring Scriptures and Wandering Wilds: Ancient Meditation

How "hagah" (murmuring scripture) and "hitbodedut" (wilderness solitude) reveal meditation hidden in the Bible.

neurodivergencechild-developmentspeech-recognition

#3443: What Makes a Pediatrician's Diagnostic Skill Unique

How pediatricians diagnose without patient history, reading cries, body language, and parent-child dynamics.

child-developmentsensory-processingspeech-recognition

#3363: Why the Teletubbies Sun-Baby Makes Infants Cry

The Teletubbies was engineered for pre-verbal brains. Here's why adult discomfort is a feature, not a bug.

child-developmentsensory-processingspeech-recognition

#2801: Why Baby Babble Sounds Like Foreign Languages

Your baby isn't speaking Korean — but here's why the overlap isn't a coincidence.

child-developmentlinguisticsspeech-recognition

#2754: Why Your Dictation Setup Might Be Wrong

Modern ASR is shockingly robust. The biggest predictor of accuracy? How well your audio matches its training data.

automatic-speech-recognitionspeech-recognitionaudio-processing

#2643: How Stenographers Type 300 Words Per Minute

Court reporters don’t type letters—they chord syllables at 300 words per minute. Here’s how it works and why AI can’t replace them yet.

speech-recognitionaudio-processingaccessibility

#2618: Text Normalization's Hidden Complexity

How to handle acronyms in text-to-speech pipelines using BERT models, lexicons, and layered preprocessing.

text-to-speechspeech-recognitionaudio-processing

#2590: The Uncanny Valley of Clean Speech

How transformer models distinguish "um" from meaningful speech — and why removing too much makes you sound like a robot.

speech-recognitionaudio-processingautomatic-speech-recognition

#2582: What Your Browser Does to Mic Audio Before It Reaches Your Server

getUserMedia returns audio, but not raw audio. Here's what browsers actually do to your mic feed before it hits your server.

audio-processingspeech-recognitionbrowser-audio-pipeline

#2563: How Audio Fingerprinting Actually Works

Spectrogram peaks, constellation maps, and hash matching — the elegant mechanics behind identifying any song in seconds.

audio-processingsignal-processingspeech-recognition

#2543: Why Base64 Adds 33% Overhead (And Why You Still Need It)

Base64 isn’t compression — it’s a safe transport encoding. Here’s how it works with audio APIs and where its limits are.

audio-engineeringspeech-recognitionapi-integration

#2510: The Design That Makes Voice Agents Tolerable

Drive-thru accuracy, healthcare triage, and the design secret that makes people *want* to talk to a machine.

voice-firstaccessibilityspeech-recognition

#2486: Why Noise Reduction Can Ruin Transcription Accuracy

Cleaning audio before transcription can increase errors by up to 46%. Here's the right approach for your voice app.

speech-recognitionaudio-processingautomatic-speech-recognition

#2479: The Screaming Baby Stress Test

Choosing the right headset and control method for dictation when you're holding a baby who won't stop screaming.

speech-recognitionvoice-firstdiy

#2443: How Podcast RSS Feeds Can Speak Every Language

One RSS feed, a transcript tag, and TTS voice cloning — the emerging standard for letting any podcast speak any language.

speech-recognitionvoice-cloningaudio-processing

#2337: When Diarization Fails Silently

Discover how PyAnnote and other tools tackle the critical task of identifying "who spoke when" in audio—and why it’s harder than it sounds.

audio-processingspeech-recognitionautomatic-speech-recognition