Audio & Speech

Speech recognition, TTS, voice cloning, audio engineering

74 episodes Page 3 of 4

#1546: The Death of Latency: Three Pillars of Modern Voice AI

Say goodbye to the "digital sandwich." Explore the three architectural pillars closing the latency gap in modern speech recognition.

#1540: Why Gnome 50 is Breaking Your Voice-to-Text Tools

Explore the engineering battle to bring low-latency AI voice input to Linux while navigating the strict security of Wayland and GNOME 50.

voice-to-textlocal-inferencelatency

#1539: Escaping the Cloud Dictation Trap

Stop shouting at your phone. Discover how dedicated hardware and local AI are making instant, private voice-to-text a reality.

speech-recognitionedge-computinghardware-engineering

#1218: The Architectural Divide Between Batch and Live Speech

Why does voice typing feel so clunky compared to recording a memo? We explore the technical hurdles of real-time AI transcription.

#992: Beyond the Digital Sandwich: The Future of Voice AI

Is speech recognition dead? Explore how multimodal models are replacing the "digital sandwich" with true intent-based reasoning.

local-aiquantizationvoice-ai

#947: Pro Audio in Acoustic Nightmares: Mobile Recording Tips

Learn how to turn a marble-floored room into a studio using your phone, simple blankets, and the right USB-C gear.

audio-engineeringmobile-recordingacoustic-treatment

#877: When Bloating Steals Your Breath

Learn how chronic bloating impacts vocal performance and discover mechanical workarounds to reclaim your breath and resonance.

vocal-performancedigestive-healthrespiratory-mechanics

#875: The Single-Ear Solution: Audio for Situational Awareness

Discover why a dedicated mono earbud is the ultimate tool for situational awareness while parenting or multitasking.

situational-awarenesstelecommunicationssensory-processing

#868: When Your Phone's Mic Beats Your Expensive Gear

Stop holding your phone like a piece of toast. Explore the best mobile microphone setups for high-quality AI voice transcription.

telecommunicationsaudio-engineeringspeech-recognition

#749: The Live vs. Scripted Trade-Off in AI Podcasting

Can AI podcasts move from polished scripts to raw, real-time conversation? Explore the technical and financial shift to live multimodal models.

large-language-modelsarchitecturemultimodal-ai

#732: Why Your Recorded Voice Sounds Wrong

Use AI to find your perfect EQ profile and build a pro vocal chain. Fix nasality, master de-essing, and sound your best on any device.

audio-engineeringaudio-processingaudio-qualitycomputational-audio

#727: The Math of Immersion: How 360-Degree Sound Actually Works

Learn how object-based audio and clever math trick your brain into hearing 360-degree sound from even the smallest mobile devices.

sensory-processingspatial-audiocomputational-audio

#725: Finding a Speaker That Loves Voices

Stop listening to podcasts through tinny speakers. Learn how to choose hardware optimized for the human voice and clear, room-filling audio.

smart-homeaudio-engineeringcomputational-audio

#720: Why Your Ears Prefer Imperfect Plastic to Perfect Pixels

Why do we still buy plastic discs in an age of neural-link streaming? Explore the science of analog warmth and the "ritual" of the record.

sensory-processinganalog-audiodigital-compression

#682: Why Your Phone Mic Beats Your Studio Headset

Why does a phone mic outperform a pro headset for AI transcription? Herman and Corn dive into the physics of MEMS and the truth about audio quality.

speech-recognitionaudio-hardwaresemiconductorssignal-processinghardware-engineering

#660: The Bit Rate Dilemma: How Much Audio Data Do You Need?

Herman and Corn explore the science of audio compression, psychoacoustics, and finding the perfect bit rate for podcasts and AI.

audio-processingdata-integritypsychoacoustics

#647: The Golden Rule of Audio Engineering

Why does digital data need to become analog? Explore the physics of sound and the critical role of the DAC in modern audio engineering.

audio-engineeringsignal-processingdigital-to-analog

#598: Audio Engineering as Prompt Engineering: Better Sound, Better AI

Can better audio quality actually make an AI smarter? Discover how audio post-production functions as a new form of prompt engineering.

prompt-engineeringlarge-language-modelsaudio-engineering

#233: How Math Gives Microphones Directional Ears

Discover how math and physics turn simple microphones into "sound spotlights" that can isolate a single voice in even the noisiest environments.

beamforming-technologymicrophone-arraysdigital-signal-processing

#196: Why Your Irish Accent Sounds American

Herman and Corn dive into the mechanics of neural text-to-speech, exploring how AI masters human prosody and the "average voice" accent problem.

neural-text-to-speechvoice-cloninggenerative-modeling