Audio & Speech

Speech recognition, TTS, voice cloning, audio engineering

102 episodes Page 4 of 6

#1808: The Architecture That Made AI Voices Run on a Raspberry Pi

How a model the size of a tweet outperforms billion-dollar giants in the race for perfect AI speech.

open-source-aismall-language-modelstext-to-speech

#1800: Hacking the Brain's Alarm System

Why some sounds make your skin crawl: the science of emergency alerts.

audio-processinghuman-computer-interactionemergency-preparedness

#1778: Audio Is the New "Read Later" Graveyard

Why listening to AI conversations beats reading dense PDFs, and how serverless GPUs make it cheap.

audio-processingserverless-gpurag

#1752: Whisper Small Beats Whisper Large in Speed & Accuracy

A 4GPU benchmark on Ubuntu shows the 1.5B parameter Whisper Large is slower and less accurate than the tiny Whisper Small.

speech-recognitiongpu-accelerationlatency

#1724: When AI Dubbing Swaps Your Gender

How does YouTube translate a video with one click? We explore the tech behind auto-dubbing, from sandwich models to voice cloning.

speech-to-speechvoice-cloningmultimodal-ai

#1568: The Signal Versus Symbol Gap

Is Gemini a brilliant audio engineer or just a talented lip-reader? Explore the "signal vs. symbol" gap in AI audio processing.

multimodal-aiaudio-processinghallucinations

#1564: The Death of the Cascaded Pipeline

Forget basic transcription. Explore how native omni-modal models are capturing the "soul" of speech with near-instant latency.

multimodal-aispeech-to-speechvoice-first

#1555: Beyond Whisper: NVIDIA’s Real-Time Speech Revolution

Move over Whisper. NVIDIA's new models offer 10x speed increases and better accuracy for real-time speech-to-text.

#1546: The Death of Latency: Three Pillars of Modern Voice AI

Say goodbye to the "digital sandwich." Explore the three architectural pillars closing the latency gap in modern speech recognition.

#1540: Why Gnome 50 is Breaking Your Voice-to-Text Tools

Explore the engineering battle to bring low-latency AI voice input to Linux while navigating the strict security of Wayland and GNOME 50.

voice-to-textlocal-inferencelatency

#1539: Escaping the Cloud Dictation Trap

Stop shouting at your phone. Discover how dedicated hardware and local AI are making instant, private voice-to-text a reality.

speech-recognitionedge-computinghardware-engineering

#1218: The Architectural Divide Between Batch and Live Speech

Why does voice typing feel so clunky compared to recording a memo? We explore the technical hurdles of real-time AI transcription.

#992: Beyond the Digital Sandwich: The Future of Voice AI

Is speech recognition dead? Explore how multimodal models are replacing the "digital sandwich" with true intent-based reasoning.

local-aiquantizationvoice-ai

#947: Pro Audio in Acoustic Nightmares: Mobile Recording Tips

Learn how to turn a marble-floored room into a studio using your phone, simple blankets, and the right USB-C gear.

audio-engineeringmobile-recordingacoustic-treatment

#877: When Bloating Steals Your Breath

Learn how chronic bloating impacts vocal performance and discover mechanical workarounds to reclaim your breath and resonance.

vocal-performancedigestive-healthrespiratory-mechanics

#875: The Single-Ear Solution: Audio for Situational Awareness

Discover why a dedicated mono earbud is the ultimate tool for situational awareness while parenting or multitasking.

situational-awarenesstelecommunicationssensory-processing

#868: When Your Phone's Mic Beats Your Expensive Gear

Stop holding your phone like a piece of toast. Explore the best mobile microphone setups for high-quality AI voice transcription.

telecommunicationsaudio-engineeringspeech-recognition

#749: The Live vs. Scripted Trade-Off in AI Podcasting

Can AI podcasts move from polished scripts to raw, real-time conversation? Explore the technical and financial shift to live multimodal models.

large-language-modelsarchitecturemultimodal-ai

#732: Why Your Recorded Voice Sounds Wrong

Use AI to find your perfect EQ profile and build a pro vocal chain. Fix nasality, master de-essing, and sound your best on any device.

audio-engineeringaudio-processingaudio-qualitycomputational-audio

#727: The Math of Immersion: How 360-Degree Sound Actually Works

Learn how object-based audio and clever math trick your brain into hearing 360-degree sound from even the smallest mobile devices.

sensory-processingspatial-audiocomputational-audio