So my cousin Miriam, the one who had meningitis as a kid, she didn't just lose the hearing in her left ear and move on. She spent about two years actively learning to read lips, and she still does it now without thinking. I remember visiting her when I was in medical school and she'd be watching the television with the sound nearly off, just following along.
Daniel picked up on that detail from the meningitis episode and it sent him down a whole path. He's been wondering about lip reading ever since, and his question this week is sort of three questions stacked under one umbrella. Is lip reading actually a learnable skill, or is it one of those things that looks real in films but falls apart in practice? Second, when we see heads of state leaving events and being rushed into cars, is there someone somewhere trying to decode what they muttered to each other? And third, the part I think he enjoyed writing most, do public figures ever deliberately say inane things, his phrase was mutterings about cabbage and onion, just to throw off anyone watching their lips?
That's a very Daniel question.
It is, and it's worth taking seriously because the fantasy reveals something about how we think about privacy. So today we're asking whether lip reading is real, how it works, who's actually doing it, and whether the cabbage and onion countermeasure has any basis in reality.
And I should say up front, the term lip reading is a little bit like calling surgery cutting. It's not wrong, but it undersells what's actually happening. The clinical term is speechreading, and the distinction matters. Lip reading in the narrow sense is visual only, just the movements of the lips and tongue and jaw. Speechreading is the full package, the lips plus facial expression, gesture, body language, context, and whatever residual hearing the person has. My cousin Miriam wasn't decoding phonemes from lip shapes in isolation. She was combining a slightly degraded visual signal with one fully working ear and a lifetime of knowing how people talk.
So when we say lip reading is hard, part of what we mean is that the narrow version barely exists in practice. The assistive version is always multimodal.
Right. And the pop culture version, the spy across the room with binoculars decoding every word through a window, that's the narrow version with impossible accuracy. That's the one that doesn't exist.
So before we get to heads of state and cabbage, let's talk about what's actually happening when someone reads lips. What does the visual signal actually contain?
Less than people assume. The core problem is something called the viseme. A phoneme is the smallest unit of sound that distinguishes meaning, like the difference between pat and bat. A viseme is the visual equivalent, the mouth shape that corresponds to a group of sounds. And here's the problem. Multiple phonemes map to a single viseme. The sounds puh, buh, and muh are all made by pressing the lips together. They look identical. You cannot tell them apart visually.
So pat, bat, and mat are the same word on the lips.
And it gets worse. The sounds kuh and guh are made at the back of the mouth, the velum. The tongue movement is almost invisible. So cat and gat, if gat were a word, look the same. The standard estimate is that only about thirty to forty percent of speech sounds are visually distinguishable. The rest are what linguists call homophenous, same face, same look.
Thirty to forty percent is a brutal ceiling. That means even a perfect lip reader, with perfect lighting and a perfect angle, is getting less than half the raw signal.
And that's before you account for the second problem, which is coarticulation. The lips don't move in discrete units. When I say the word spoon, my lips are already rounding for the oo sound while I'm still producing the s. The movements blend together. There's no clean segmentation. So the lip reader isn't seeing a sequence of visemes like beads on a string. They're seeing a continuous flow of movement where every shape is influenced by the shapes around it.
This is why the classic demonstration works. The phrase elephant juice and the phrase I love you look nearly identical on the lips.
They do. If you mouth elephant juice at someone, most people will read it as I love you. It's a party trick, but it's also a perfect illustration of the ambiguity problem. The visual signal underspecifies the message, so the brain fills in the gaps using context and expectation.
And the brain fills those gaps so aggressively that it will override what the ears are actually hearing. That's the McGurk effect.
The McGurk effect is the single best piece of evidence that visual speech isn't a separate channel from auditory speech. It's integrated at a very low level. You play a sound, the syllable ba, and you pair it with a video of a person saying ga. What people report hearing is da, or sometimes something else entirely. The visual information changes what you perceive auditorily. Your brain takes the two conflicting signals and produces a compromise that matches neither.
Which means the visual signal isn't just a supplement for when hearing fails. It's part of how hearing works in the first place.
That's the key insight. We're all lip reading all the time, just below conscious awareness. If you've ever struggled to understand someone at a loud party and then turned to look at their face and suddenly the words snapped into focus, that's visual speech integration doing its job. The difference for someone like my cousin is that she's learned to make that process deliberate and effortful.
So let's talk about the learning part. Daniel's first question was whether this is actually learnable. Can an adult decide one day to learn lip reading and make meaningful progress?
The honest answer is yes, but with a significant asterisk. There was a review published in the American Journal of Audiology a few years back, Bernstein and colleagues, that looked at the training literature. And the finding was that training does produce measurable gains. People who practice speechreading get better at it. But the gains are modest, and they plateau. You don't get a linear improvement curve where every hour of practice makes you five percent better forever. You improve for a while, then you hit a ceiling that's set by the fundamental ambiguity of the visual signal.
So it's learnable in the sense that chess is learnable. You can get better, but you're never going to see the whole board.
There's a developmental angle too. Children with hearing loss often acquire speechreading implicitly, the way hearing children acquire spoken language. They don't take lessons. They just watch faces constantly and their brains wire themselves to extract maximum information from the visual channel. That's different from an adult who loses hearing later in life and has to learn the skill deliberately. The adult brain can do it, but it's effortful in a way that the child's acquisition never was.
My cousin, Miriam, she lost the hearing in one ear when she was about seven. So she was right at the edge of that developmental window. And she had one good ear, which is a huge advantage. She wasn't trying to reconstruct speech from vision alone. She was using vision to fill in the gaps that the one bad ear left.
That's the realistic assistive use case. Not the silent film star reading lips across a crowded room, but someone with partial hearing using visual cues to disambiguate what the remaining hearing can't quite resolve.
And that's where the fatigue comes in. Speechreading is not passive. It demands sustained attention. You're holding multiple hypotheses in working memory, trying to match the lip movements against possible words, checking them against context, discarding the ones that don't fit. It's cognitively expensive. People who rely on speechreading report genuine exhaustion after long conversations. It's the difference between walking on a paved path and walking on a tightrope. Same direction, very different energy cost.
So the mechanism is clear. The visual signal is radically underspecified, the brain compensates with context and expectation, and the whole process is effortful and error-prone. Now let's talk about who's actually using this skill in the wild, and for what.
The assistive side is the most straightforward. People with hearing loss use speechreading constantly, often without thinking about it. But there's also a professional tier. There are people who get paid to transcribe speech from video when there's no audio. And this is where Daniel's heads of state question gets interesting.
Because the fantasy is that there's someone with a telescope lens aimed at the president's mouth as he walks to the car, decoding every word. And the reality is both more mundane and more complicated.
There are professional lip reading services. The one that gets cited most often is a UK company called 121 Captions, which has a whole section on their site about professional lip reading. They work with law enforcement, journalists, private clients. And the key detail is that they work from high quality footage. Good resolution, good lighting, a clear angle on the face. They're not decoding grainy phone video from across a street.
And even with good footage, the error rate is significant. The transcripts they produce come with caveats. They're interpretations, not verbatim records.
Right. The claim you sometimes see is that a good lip reader can achieve eighty percent accuracy. But that number, when it's cited at all, is for controlled conditions with known context. You know the topic, you know the speaker, you know the vocabulary. In the real world, with unknown context and imperfect footage, the accuracy drops dramatically.
This is where the viral video problem comes in. We've all seen clips of politicians or celebrities supposedly caught saying something outrageous, and the caption says something like lip reader reveals what the president really said. And the transcript is always incredibly juicy.
And often wrong. The incentive structure is completely backwards. A viral video needs a juicy quote. The juicier the quote, the more engagement. So there's enormous pressure to produce a transcript that's interesting, not one that's accurate. And because the visual signal is so ambiguous, there's always room to nudge the interpretation toward the more sensational reading.
The same viseme ambiguity that makes elephant juice and I love you indistinguishable is what makes it possible to read almost anything into a politician's lip movements if you try hard enough.
And once a transcript is out there, it's very hard to correct. The correction never travels as far as the original claim. So you get this situation where a fabricated or heavily edited lip reading transcript becomes the accepted version of what someone said, and no amount of debunking dislodges it.
So now we get to the cabbage and onion question. Daniel's hypothetical is that a head of state, aware that lip readers might be watching, deliberately says something inane to throw them off. Is there any evidence this actually happens?
I haven't found a verified case of a head of state systematically doing this. There are anecdotes about people covering their mouths when they talk, which is a real and sensible countermeasure. If you're having a private conversation in public and you don't want anyone reading your lips, you put your hand over your mouth. That's basic tradecraft. But the deliberate nonsense utterance, the cabbage and onion strategy, that's more of a cultural trope than a documented practice.
Which is interesting, because the trope persists for a reason. The fantasy that someone might be decoding your private mutterings is compelling. It speaks to a deeper anxiety about legibility. We want to believe that our private moments could be read by someone skilled enough, and simultaneously we want to believe we could defeat that reading with a simple trick.
The cabbage and onion thing is a defensive fantasy. It's the idea that you can opt out of surveillance by being absurd. And it's appealing because it suggests the watchers are powerful but also easily fooled. You don't have to change your behavior, just your content. Say something meaningless and the decoder's skill is neutralized.
But that only works if the decoder is actually decoding. If the real threat isn't accurate lip reading but the perception that lip reading might be happening, then saying cabbage and onion doesn't help. The watcher just transcribes cabbage and onion, or more likely, transcribes something juicier and claims you said that instead.
This is the key point Daniel's question is circling. The real privacy implication of lip reading isn't accurate decoding. It's the chilling effect. If people believe their lips can be read, they self-censor. They cover their mouths. They stop having private conversations in public. The perception of surveillance does the work that the surveillance itself can't actually do.
So the assistive and surveillance applications are the same skill pointed in different directions. The hard of hearing person uses speechreading to participate more fully in conversation. The paranoid public figure uses the fear of speechreading to withdraw from public conversation. Same perceptual mechanism, opposite social effect.
The technology is neutral in the middle. The visemes don't care whether you're trying to understand your grandmother or decode a diplomat. The ambiguity is the same. What changes is the application, the intent, and the incentives.
There's one more piece here I want to pull out. Daniel's image of the head of state being rushed into a car, with someone somewhere trying to decode the mutterings. That image assumes a level of competence that the research just doesn't support. At the distance those videos are usually shot, with the angle and the motion and the obstruction, the visual signal is nearly useless.
The car door is closing, the head is turned, the hand is partially covering the mouth, the lighting is bad. A professional lip reader looking at that footage would tell you they can maybe get one word in ten, and that word would be low confidence. The idea that someone is extracting full sentences from that scenario is fantasy.
But the fantasy persists because it's narratively satisfying. It's the same reason we believe in codebreakers who can crack any cipher. We want there to be a hidden skill that makes the world legible.
The people who sell that fantasy, the viral video makers, the tabloids, the spy thrillers, they have every incentive to keep it alive. The truth, which is that lip reading is hard and unreliable and mostly useful as a supplement to hearing, is much less exciting.
Let's sit with the actual answer to Daniel's question for a moment. Is lip reading real? Yes, it's a real perceptual skill with a solid scientific basis. Is it learnable? Yes, adults can improve with training, but the gains are modest and they plateau. Is someone decoding heads of state from video? Probably not with any accuracy, though the attempt is made. Do public figures say cabbage and onion to throw off lip readers? No verified cases, but the fantasy is revealing.
The assistive use case is the one that actually matters. My cousin Miriam isn't a spy. She's a woman who watches television with the sound low because she's learned to extract meaning from faces in a way that supplements her remaining hearing. That's the real story of lip reading. It's a tool for participation, not a weapon for surveillance.
The surveillance fantasy is a projection. We assume our private moments are being decoded because we can imagine decoding someone else's. But the actual skill doesn't support that fantasy. It supports something more modest and more human.
The McGurk effect shows that we're all already reading lips, all the time, without knowing it. The question isn't whether lip reading is real. It's whether we can accept that it's real but limited.
That's the uncomfortable part. We want the superpower. We want the codebreaker. We don't want the tired person at the dinner party straining to follow a conversation.
Hilbert: It was fifty dollars a game.
What was?
Hilbert: What they paid me. Local news station, mid nineties. I'd sit in the editing bay with a tape of the football game and transcribe what the coaches said on the sidelines. They called it professional lip reading. I called it guessing with a timecode.
You were the lip reader.
Hilbert: One of them. They had a rotation. I got the Sunday games because nobody else wanted to come in. Fifty dollars a game, cash, and they'd feed me the footage on VHS. I'd watch the coach on the sideline, hand over his mouth half the time, and I'd write down what I thought he said.
How often were you right?
Hilbert: No idea. Nobody checked. The producer would look at my transcript and circle the lines he liked. Then he'd run them under the highlight reel. The coach would be mouthing something and the caption would say what I'd written. The station didn't care if it was accurate. They cared if it was good television.
The incentive was to make it juicy.
Hilbert: The incentive was to keep the job. If I turned in a transcript that said the coach was talking about blocking assignments, that didn't air. If I turned in one that said he was yelling at the referee, that aired. So I learned to write down the version that would air.
You were manufacturing the viral videos before viral videos existed.
Hilbert: I was manufacturing captions. There's one I still remember. Coach named Delvecchio, worked at a high school upstate. He was mouthing something on the sideline, and I wrote down that he said we need to run the ball more. That's what I thought he said. The producer looked at it and said, can you make it about the refs? So I wrote we need to run the ball more, these refs are killing us. And that's what aired.
The coach never said that.
Hilbert: The coach never said any of it. I watched his lips and I guessed. The refs line was pure invention. They gave me an extra twenty five dollars for it.
You got a bonus for the fabrication.
Hilbert: I got a bonus for the entertainment value. That's what the producer called it. Entertainment value.
Did anyone ever call the station to complain?
Hilbert: The coach did. Called the news director and said he never said anything about the refs. The news director told him the caption was based on professional lip reading analysis and stood by it. I was the professional lip reading analysis. I was nineteen years old and I'd never taken a class in my life. I had a VCR and a notepad.
This is the perfect illustration of everything we've been saying. The skill is real, but the incentives corrupt it.
Hilbert: The skill is real if you're trying to help someone hear. If you're trying to make television, the skill is whatever the producer says it is. I still have the tape of that broadcast. I've been meaning to write to Coach Delvecchio for thirty years. Apologize. Explain that I was a kid and I needed the money and I didn't understand what I was doing.
Have you written it?
Hilbert: I've written it about forty times. Never sent it. He probably doesn't remember. Or maybe he does. Either way, the apology would be for me, not for him.
The thing that strikes me is that you were doing exactly what the viral video makers do now. Same incentive, same ambiguity, same fabrication. The technology changed but the dynamic didn't.
Hilbert: The dynamic is older than the VCR. Somebody wants a quote, somebody provides one. Whether it was said or not is a secondary concern. I can still lip read, by the way. But only if the person is saying cabbage and onion.
Why cabbage and onion?
Hilbert: Because that's the phrase I practiced with. My sister and I used to sit across from each other and mouth words. She'd do cabbage and onion and I'd guess. It was the only phrase I could get right every time. The bilabial sounds are easy. The rest of it is just guessing.
The countermeasure Daniel imagined, the nonsense phrase, is the one phrase you can actually read.
Hilbert: That's the joke, isn't it. The only thing I can reliably lip read is the thing designed to defeat lip reading.
There's a lesson in there somewhere about the arms race between concealment and detection.
Hilbert: The lesson is that concealment doesn't work. If someone wants to put words in your mouth, they'll put them there whether you said cabbage or onion or the entire periodic table. The lips are ambiguous enough to accommodate any transcript.
The only defense is to not have your lips filmed in the first place.
Hilbert: Or to not care what they write. Coach Delvecchio cared. I don't blame him. But caring didn't change what aired.
The perception of being decoded is the real weapon. The coach knew he'd been misquoted, but the viewers didn't. They saw a caption and assumed it was true. That's the chilling effect we were talking about.
Hilbert: The viewers probably forgot by the next commercial break. But the coach didn't. He had to go to work the next week and face his players and the other coaches and the parents. He knew what he said and what was aired, and he couldn't prove the difference.
That's the part that sticks with me. The asymmetry. The station could invent a quote and face no consequences. The coach had to live with the invented quote.
Hilbert: That's television. That's also the internet now. Same thing, faster.
The one thing I want to carry out of this is that lip reading is real but radically oversold. The viseme problem means the visual signal is always underspecified. The McGurk effect shows the brain is always filling gaps. And the professional applications, whether assistive or journalistic, are always shaped by incentives. My cousin uses speechreading to watch television with the sound low. The news station used it to invent quotes. Same skill, different purposes.
The cabbage and onion fantasy is really a fantasy about control. We want to believe we can defeat surveillance by being absurd. But the actual lesson from Hilbert's story is that the decoder writes the transcript. The lips don't get a say.
As video resolution improves and AI lip reading gets better, the accuracy problem will shrink but never disappear. The visemes are still ambiguous. The incentives will still push toward the juicy quote. The only thing that changes is how convincing the fabrication looks.
The open question is whether we'll get better at accepting uncertainty, or whether we'll just get better at manufacturing certainty. I suspect the latter.
This has been My Weird Prompts. Thank you to our producer Hilbert Flumingtop for keeping the show running.
If you enjoyed this episode, leave a review wherever you listen. It helps other people find the show.
We'll be back soon.