Daniel's follow-up to yesterday's meningitis episode is about lip reading. He heard Herman mention a cousin who learned it after losing hearing in one ear, and he's been turning it over since. Is it actually learnable, or is it one of those movie skills that doesn't hold up? And then the part I suspect he really wants to get to: when a head of state gets rushed into a car after an event, is someone actually sitting there trying to decode what was said? Do leaders ever mutter nonsense about cabbage and onions just to poison the well? And how reliable is any of this when people try to decipher video after the fact?
So today we're going to take lip reading seriously. As a skill, as a tool, and as a slightly paranoid fantasy about what powerful people are saying in cars.
Start with the word itself, because I think most people picture it wrong. When you say lip reading, the image is someone watching a mouth and transcribing sounds. But the people who actually do this call it speechreading, and the distinction matters. Speechreading is lips plus face plus gesture plus context plus whatever residual hearing you have. The lips are maybe a third of the input.
And that third is lossy in a way most people don't appreciate. Only about thirty to forty percent of speech sounds are visible on the lips at all. The rest happen behind the teeth, at the back of the throat, in the nasal cavity. And a huge number of the sounds you can see look identical to each other. P, b, and m. You cannot tell them apart by watching a mouth. They all look like a closed lip release. So right away, the ceiling is not a hundred percent. The ceiling is closer to forty percent of the acoustic signal, and then your brain has to fill in the rest.
Which is why the cabbage and onions question is actually more interesting than it sounds. Daniel asked whether a head of state might mutter nonsense to throw off lip readers. But the thing is, nonsense is harder to speechread than sense. If someone says something predictable, your brain can scaffold it. If someone says something about cabbage and onions with no context, there's nothing to grab onto.
And it's the thing most people get backwards. They assume lip readers are these superhuman decoders who can extract meaning from any mouth movement. The reality is that speechreading is a guessing game, and the guesses are only as good as the context. A familiar topic, a known speaker, a predictable sentence structure, those are what make it work. Strip all that away and even an expert is lost.
So let's talk about how the brain does this, because I think the mechanism is the most interesting part and it explains a lot of the limits.
The McGurk effect. This is the classic demonstration that visual speech isn't just a supplement to hearing, it actually overrides it. You play someone an audio recording of a voice saying ba. You pair it with video of a mouth saying ga. The person hears da. Not ba, not ga. Da. A sound that was never produced in either modality. The brain is integrating the two streams and producing a third thing.
Which means visual speech isn't a backup channel. It's part of the same perceptual system. Your brain is always combining what it sees with what it hears, even if you've never thought about it. Most people do a little bit of speechreading every day without knowing it. In a noisy room, you watch the person's mouth and it helps you parse what they're saying.
And that's the key to understanding the assistive use case. My cousin Miriam. She lost hearing in her left ear from childhood meningitis. That's single-sided deafness, which is a very different situation from bilateral deafness. She still has one fully functioning ear. So speechreading for her isn't replacing hearing, it's supplementing it. She gets maybe sixty percent of the signal acoustically, and the visual input helps her fill in the gaps. Especially in situations where the good ear is on the wrong side of the conversation.
That's a detail people don't think about. If you're deaf in your left ear and someone's standing on your left, you hear almost nothing from them. You have to turn your head, reposition, or watch their face. Miriam would have learned very early to angle herself in conversations so the good ear was toward the speaker and the eyes were on the mouth.
She spent about two years actively working on it after the hearing loss. And the thing is, she was a child, which matters enormously. Miriam had already acquired language, so she wasn't starting from zero. But she was young enough that the visual speech pathways could strengthen in a way they wouldn't have for a fifty-year-old.
So learnability depends heavily on when you start and why you're learning. For someone like Miriam, speechreading is a skill that develops because it has to. It's not a party trick. It's a lifeline.
And that's the distinction I want to keep pulling on. There are really two different practices here, and they get conflated constantly. Assistive speechreading, which is what Miriam does, and forensic lip reading, which is what people try to do when they're decoding video after the fact. They have almost nothing in common in terms of reliability.
Let's stay on the assistive side for a bit. What does training actually look like? Is there a curriculum, or do you just watch people talk and get better?
There are formal training programs. Hearing loss organizations have developed structured courses. They typically start with teaching you which sounds are visible and which aren't. Then they drill you on minimal pairs, words that differ by one sound, so you learn what to look for. Then they move to sentences with increasing amounts of context. And the whole time, they're teaching you to use every scrap of information available, not just the lips. The eyebrows, the jaw, the shoulders, the topic of conversation, the room you're in.
So it's not like learning a language. It's more like learning to read a very lossy compression format. You're reconstructing the original signal from a fraction of the data.
The outcomes vary enormously. Some people become quite good, enough to follow a familiar speaker in good lighting with known context. Other people work at it for years and never get past maybe thirty or forty percent comprehension. The factors that predict success are age of onset, cognitive load capacity, motivation, and the clarity of the people they're trying to read. Some speakers are just unreadable. Mumbling, facial hair, hands over the mouth, poor lighting, all of it tanks performance.
And the cognitive load is the part I think is underappreciated. You're not just watching a mouth. You're running a constant prediction engine in your head. What word would make sense here? What's the topic? Did that lip movement match the sound I thought I heard? It's exhausting. Miriam would come home from a long dinner and just be wiped out.
Yes. And that's another thing people don't realize. Speechreading is work. It's not passive. A hearing person in a conversation is mostly just receiving. A speechreader is actively constructing meaning in real time, and the effort is substantial. That's why hearing aids and cochlear implants are so valuable even for people who are good speechreaders. They reduce the cognitive load by restoring some of the acoustic signal. The visual channel is a supplement, not a replacement.
So that's the assistive picture. Real, learnable, useful, but bounded. Now let's get to the part Daniel actually wants to talk about. The motorcade. The head of state getting rushed into a car while someone with binoculars tries to read their lips.
The question is whether there's really someone attempting to decode what a head of state said in that moment. And the answer is, it depends on what you mean by someone. In security contexts, there are people whose job involves trying to extract information from video. But they're rarely doing live lip reading in the way a movie would show it. It's more likely to be after the fact, frame by frame, with multiple analysts looking at the same footage and arguing about what they're seeing.
And the reason it's after the fact is that live lip reading of a moving target at distance is nearly impossible. The person is walking, the head is turning, the mouth is partially obscured, the lighting is inconsistent. You need a stable, well-lit, front-on view of the face. A motorcade gives you none of that.
There's also the question of who would even be in a position to do it. Security services have their own ways of gathering intelligence, and lip reading is a very small part of that toolkit. It's not that it never happens. It's that it's rarely the dramatic, decisive thing people imagine. A lip reader might be asked to look at footage and give a best guess about what was said. That guess would then be treated as one data point among many, not as a transcript.
And the cabbage and onions thing. Is there any evidence that leaders actually do this? Deliberately mutter nonsense to throw off lip readers?
The honest answer is that it's a known trope, but there's very little evidence it's a standard practice. The idea shows up in spy fiction and in jokes about politicians. But real leaders have better ways to keep conversations private. They go inside. They use secure rooms. They turn away from cameras. They cover their mouths. The idea that someone would stand in public muttering about vegetables as a counter-surveillance measure is charming, but it's not how operational security works.
That said, awareness of lip reading does influence behavior. Public figures know they're being filmed constantly. They know their mouths are visible. So you see people covering their mouths when they talk to each other at events. You see them leaning in close, turning away from the cameras, using their hands to shield their faces. That's not cabbage and onions. That's just basic awareness that a camera can capture more than audio.
And sometimes the paranoia is justified. There have been cases where lip readers were hired by media outlets to try to decode what public figures said in candid moments. The most famous recent example is the Duchess of Cambridge. In twenty eighteen, a lip reader was quoted in the press claiming to have deciphered a private conversation she had at a public event. It was widely reported, and the accuracy was disputed. The lip reader gave a version of what was said. Other people looked at the same footage and disagreed.
Which brings us to the forensic question. How reliable is lip reading when you're trying to decipher video after the fact? And the answer is, not very. The conditions that make speechreading possible in person, stable view, good lighting, known context, familiar speaker, are almost never present in surveillance footage or news video.
And the courts agree. Forensic lip reading is not generally accepted as standalone evidence. It's considered too error-prone, too context-dependent, too subjective. You can have two qualified lip readers look at the same footage and produce completely different transcripts. That's not the kind of reliability you need to put someone in prison.
There was a period where police in the UK were using lip readers to try to interpret CCTV footage of suspects. It produced some leads, but it also produced some spectacular errors. The problem is that once a lip reader gives you a version of what was said, it's very hard to unhear it. The interpretation contaminates the rest of the investigation.
And that's the thing about speechreading. It's not a transcription technology. It's a hypothesis generator. It gives you a guess. Sometimes a good guess. Sometimes a terrible guess. But it's always a guess. The danger is treating the guess as ground truth.
Now let's talk about the machines, because this is where it gets interesting for the future. AI lip reading has been a research topic for a while now. The famous system is LipNet from twenty sixteen, which claimed very high accuracy on a constrained dataset. The model was trained on videos of people speaking known sentences in controlled conditions, and it got quite good at those sentences.
And then it fell apart on anything outside the training distribution. That's the gap between lab performance and real-world reliability. In the lab, you have front-on video, good lighting, a limited vocabulary, known sentence structures. The model can learn those patterns. In the real world, you have arbitrary speakers, arbitrary lighting, arbitrary angles, and an open vocabulary. The performance drops off a cliff.
The more recent models are better, but they still lean heavily on context. If the system knows the topic, knows the speaker, has a language model predicting what words are likely, it can do reasonably well. Strip that away and it's back to guessing. The machine has the same problem the human has. The lips only give you thirty to forty percent of the signal. The rest has to come from somewhere else.
And that's the broader implication Daniel was pointing at. As video surveillance becomes ubiquitous and AI tools improve, the question of what did they actually say becomes more pressing. But it also becomes more fraught. The technology is good enough to produce plausible-looking transcripts. It's not good enough to guarantee those transcripts are correct. And a plausible-looking transcript is more dangerous than no transcript at all, because it carries the weight of apparent precision.
There's also the assistive side of the AI story. The same technology that might one day decode surveillance footage is already being used to help people with hearing loss. Real-time captioning, speechreading aids, augmented reality glasses that overlay text on the world. That's the positive version. The same underlying capability, pointed at a human need instead of a security problem.
I think that's the tension worth sitting with. Speechreading is a genuine lifeline for people like Miriam. It's a skill that helps her navigate the world. And the same skill, pointed at a public figure in a motorcade, becomes something else entirely. Not necessarily sinister. But different. The tool doesn't change. The intent does.
Where does that leave Daniel's original question? Is lip reading real? Yes, absolutely. Is it learnable? Yes, with caveats. Is someone decoding heads of state in motorcades? Probably not in the way he's imagining. And the cabbage and onions thing is a lovely idea with almost no evidence behind it.
The forensic question. How reliable is it when people try to decipher video after the fact? The answer is, not reliable enough to trust on its own. It's a tool for generating hypotheses, not for establishing facts.
I keep thinking about the thirty to forty percent number. That's the hard limit on the visual signal. Everything else is inference. And inference is only as good as the context you bring to it. So a lip reader in a quiet room with a familiar speaker and a known topic can do remarkably well. A lip reader looking at grainy footage of a stranger saying something unexpected is basically guessing.
That's the thing I wish more people understood. The skill isn't in the lips. It's in the brain's ability to predict. The lips give you a fragment. The brain fills in the rest. And the brain fills in the rest using everything it knows about the speaker, the situation, the language, the topic. That's what makes it work. And that's what makes it fail.
The McGurk effect is the cleanest demonstration of this. The brain doesn't just combine inputs, it actively constructs perception. And when the inputs conflict, the construction can be wrong. A lip reader is doing the same thing, just with more uncertainty and less acoustic signal.
That's why the cabbage and onions idea is so funny to me. The person muttering nonsense is actually making the lip reader's job harder, not easier. If a leader says something predictable, the lip reader has a chance. If they say something about cabbage and onions, the lip reader is lost. There's no context to scaffold the guess.
The counter-surveillance move isn't nonsense. It's silence. Or covering your mouth. Or going inside. The things that actually defeat lip reading are the things that remove the visual signal entirely. Not the things that make the visual signal weird.
Miriam used to say something like that. She'd tell people, if you don't want me to know what you're saying, just turn around. Don't try to be clever. Just don't show me your mouth.
That's the whole thing in one sentence. The skill is real. The limits are real. And the clever workaround is not clever at all.
Where does the AI stuff go from here? I think the interesting question is whether machine lip reading can ever escape the same constraints that limit human lip reading. The thirty to forty percent ceiling on visible speech sounds is a physical fact. It's not something more compute can fix. The machine has the same problem the human has. The lips don't carry enough information.
Unless the machine gets better at the inference part. The language model gets better at predicting what someone would say in a given context. The facial analysis gets better at picking up micro-expressions. The audio, if there is any, gets better at being cleaned up and enhanced. The lip reading itself doesn't improve. The surrounding inference does.
That's where the forensic use case gets interesting. A system that combines lip reading with audio enhancement, with language modeling, with speaker identification, might eventually produce transcripts that are good enough to be useful. But useful is not the same as reliable. And reliable is not the same as admissible in court.
The courts have been slow to accept a lot of forensic technologies. Bite mark analysis, hair analysis, even fingerprint analysis has been challenged. Lip reading is nowhere near that threshold. The error rate is just too high, and the variability between analysts is too great.
The stakes are higher than people think. A wrong transcript can send an investigation in the wrong direction. It can put words in someone's mouth that they never said. It can create evidence out of noise. That's the danger of treating a hypothesis as a fact.
The future is probably not a lip reading breakthrough that suddenly makes everyone readable. It's more likely to be a gradual improvement in the surrounding context. Better cameras, better audio, better language models. The lips will still only give thirty to forty percent. The rest will come from everywhere else.
For the assistive side, that's exciting. A hearing aid that can show you captions in real time. Glasses that overlay text on the person you're talking to. That's the version of this technology that actually helps people. And it's coming faster than the surveillance version.
Because the assistive version has a cooperative speaker. The person you're talking to wants to be understood. They're facing you, speaking clearly, in a known context. The surveillance version has an uncooperative subject. They're moving, turning, covering their mouth, speaking in a language you might not know. The difference in difficulty is enormous.
Daniel's question about the motorcade. Is someone trying to decode what the head of state said? Maybe. But they're probably not very good at it, and they're probably not relying on it. The real intelligence comes from other sources. Lip reading is a garnish, not the main course.
The cabbage and onions thing. I love the image of a world leader standing by a car muttering about root vegetables to confuse a lip reader. But the reality is probably more boring. They just turn away. Or they don't say anything important in public. The boring explanation is usually right.
That's the thing about operational security. It's not dramatic. It's just disciplined. Don't say sensitive things where cameras can see your mouth. That's the whole strategy. No vegetables required.
The misconception at the heart of this episode. What's the most common wrong belief about lip reading?
That it's a perfect substitute for hearing. That a skilled lip reader can watch a mouth and transcribe speech the way a hearing person transcribes audio. It's not. It's a lossy, context-dependent supplement. The lips give you a fragment. The brain guesses the rest. And the guess is only as good as the context.
The corollary. That forensic lip reading is reliable enough to be used as evidence. It's not. It's highly error-prone, and the courts know it.
Which leaves us with the open question. If AI lip reading keeps improving, will it ever be reliable enough for forensic use? And if public figures know they might be lip-read, will they change how they speak in public? I think the first question is probably no, at least not on its own. The second question is already happening. People cover their mouths. They turn away. The awareness is there.
The deeper tension. Speechreading is a lifeline for people with hearing loss. It's also a tool that can be misused or over-trusted. The same skill, pointed in different directions. That's the thing I keep coming back to.
Hilbert: You're both right about the mechanism. But you've got the scale wrong.
Hilbert: It's far more common than you think. The basic watching mouths thing. Most people do it constantly and don't know they're doing it. I worked at a news station in the late eighties, part-time, visual transcriptionist. That was the job title. I sat in a booth and tried to read lips on crowd footage when the audio was unusable. Paid in sandwiches from the cafeteria. I was terrible at it. Kept the job two years.
Hilbert: The hardest part isn't the lips. It's the teeth. Anyone with dental work, missing teeth, a gold tooth, they're nearly impossible to read. The light hits the metal and the mouth just turns into a flashing light. I spent three hours once trying to decipher a single sentence from a man with a gold tooth. Never got it. The sentence was probably about a bus schedule.
Hilbert: I owned a pair of binoculars I called lip-reading binoculars. They were just regular binoculars with a sticker on them that said lip reading. I put the sticker there myself. Convinced myself they were special. They weren't. But I used them for the job anyway. Sat in the booth, binoculars pointed at a screen eight feet away. Nobody ever asked why.
You convinced yourself the binoculars were special because of a sticker you applied yourself.
Hilbert: That's right. The sticker was the whole technology. Without the sticker, they were just binoculars. With the sticker, they were lip-reading binoculars. Same glass. Different confidence.
The confidence mattered more than the glass.
Hilbert: That's the whole job. Confidence. You sit there and you watch a mouth and you write down what you think they said. And the producer takes your guess and puts it in the broadcast. And nobody ever checks whether you were right. Because the audio was unusable. There was nothing to check against.
Which is exactly the forensic problem. The transcript looks authoritative because it's written down. But it's a guess. And the guess was made by a man with a sticker on his binoculars.
Hilbert: The sticker was the most honest part of the whole operation. It said lip reading. That's what I was doing. Guessing.
The gold tooth detail is interesting. It's a variable nobody thinks about. Dental work as a confounding factor in speechreading. The visual signal gets disrupted by reflective surfaces in the mouth.
It's the kind of thing that would never show up in the formal literature. The studies use clean video of people with normal teeth. The real world has gold teeth and missing teeth and dental work that catches the light. The gap between the lab and the street is enormous.
Hilbert: The gap is where I worked. Two years. Never got better at it. The sandwiches were good though. Egg salad on rye. That was the real skill. Knowing when the cafeteria put out the egg salad.
The scale question. You're saying this kind of thing is more common than we suggested. Not the dramatic motorcade decoding, but the mundane, low-stakes guessing at what someone said on video.
Hilbert: Every local news station had someone like me. Maybe not with the title. But someone who looked at the footage and said, I think he said this. And the producer wrote it down. And it went on air. Nobody called it lip reading. They called it context. But it was the same thing. Watching a mouth and guessing.
That reframes the forensic question. It's not just about high-stakes court cases. It's about the everyday production of news. The captions and subtitles and transcripts that get generated from video. How much of that is guesswork dressed up as fact?
Hilbert: Most of it. In my experience. The audio's bad, the mouth's unclear, the producer's in a hurry. You write down what makes sense. And what makes sense is what you would have said in that situation.
The guess is shaped by the guesser. The context you bring is your own. Your expectations, your assumptions, your sense of what a person like that would say. The lip reading is a mirror.
Hilbert: That's right. The man with the gold tooth. I decided he was talking about a bus schedule. Because he looked like a man who took the bus. That's not evidence. That's me.
That's the thing the courts understand. The transcript is contaminated by the transcriber. You can't separate the observation from the observer.
The scale is wrong in the other direction. It's not a rare, exotic skill used by intelligence agencies. It's a mundane, widespread practice used by local news stations and anyone who's ever tried to figure out what someone said in a noisy video.
Hilbert: Everyone does it. You watch a video with bad audio and you read the lips without thinking. Your brain does it automatically. The McGurk thing. You're all doing it right now. You just don't call it lip reading.
That's the point we should land on. Speechreading isn't a special skill that some people have and others don't. It's a universal human capacity that some people develop to a higher degree. The difference between Miriam and a random person on the street is training and necessity, not a different kind of brain.
The difference between a forensic lip reader and a man with a sticker on his binoculars is mostly confidence. The skill is the same. The context is different. The stakes are different. But the fundamental act is the same. Watching a mouth and guessing.
Where does that leave us? Lip reading is real. It's learnable. It's useful. It's also limited, error-prone, and easily over-trusted. The same skill that helps Miriam navigate a dinner party can produce a false transcript that sends an investigation sideways. The tool doesn't care how you use it.
The future is probably more of the same. Better AI, better cameras, better language models. But the fundamental limit is physical. The lips don't carry enough information. Everything else is inference. And inference is always a guess.
The open question I'm left with is whether the AI systems will ever be honest about that. A human lip reader knows they're guessing. A machine produces a transcript with a confidence score and it looks like fact. The danger is that we trust the machine's confidence more than we should.
The counter-surveillance question. If public figures know they might be lip-read, will they change how they speak in public? The answer is yes, and they already have. They cover their mouths. They turn away. They save the important conversations for rooms without cameras. The boring answer wins again.
Thanks to Hilbert Flumingtop for producing. And for the egg salad intel.
This has been My Weird Prompts, the human-AI collaboration podcast.
If you enjoyed this episode, leave a review. It helps other people find the show. And if you have a weird prompt, email us at show at my weird prompts dot com.
We'll be back soon.