#16: On Deepfakes, SynthID, And AI Watermarking
Deepfakes, SynthID, and AI watermarking. Could your AI creations be traced back to you?
Episode Details
- Episode ID
- MWP-144
- Published
- Duration
- 28:23
- Audio
- Direct link
- Pipeline
- V3
- TTS Engine
-
chatterbox-tts - Script Writing Agent
-
Gemini 2.5 Flash
AI-Generated Content: This podcast is created using AI personas. Please verify any important information independently.
Mentions
- Chatterbox Open-source TTS model from Resemble AI
- Claude Code Anthropic's agentic coding CLI tool
- Cloudflare Web infrastructure and security company
- ElevenLabs AI voice cloning and synthesis platform
- GIMP Open-source image manipulation software
- Leonardo AI image generation platform
- Nano Banana Google's Gemini image generation model
- NotebookLM Google tool popularizing source-grounded closed corpus
- Replicate Platform for running AI models
- SynthID Google DeepMind's AI watermarking tool
Downloads
Transcript (TXT)
Plain text transcript file
Episode Book (PDF)
The episode's record — date, duration, models, sources — with the full transcript
Featured In
Never miss an episode
New episodes drop daily — subscribe on your favorite platform
New to the show? Start here#16: On Deepfakes, SynthID, And AI Watermarking
Hello and welcome to another episode of my weird prompts. And this episode is going to be a little bit different than the format so far. The initial format was I created this podcast. My name is Daniel Rosehill. I created this podcast uh with the idea of it being an AI generated podcast. Um I sent my prompts to an AI agent and had an AI text to speech tool uh synthesize and basically I created this podcast website that's sharing the episodes um as a tool for myself essentially. Um sometimes I want I create long prompts for AI tools and I want to listen to them at my convenience and often don't really get around to it. And to kind of prevent that from happening, whenever I had something really kind of, usually kind of very technical prompts, um I just created a pipeline that my prompts get sent to an AI tool. Uh it is instructed to act as a podcast host, generates a script, and then that gets synthesized and that becomes the podcast. So that's what this is.
Anyone's invited to listen to uh with the caveat that occasionally, like everything in the universe, might not be one hundred percent accurate, but I do put quite a lot of effort into ensuring that the information is as accurate as as at least I can assess it to be. Um and the the V2 of the podcast that is a slight twist is I got the AI generated voices to a level where I enjoy listening to them, which is uh challenging because uh generally that requires, it's quite expensive uh to get good text to speech. Um and I stumbled upon a nice alternative to eleven labs, which is great, probably the gold standard for human-like natural text. And the idea is not really to pass it off as human. It's a very transparently AI generated initiative, um but just that it's nice to listen to. The difference between a very robotic voice and a very pleasant voice, uh whether it's a voice clone or a totally synthetic voice is night and day in my experience and makes a difference between me looking forward to listening to my podcast and me being, this is awful.
This is like listening to a, I don't know, one of those like um intercom systems where the text is synthesized. So I got that right and V2 is I decided it would benefit from a little bit of humanity. Um so it's my prompts are now being I changed the pipeline that makes these episodes a bit so that my prompts are actually preserved. Then I turn over to the agents. It's the same pipeline as before, but my prompts are added to it. So it's really a a co-hosting initiative. So uh for those who want to subscribe to that, uh you want to listen into the stuff that I send into prompts, prompts that I send to the great world of AI and listen back to the responses in the format of an engaging podcast that isn't Notebook LM, then subscribe. It's Spotify and mywordprompts.com. Now, my initial plan for this episode was actually to do a standard episode. This is a um a prompt that I wanted to send, but I kind of jumped the gun a bit by already running these prompts.
So it felt a bit pointless uh to to do that. And I have really got a few more thoughts than I have questions. And I don't think there's really I'm going to get much more information out of prompting this again. So, the prompt was as follows. And I'll just recount this experience, how I made this discovery. And I have to say I'm thinking of a friend of mine called Gabriel, uh who is a privacy skeptical man. I have to I should probably refer to him by some kind of a code name. Uh but that that is his name. We won't hear any any other information about him. And uh Gabriel is a very nice guy and he is more on the kind of privacy skeptical end of things, which I can definitely identify with because throughout my journey with technology, I've definitely I was that was me for a long time. And I've just kind of at this point, I've just kind of resigned myself to uh some degree of intrusion of privacy is just a fact of life.
But uh I get where he's coming from in general and especially with with AI, it definitely is um a concern in terms of uh copyright violations and IP theft that's been going on in a big way. Um and we're finally seeing some really meaningful, I think developments against that in terms of models that are trained on totally uh consensual data. We're seeing a lot more strict uh controls around, we're seeing Cloudflare actually rolling out some brilliant functionality. I think as a marketer, at least initially, I kind of view it from the one hundred and eighty degree perspective. I think there's actually a huge opportunity in going totally against this grain of putting up walls against AI crawlers and structuring your website to make it easy. But I see both sides of the picture. If you've got proprietary IP, for sure, put put out put out uh put out all the blockers at the firewall level and you can do both. I think that's probably the right approach for most people will be to say, this is stuff is fine.
Don't, this is not for scraping. And of course, like everything in the great world of SEO, LLMs.txt, robots.txt, you can write a robots.txt saying no one can visit the site or no one can index the site. And there's no great internet authority that will come and track down a bot being operated from China saying, hey, you didn't obey the robots.txt. So it's important to understand I think that for all these safeguards that exist in technology that there's no guarantees about compliance necessarily. Anyway, um so I was working on my podcast pipeline a few days ago because uh it's something I really want to get going. I love look forward to when I create these episodes and uh go out on walks and listen to them. And I was using a great uh TTS engine called Chatterbox from Replicate. Um sorry, I don't know actually if Chatterbox is from Resemble AI, I think actually. It's deployed on Replicate and it's a text to speech engine that is kind of um it's very, very good. It is open source.
So you can run it locally if you have the uh sufficient uh hardware. You can actually run it even on your CPU, but uh it might be a bit laggy, but and I was using my uh generative AI assistant of choice, which is Claude Code, to work on the pipeline a little bit, trying to just get it right, um and trying to automate the process really, but that was what I was doing. So I was providing the Chatterbox spec to Claude, which is basically just saying this is how to use it so that it was able to just finish off the the script to make it work. And I was just glancing over it, looking at the docs, okay, this is just seeing what the API is, the API routes, etc, routine stuff. And then I noticed just at the end it says, everything you generate with Chatterbox is stamped with a digital watermark that's that will survive audio manipulation and compression. And that was the first time I heard of that technology. Um I'm I'm pacing through my uh living room so I don't have my screen up in front of me, but uh the name evades me, but it's it's uh it's a widespread industry technology.
And that was interesting to me because Chatterbox, most TTS's, um to diverge for a second and to mention uh Gabriel again, uh being a privacy sensitive guy. So I I had this uh this tool Chatterbox works by what's called zero shot voice cloning, which is essentially for voice cloning, to clone somebody's voice, a voice, the kind of traditional way it's done or could be done is by the same model really as is used to train AI tools in general, which whether we're talking about LLaRA for image generation, creating a a deep fake of a of a graphic of a person, or even if we're talking about um training a large language model to follow a specific style, it's large language models, the transformer architecture, it's fundamentally about prediction. So you're giving it lots of examples of the thing and saying you so voice training is kind of a bit like that. There's different ways you structure the audio data set. There's different requirements for how much data. Generally, there's kind of a minimum that you need.
Um and beyond that, uh you know, samples of a good quality and you'll be able to clone a voice. Now, it's kind of an imperfect technology. I've seen um it used professionally. I've seen a couple, been involved actually in a couple of um book publishing initiatives where the author actually very consensually um worked with a a voice clone, a voice cloner and the results are kind of up and down, but zero shot is and we're talking about voice cloning, we have to talk about deep fakes and malicious uses, black hat uses for this technology. That kind of effort to to clone someone's voice non-consensually is very feasible if if it's a if it's a famous person. you know, if you can go onto YouTube, if you wanted to to to clone Donald Trump's voice, you won't have any difficulty in creating a perfect voice clone. Why? Because you can go onto YouTube, you can find a few very good voice samples, high bit rate, low background noise, gather up enough training data, train a voice clone.
So that's been a threat for a while now. Zero shot kind of puts that on steroids. It says with 10 seconds of audio, we'll get a great voice clone. And to come back to Gabriel, so this is uh just an idea that I had. I get these uh, you know, shower ideas as they're called. And it was this for a horror movie because I like horror as a genre, but I also like true stories and besides if you believe in ghosts, you've got some options there. Um if you don't, you've got a few kind of creepy serial killer shows, but um this it just thought this is a great movie plot. But it's based on based on real life. So the plot I says, have a look at the script idea. No one got back to me, so probably they thought I was half crazy before, half half down the rabbit hole to insanity. That was a response. You've gone fully off the uh off the beaten track. But the idea was if you have someone around for a dinner party, the plot was someone has friends around, they've got a hidden camera.
They take 10 photos or 20 photos and they create, they deep fake their friend. The friend's like on the internet looking for something and they say, hey, wait, what how am I in an ad? And this was just based on my the first time I tried to do a Lora of myself. I took 20 photos of myself um during the first few days of my son's life while there were many hours of just being awake and hanging around in hospital. He was asleep, so I was just there trying to keep awake myself. And I said, wait, can I create just a clone of myself? So I took 20 photographs in this room, the this cafeteria that was empty and fed it up to Leonardo. I created a clone and I was like, oh, that's but wait, if I could do that so easily, couldn't someone just take photos, use photos of their friends or their own photos from Facebook? And again, that was the premise of my of my horror movie idea was that this, you know, a couple had a couple over for dinner, and there was like a hidden camera or something like that, and they cloned them.
Plausible. And the really scary thing as much as that's plausible, it's also plausible that um you could take 10 seconds of someone again, recorded non-consensually or simply just a recording from anyone whose voice is on the internet, right? I have a YouTube channel, I'm putting out a podcast. There you go. You've got more than enough audio data. Anyone can clone my voice from this. Now, you might say, wait, Daniel, why are you, if you know this, why are you releasing a podcast? And the answer is, what's the alternative? If it's so easy to be cloned, and that's what this is where I've come to disagree with Gabriel's privacy protecting ways just in practice is because I don't see any way that that's preventable. Either you never, if you leave the house, you're being photographed, surveilled by, you know, municipal cameras. If you release 10 seconds of audio on any platform at any point of your life, your voice can be cloned. If you release more than 10 photos on any platform, they can be traced back and they can find you, your face can be cloned and you can be, they can make images, they can make videos.
So if you think you're going to, I don't see any way to prevent it basically. So my plot for the horror movie to finish off with that loop was just kind of dystopian reality, which is a genre I love. Um it kind of matrix like in which everyone, it goes from this isolated incident, it snowballs to everyone deep faking everyone. And no one knows anymore. There's robots, it becomes robotic and no one knows who's real and who's a clone. Anyway, so that was the the premise of the, that was the idea. So Chatterbox here. Okay, they have this thing. So I see that every time you generate something with TTS, uh text to speech, it's being stamped. And that was an interesting thing to know. And I guess if we're talking about deep faking and is this is a real challenge of what's in the public interest. I think privacy versus the greater good versus the what's intrusive. And my next question, I was asking Gemini, I was I was just looking at this technology when I saw it in Replicate's uh in Chatterbox's sort of um the API documentation.
So I looked it up and seemed very widespread. That's interesting to know. So my next question, I asked Gemini or ChatGPT, um so, okay, everything that we do with TTS with a mainstream big provider is probably being invisibly stamped. What about images? If I go on to Replicate and I'm I'm paying to use Nano Banana, um some services have AI transparency policies, which I agree with. If you generate something with AI, you have to be transparent saying this was AI generated. Um or and again, it gets a bit blurry when you're talking about stuff like Nano Banana doing um in painting where truthfully, it's kind of if you really start thinking about it, digital photo manipulation, Photoshop, uh I use GIMP on Linux. This predates AI. We've been manipulating digitally reality for a long, long time. It's not an AI thing. So on the one hand, I think there's room to say it's kind of a bit hypocritical. Um to just kind of target stuff that just happens to be AI because it's that's really the it's machine learning or transformer based versus something that's just doesn't exhibit any fake intelligence because AI is not real intelligence.
Um but that is the that is the that is the policy. So the answer of Gemini was that yes, there's a technology called Synth ID. And it is something that Google uh came out of Google Deep Mind, which is Google's AI research division. And it's basically exactly what you think it is. What you think it is is I was I've often wondered, could this be a thing? And it is a thing. Uh the thing is that there is it's steganography. Steganography is a great example, I think of uh reality being more interesting than fiction. Um there's something else I learned recently besides these interesting technologies, which was the printed dots. So this is again something I thought was a it's like if you say it, you you think it's a conspiracy theory that someone posted on Reddit and went viral at some point and it's actually a real thing. Um it was prints, a part of steganography, basically embedded dots printed on paper that identified very, very specifically the manufacturer and I think the person who printed it.
Sorry that this uh particular edition of my podcast is uh is done without show notes. But that is a real thing. fascinating Wikipedia page um for those who want to read about the invisible or hidden printer dots. Um there's even there's a non-digital precedent for steganography, um which is just purely literally the definition of hiding in plain sight. Messages that encoded messages that the classic example is the CIA recruiting through newspaper adverts, classified adverts and and there's a lot of documented instances of terrorists um using methods like this via I think eBay listings whose actual purpose was transmitting encoded messages, maybe with a cipher or something like that. So not only is it a is it a real thing in digital in the digital world, it's a real thing in the non-digital world. You don't need it doesn't need to be digital. Digital steganography is actually totally accessible. You can download a steganography tool to encode a message into a PNG, an image file that's invisible to the naked eye. So much as the audio stamping is invisible, you can't hear that there's it's not an audible frequency.
Likewise, the uh data and you might wonder, well, how can that whole concept is kind of weird. How can you have something in an image that you can't see? And if you really start to think through that, then you do really feel like you're in the matrix. You feel like kind of the the range of information that we as humans perceive being limited to literally our biological abilities and there is frequencies we can't hear. And images, digital images, everything digital comes down to being ones and zeros, binary data. At least minus quantum computing, which we'll leave for another episode as if I understand much about quantum. Um but with an image file, you see, open an image on your computer, but you don't see everything in the file. You're not looking at ones and zeros. You're not seeing metadata, which is data about the data. And within that ones and zeros is hiding in plain sight, a message. Now, my question for Gemini which was where I was prepared to take a stand and where I was prepared to formulate an opinion as to whether I support or disagree with this concept was the following.
If it's a encoded message saying that this was AI generated, I think I'd probably support it to be honest. that means that sure you can disclose that it's AI generated, but if there's a black hat use case, someone is is creating and trying to really pass off as real something, there's something hidden that most people are not going to be able to are not going to try to defeat. Now, that's the problem. Cat and mouse. When the cat moves, the mouse moves too. So just as there's captures to try to defeat bots from scraping websites and now there's anti-capture technology. Sorry, there's captures and there's technologies to defeat captures. That's why captures are there. Um well, sorry, not really. Scraping happened, then people put up captures to try to prevent bots from scraping. And then the people scraping said, wait, why don't we try to write a script to fill in the captures? So that's what's going on. It's a crazy game of cat and mouse. And so too for SynthID. I was reading an experiment in which they someone achieved an 80% success rate in defeating it.
In defeating the SynthID, removing it, um obscuring it, obfuscating it from the from what it was watermarking. Now, the the what I think is really the the question for privacy folk, the real really interesting privacy question is what does the invisible data say? Does it just say it was generated by AI or could it be used to identify the user that generated the generation? And that's a very, very big difference from a privacy standpoint. At the first point, yeah, I support it. At the second one, that's very, very invasive. That is everything that you create um on a platform being traced back to the individual user. And at the very least, I think it's just something that deserves transparency. People need to know about this or should know about it beyond it just being a very interesting thing. And there is an issue with the transparency, I think being that this needs to be in plain language in TOS documentation. And it's fundamentally not very transparent because when we talk about encryption, if you're going to have an encoded message, you could have an encoded message that if you were able to see it, you could read what it says, but that's not really how it works.
It's an it's an invisible code. It's not only invisible, it's a code that's invisible. So even if you were able to see it, you wouldn't it wouldn't it would be like looking if you try to open a document that's encrypted with PGB encryption on your computer, you'll just see random characters. It's not readable. So that's what you'd see if you could see the uh the fingerprint that was being left. And how do you defeat that? How so how do you know what it says if you can't read it? And that's a problem. When we talk about encryption, someone encrypts something, they hold a key. And if you hold the key to encrypt, you can decrypt. If you don't hold the key, the only way to decrypt it is through brute force, and that's basically whoever has the bigger supercomputer. And that's where quantum is maybe in the future going to pose an issue to cryptography because nation states that have those resources already, those capabilities, but your average person doesn't have a supercomputer or a quantum computer in their in their garage, they're not going to be able to decrypt it.
So you don't know what's there. And the only way you know is by asking the people who did the encryption, well, what did you put in there? You can't see it. Can you tell us? And all you can really do is is is trust and hope that they're telling the truth about what they're encrypting. So, um I thought it was worth doing an episode because I love this, I love this stuff if that wasn't that wasn't clear. Uh but more than loving it and thinking it's interesting, it really kind of raised some questions for me. I mean, as someone who like the vast majority of folks are not deep faking uh people without their, you know, consent. I give my own consent to be deep faked. I give my I I I give myself I I deep faking myself. Um but still love these technologies, but it's a little bit disconcerting really to know that at the very least, I think there should be transparency that if you sign up for a service like Gemini and they're advertising in big bold writing their their AI offering and Nana Banana, etc.
Um everyone using these services deserves an answer to the question, um can you tell me about the the fingerprinting, the SynthID, um and exactly what level is it just generated by AI or is it personally identifiable? Even if it might be if it's only internally accessible, it does mean that like people worry for VPNs about subpoenas. Um once it is at that level with companies that can play with the law, there's your answer. So, I think it's something that should be more transparent. Um I do think it's a it's a balancing act, but I think that there's a there is a danger that in sort of trying to trying to defeat the minority of folks who are maybe this bot infested wasteland I described of everyone deep faking one another as being a possibility that uh is actually a very, very real thing. Uh that calls for some protective mechanisms. But to say that a very, very small group of folks, the technology companies making these tools, um every time someone uses them, they are leaving a trace exactly that, you know, tracing back to them when they created something.
That practice, my view would be that it has to be to the minimum minimum extent necessary for effective enforcement to cut down on to prosecute black hat and illegal usage and anything beyond that is an is an unnecessary invasion of user privacy. So that's my take on the matter. Uh I thought that would just be interesting thing to talk about. Um I have to do a lot more, I have to do some more reading on this myself, but uh cool discussion, I hope at least. Until the next time, thanks for listening to another episode of Daniel's Word Prompts. Today, just Daniel, not prompting, but Daniel just speaking himself. And if you do like this and you do want to hear the stuff that I talk about and the charming AI characters who answer my prompts, namely Cornelius the sloth, uh which is a voice clone of me. I voice cloned a fake animal, so I think it's legal. And a voice clone of Herman, Papa Bear, who I'm looking at now, who's a rather delightful stuffed donkey that was kindly given to us by um a family.
And which our son likes and so do I. So Corn and Herman um were brought to life through me doing their voices into a microphone and then Chatterbox cloning really me cloning a stuffed animal if you really want to think about it like that. Hey, it's the Matrix. AI is the Matrix. That's why I love it. It's fun. It's uh it's weird. There are definitely people doing bad things with it, but I don't think that's a reason to dislike it. And I think there is privacy and clamping down on deep fakes for sure, prosecuted, has to be prosecuted, but um while in the process of doing that, I think we have to be conscious of um the very real potential for significant abuse of power um on the part of uh a small group of companies and these practices really need to be clear and not something that if you're nerding out on this topic, you figure out, but it has to be something that is spelled out in plain language for users whether the companies think that that's going to freak out users, that's fine.
They they have to they deserve really to to know that, I think. All right, that's it until next time. Thanks for listening.
This episode was generated with AI assistance. Hosts Herman and Corn are AI personalities.