From deepfakes to DNA: the science of watermarking AI
Google DeepMind · 2026-10-01 · official · 4,046 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
In this episode of Google DeepMind: The Podcast, host Professor Hannah Fry interviews Pushmeet Kohli (VP of Science Research at Google DeepMind) and Jeremy Ratcliff (Research Scientist at Google DeepMind). They discuss the principles and implementation of AI watermarking across media (text, images, video, audio) through SynthID, as well as its expansion into biology with SynthID Bio to watermark AI-designed protein sequences and 3D structures for biosecurity and provenance.
What is shown
- [00:00] Hannah Fry introduces the challenge of identifying synthetic content and dangerous AI-designed biological molecules.
- [01:11] In-studio interview with Pushmeet Kohli and Jeremy Ratcliff discussing the definition and motivation of watermarking for provenance.
- [02:18] Overview of essential watermarking criteria: human imperceptibility, high detectability, and cross-system generalizability.
- [03:40] Pushmeet outlines the three design pillars of SynthID: quality preservation, robustness against adversarial removal/transformations, and scalability/ease of integration.
- [11:38] Technical explanation of text watermarking: leveraging token selection optionality to bias probability distributions according to a pseudorandom secret key.
- [15:24] Explanation of image watermarking architecture: co-training an encoder neural network and a detector neural network against an adversarial agent that applies cropping, rotation, scaling, and noise.
- [18:17] Discussion of consumer-facing verification features integrated into the Gemini app and developer portals.
- [19:53] Deep dive into SynthID Bio:
- SynthID Bio Structure: Watermarking predicted 3D atomic coordinates (building upon AlphaFold 3) by slightly perturbing atom positions and angles.
- SynthID Bio Sequence: Watermarking amino acid sequences via generative models like ProteinMPNN by substituting functionally equivalent amino acids.
- [22:36] Explanation of biosecurity screening mechanisms at commercial DNA synthesis providers, and how generative models could bypass sequence-matching databases without provenance tracking.
- [27:08] Details on physical wet-lab validation testing watermarked protein binders against unwatermarked designs to evaluate hit rates and binding affinities.
Claims & numbers
- Pushmeet Kohli states that DeepMind's watermarking initiative began nearly eight years ago.
- Pushmeet states a standard 1-megapixel image contains approximately 1 million to 3 million numerical values depending on encoding, providing substantial bandwidth to hide imperceptible signals without degrading quality.
- Pushmeet states that text watermarking depends on token entropy/optionality; very short or deterministic snippets (such as "What is the capital of France? Paris") cannot be watermarked without altering meaning.
- Pushmeet notes that SynthID technology is used by industry partners including NVIDIA and OpenAI, and detection capabilities are deployed directly within the Gemini application.
- Jeremy Ratcliff describes two distinct biological watermarking modalities: SynthID Bio Structure (altering 3D atom positions/angles) and SynthID Bio Sequence (selecting alternative amino acid residues).
- Jeremy claims wet-lab experiments physically synthesized and tested watermarked versus unwatermarked protein binders, demonstrating near-identical binding hit rates and quantitative binding metrics in vitro.
- Jeremy states Google DeepMind is open-sourcing the SynthID Bio technology and code to enable community adoption and integration with DNA synthesis screening providers.
Notable quotes
- [02:21] Jeremy Ratcliff: "I think for us, in particular, human imperceptibility is one... having high detectability... and then also good generalizability."
- [12:25] Hannah Fry: "Behind the scenes, there is a secret key which says bias these words over others, and then when you look back at all of the text together, if you can see those words appearing over and over and over again, you can be confident that this is AI generated."
- [36:15] Pushmeet Kohli: "The ideal situation will be hard to take away to an extent that if you want to get rid of the detection signal, that you have to change the function."
Assessment
This is an official podcast discussion produced by Google DeepMind detailing the theoretical mechanics, rollout, and experimental wet-lab validation of SynthID and SynthID Bio. The conversation presents authentic research methodologies and laboratory results without dramatized demonstrations, contextualizing both the capabilities and the inherent limitations (such as low-entropy short text) of AI watermarking.
Described by gemini-3.8-flash on 2026-10-01 from the video's audio and frames.