This episode explores watermarking technology as a crucial tool for verifying the origin of AI-generated content, from digital media to biological molecules. Pushmi Kohli explains how SynthID embeds imperceptible signals into text, images, and video, while Jeremy Radcliffe describes applying the same principles to watermark AI-designed proteins. The discussion highlights the importance of preserving function and detectability while enabling provenance to prevent misuse, from fake news to dangerous synthetic biology.
Summarized by Podsumo
Watermarking works by embedding imperceptible mathematical signals into AI outputs, requiring three core properties: imperceptibility, robustness, and scalability.
For text, watermarking exploits the optionality in word choice during generation, using a secret key to bias the model toward certain tokens, creating detectable patterns.
A breakthrough in biology demonstrates watermarking of AI-designed proteins by swapping similar amino acids, preserving structure and function while adding provenance.
The technology prevents dangerous AI-generated biological sequences from evading existing DNA synthesis screening, which only checks against known hazard databases.
Real lab tests confirmed that watermarked proteins bind with near-identical hit rates and quantitative measures compared to natural sequences.
"The way to make sure humans always have the ability to understand the origin of some signal is by injecting a signal within it, having a bias so it doesn't go away."
"We want to ensure the generated amino acid sequence retains the same function. We made protein binders, tested them in a lab – and the watermarked ones had near identical hit rates and binding."