Google DeepMind Unveils SynthID Bio for Watermarking AI-Designed Proteins
Google DeepMind has developed SynthID Bio, a novel technique to embed invisible watermarks into AI-generated proteins, ensuring their integrity and aiding in biosecurity screening.

Google DeepMind has introduced SynthID Bio, a groundbreaking method for watermarking proteins designed by artificial intelligence. This innovative technique embeds a digital signature directly into the amino acid sequence and the predicted three-dimensional structure of a protein. Crucially, wet-lab tests have confirmed that these watermarks do not compromise the protein's intended function, a critical factor for its practical application.
The primary motivation behind SynthID Bio is to bolster biosecurity measures, particularly within the realm of DNA synthesis screening. As AI models become increasingly adept at generating novel protein sequences that may not resemble known biological threats, traditional screening methods struggle to identify potential risks. SynthID Bio aims to provide DNA synthesis providers with an automated way to verify the origin of designed proteins, distinguishing those generated by trusted, safeguarded AI models from potentially malicious designs.
In rigorous laboratory tests, Google DeepMind focused on protein binders – molecules engineered to attach to specific target proteins. By integrating SynthID Bio with their AlphaProteo design method and a modified ProteinMPNN sequence generator, they tested watermarked designs against three key targets: VEGF-A, the SARS-CoV-2 spike protein's receptor-binding domain, and PD-L1. The results were highly encouraging, showing that the watermarked protein binders performed comparably to their unwatermarked counterparts in terms of hit rate, binding affinity, and natural sequence diversity.
For the structural aspect, DeepMind fine-tuned a portion of AlphaFold 3's diffusion network to incorporate the watermark into the model's internal weights. This ensures that any protein structure predicted by the modified model carries the embedded signature. The technique maintains AlphaFold 3's high prediction accuracy, and the watermark has demonstrated resilience against digital noise and minor alterations in atomic coordinates, with detectability reported as near-perfect.
Beyond screening, SynthID Bio also offers potential benefits for public biological databases such as the Protein Data Bank, UniProt, and GenBank. These repositories are susceptible to mislabeled entries, which can have significant repercussions for biosecurity. The watermark could serve as an indicator for synthetic entries, flagging them for additional labeling or review during the submission process, thereby enhancing data integrity.
Despite its promising advancements, Google DeepMind acknowledges that SynthID Bio is not yet impervious to deliberate tampering. The company suggests that further security layers, such as pairing watermarks with provenance metadata or central repositories of AI-generated biological data, are necessary. SynthID Bio is envisioned as one component within a multi-layered security strategy that also includes model-level safeguards and customer vetting processes.
Early feedback from industry experts has been positive. James Diggans, VP of Policy and Biosecurity at Twist Bioscience, described watermarking as a "promising new addition to the biosecurity toolbox that could strengthen screening." Ongoing research is also exploring SynthID Bio's integration into genomic models, with initial tests on a bacteriophage genome showing functional watermarked phages.
Google DeepMind is committed to fostering community engagement and transparency by publishing the methods paper, open-sourcing the code and in vitro data, and releasing the model weights. This move aims to accelerate research and development in AI-driven protein design and biosecurity.