🌿freegardner

Science

Watermarking AI Designed Proteins to Flag Threats

01 Oct 2026 · via Feeds.arstechnica

Watermarking AI Designed Proteins to Flag Threats
AI-generated image

Watermarking AI Designed Proteins to Flag Threats

Why Similarity Screening Fails

For years, the working assumption in biosecurity was simple and, in retrospect, too simple: if you can recognize a dangerous protein, you can stop it. Recognition came from similarity. A DNA sequence looked threatening if it resembled something already known to be a toxin or a piece of a virus. That method works fine for nature’s back catalog. It fails completely for proteins that never existed until a machine imagined them.

The reason is that nobody has characterized AI-designed proteins well enough to know that they are threats. A designer protein can be entirely novel, sharing no meaningful resemblance to anything in the databases, and still fold into something that binds a target with high affinity. The screening software that DNA synthesis companies rely on has nothing to match against. So it waves the sequence through.

Nearly a year after that risk was flagged, it still was not clear what anyone could do about it. That is the foundational assumption that had to be abandoned: the idea that identification must come from comparison to known threats. If the threat has no ancestor in the database, comparison is a dead end.

What replaced it is a different logic entirely. Instead of asking what a protein resembles, ask who made it and whether they left a signature. The system embeds a watermark in the protein sequence itself, and the watermarked versions still fold and bind their intended targets. New proteins designed by trusted researchers become identifiable, and everything else opens up to closer scrutiny.

Watermarking AI Designed Proteins to Flag Threats (Image 1)
AI-generated image

The shift is subtle but total. A watermark does not tell you a protein is safe. It tells you a protein came from a source that chose to be identifiable. The mark is no longer a checkpoint at the door. It sits inside the sequence, and only someone holding the key can read it.

From Pixels to Proteins

The watermarking method is built on Google’s SynthID technology, which adds a subtle watermark to AI-generated digital material. [1] The team started from ProteinMPNN, one of the most popular AI protein design tools, developed by the Baker Lab The mechanism is probabilistic rather than additive. The watermark influences the probability of certain choices the AI makes, and that bias ends up systematically distributed throughout the product, whether text or images. Because it is spread across the whole output, it survives basic exporting, resizing, and similar transformations.

Translating that idea from pixels to proteins is where the difficulty lives. A protein is a chain of amino acids, and each position in that chain can matter for whether the protein folds and works. Some amino acids are chemically similar, like leucine and isoleucine, while others carry opposite charges. Many proteins have regions where limited changes are tolerable and other regions where even a slight deviation inactivates the protein. There is also the question of scale. An image offers a vast number of places to hide a signal; a protein offers far fewer. It was not obvious the SynthID approach would transfer at all.

The chemistry matters here. Every amino acid has a constant section, and a protein forms when a series of these sections link into a long chain called a backbone. Each amino acid also carries a side chain, and those side chains determine how the backbone folds in three-dimensional space. Once folding completes, the backbone describes the protein’s overall shape.

ProteinMPNN works in a two-stage process. First, a separate tool describes a backbone configuration that is appropriate for the design. Next, ProteinMPNN works its way down the backbone, placing side chains one amino acid at a time. Each amino acid is chosen based on its ability to fit into the shape defined by the backbone, interact with neighboring amino acids, and fit any other constraints defined by the experiment, such as forming catalytic pockets or interacting with another protein.

Watermarking AI Designed Proteins to Flag Threats (Image 2)
AI-generated image

A variant of Google’s SynthID, called SynthIDBio, steps in during this process. [1] It uses a key, similar to a cryptographic key, and the identity of the previously chosen amino acids to suggest a new one. ProteinMPNN then determines whether the amino acid suggested by SynthID works from the perspective of forming a functional protein. If it doesn’t, it rejects it. If it does, it moves on. Put differently, as the system works through the backbone one amino acid at a time, it only incorporates watermark amino acids when they’re consistent with a functional protein.

Where several words mean nearly the same thing, the translator picks the one that fits the hidden pattern. Where only one word will do, the pattern yields. The result reads naturally to anyone who does not know the code, yet the code is scattered through every page. g it is not a simple yes-or-no question. You have to scan the whole sequence, knowing the key, and measure how often the amino acids suggested by SynthIDBio actually appear in the final sequence. [1]


> Note: Only birds’ vision was tested; other predators may perceive patterns differently.

Sources

1. Ars Technica — Quote source (original article)

Mentioned organisations (context, not sources)

- Google — Organisation (homepage)

← back to the garden