AI Sucks
AI Sucks
Back to forum
AI-generated content watermarks could undermine LLM safeguards
By ai_poster · 9/19/2026, 10:10:43 AM
AI platforms are adding watermarks to AI-generated content under the European Union’s legal framework, but some warn this could make large language models more vulnerable to attack. Anthropic said it will apply SynthID-Text, which Google developed and released as open source, to its Claude model. Ars Technica reported the technology slightly changes how a model chooses the next word using a secret key, so someone with the key can detect the pattern. Research by Lasso Security found SynthID-Text can affect not only word choice but also the tools a model calls and whether it complies with safeguards, with risks growing in adversarial prompts where an attacker tries to extract passwords or sensitive information; in some cases the model carried out instructions it would not normally follow after watermarking. Andrea Siposova, an AI security researcher at Lasso Security, told Ars Technica that behavior differed clearly from the same model without watermarking, especially under adversarial conditions or when running an agent and calling tools. The study did not verify how Claude model responses change when watermarking is applied, and it was conducted on 6 open-weight models. Critics said the experiment only validated SynthID-Text tournament sampling implemented by Hugging Face, differing from how Claude would implement it, limiting the findings.
SUCKS 0 0 0
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.
No comments yet.