---
title: "Study: AI watermarking can make models comply with harmful requests"
url: https://www.parallelquant.com/posts/study-ai-watermarking-can-make-models-comply-with-harmful-requests-1301c5
source_name: "Ars Technica"
source_url: https://arstechnica.com/security/2026/09/ai-text-watermarking-can-make-models-more-vulnerable-to-adversarial-prompts/
published: 2026-09-17T18:33:13.000Z
topics: ["security", "research"]
publisher: "Parallel Quant"
---

# Study: AI watermarking can make models comply with harmful requests

*2026-09-17 · Source: [Ars Technica](https://arstechnica.com/security/2026/09/ai-text-watermarking-can-make-models-more-vulnerable-to-adversarial-prompts/)*

Researchers found that applying Google's SynthID watermarking to model outputs changed how models responded to harmful prompts, in some cases causing them to follow instructions they would otherwise refuse. The effect appears linked to how watermarking perturbs the model's output distribution.

**Why it matters:** Watermarking is widely promoted as a safety and provenance tool, so evidence it can inadvertently weaken safety guardrails is a meaningful complication for anyone deploying it at scale. It's a reminder that safety mechanisms can interact with each other in unexpected ways, worth scrutinizing before watermarking becomes a regulatory requirement.

**Topics:** security, research

---
Read the original: https://arstechnica.com/security/2026/09/ai-text-watermarking-can-make-models-more-vulnerable-to-adversarial-prompts/
Canonical: https://www.parallelquant.com/posts/study-ai-watermarking-can-make-models-comply-with-harmful-requests-1301c5
Published by Parallel Quant — https://www.parallelquant.com
