Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> Honest note: Anthropic has not shipped a public Claude watermark detector yet. This tool uses rewrite-based neutralization — a meaning-preserving paraphrase with a non-Claude model — which is the attack path watermark research points to. Not affiliated with Anthropic.

Well, they should have run their own AI slop website through their tool...



This attack was actually pointed out in the watermarking paper linked above. The researchers added an instruction to the prompt that switches letters like a Caesar Cipher. It lowers the quality of the output from the LLM but alters the "red list" enough for a watermark detection tool to fail at detecting the watermark.


also their example for rewriting just completely changes it. might as well redo it in this case (with another model or by hand)




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: