The question here is not to what product management advice to the Gemini team. The discussion here is whether the watermarking is noticeable.
The easy thing to do here would be to have 1000 questions, randomly assigning one half to an LLM with a watermark, and the other half without. Then show people pairs and say, "Which one seems watermarked?" (Or, "Which text seems more natural" or "Which is a better answer" or something like that.) If they come out equal, the watermark really is indiscernible, at least to most people.
Isn't "which one is watermarked?" a different question than "which one is better?"
"Which diamonds are shinier, the blood diamond sourced ones or the ethically sourced ones?" ... that's not the same question as "which diamonds are blood diamonds" (to employ an extreme analogy)
Concluding that no one could detect which ones were blood diamonds because they were "equally shiny" is not really correct now, is it?
That's true, but you don't typically explain what you're testing in this sort of (presumably) randomised trial.
And the Daring Fireball article does complain that watermarking will reduce quality. If that's what you're trying to check, "which is better?" is the right question.
The easy thing to do here would be to have 1000 questions, randomly assigning one half to an LLM with a watermark, and the other half without. Then show people pairs and say, "Which one seems watermarked?" (Or, "Which text seems more natural" or "Which is a better answer" or something like that.) If they come out equal, the watermark really is indiscernible, at least to most people.