> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models
That basically means, we don’t know, and we hope the model didn’t look up user conversations, and the best thing we can do is hope.
That’s seriously disgusting. I can understand why on a technical level why perhaps it is impossible to answer what exactly the model had access to, but it still is disgusting.
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models
it just means that they've been working with paraphrased user data everywhere, which any smart person can figure out is how anthropic and openai train on so called non-retained data.
How could they possibly know? If Tristan posted on r/math and they slurped that up as training data, that would count, no? They might never even know. I can’t envision any absolute statement by them claiming that they didn’t use his work that survives legal rigor. That is, this statement was never not going to be in this post in any of the infinite multiverses.
This is private work that is being done with Codex. Not drafts shared on r/math. I suggest you reread the article.
This is more of an expectation of privacy issue. If you write on an envelope and USPS has a copy, you have no reason to be mad; if you write on a letter inside the envelope and USPS still has a copy, you could rightfully be mad and say this is disgusting.
Every other paper in existence has been ingested with 99% of writers not knowing it will be retroactively used for training. But session data which is disclosed as being used in terms of service is disgusting?
If the work is duplicative/derivative then the preprints they put in sessions can be shown by the users and we can see.
What's disgusting is that they refuse to acknowledge what they and only they know: were these chats fed in to the new model that found these results?
Nobody would be disgusted at following stated policy, that's doesn't make sense, without first objecting to the policy. But you're bringing up distractions from the actual concerns: did OpenAI use the private chats and why won't they confirm or deny it?
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models
That basically means, we don’t know, and we hope the model didn’t look up user conversations, and the best thing we can do is hope.
That’s seriously disgusting. I can understand why on a technical level why perhaps it is impossible to answer what exactly the model had access to, but it still is disgusting.