I stopped reading when it said the repos had once been public and were in a Bing cache. That is probably true of a lot of once public, now private, data. Not just in Bing, but also in archive.org, google. It said the repo in question had a “public phase”.
Anyway, there is no indication that your data would be exposed if you had a private repo that was always private.
You've got to treat any repository made public - no matter how briefly - as compromised / downloaded / accessed / viewed.
I found the article title pretty misleading. There's perhaps an interesting conversation to have around the possibility of wanting to be able to scrub training data after-the-fact and how (or if) that could work - but that's not what the article's headline tried to convey.
Neither of these expunging mechanisms work if you're not the domain owner, however, so this is just one more reminder that any content you upload to somebody else's website is never fully under your control.
Is it just me or is it painfully obvious that article was written by AI? Or am I being mean to the last real person writing titled bullet points every 3 paragraphs.
Or the wronged parties were those who entrusted someone else not to put their data (e.g., closed software, NDA material, proprietary internal info shared with contractor, access control secrets) in a public repo.
Anyway, there is no indication that your data would be exposed if you had a private repo that was always private.