Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> I can imagine finer-grained exclusions, such as allowing full-text indexing but only for accounts on the same instance, or allowing use for search but no other applications. (No ML model building!)

I think it's unlikely that you can prevent ML model building with a carefully designed license. The most common legal position (though not something that has been tested in court yet) is that training models is sufficiently transformative to count as fair use, and does not require any sort of license to the data.

You can see this in all the state of the art tools that are trained on all the publicly available data that they can scrape, without regard for license: translation (text), GPT-3 (text), Stable Diffusion etc (images), Co-Pilot (code).

For preventing trolling and harassment a licensing approach is an even worse fit, since those are not people who care about respecting licenses.



None of those tools have actually been legally tested, and there is a reason much of it has been using data sets laundered through academics.The companies behind them know this is not at all a given, and academics make for more sympathetic defendants than billionaires. Transformative use, as one part of a fair use consideration, is a defense against copyright infringement. The proposal is to first require agreement to a separate license before even being able to access the content. This is an additional layer, which may or may not be enforceable but would definitely establish either negligence or intent, and also brings things like unauthorized access into it. The fact that all of this also includes huge amounts of PII means that in a growing number of jurisdictions misuse of it will not be protected by any copyright exceptions. It would be an endless battle to stop smaller abusers, but you could definitely prevent GPT-3, Stable Diffusion, and Co-Pilot, since they are all coming out of well-defined legal entities with assets and identifiable humans to go after.


There are jurisdictions (the EU, the UK, Singapore, Japan) with copyright exceptions specifically for text and data mining for AI purposes.

https://www.twobirds.com/en/insights/2021/singapore/coming-u...


That's interesting thanks. I wasn't aware of the Singapore one. It seems to be the broadest, but based on the linked page, it's not clear to me how it would come down here. It requires legal access to the material first, but also says you can't contractually override the copyright exception. I don't know how they would weigh it if you're only granted access based on that contract (rather than it being a small part of a broader contract).

For the case of the EU, based on the way the GDPR and related digital laws are drafted very much in a "spirit of the law" and with individual rights and agency trumping corporate interests, it seems fairly likely it would not just allow coopting personal social media content over the explicit wishes of the creators, regardless of whether any access was deemed legal. For the moment, I think the UK digital laws are still basically just copies of the EU ones as well (with some search and replace) but I guess they'll drift apart over time if there's no de-brexit.

It's also worth noting, especially given how it's written about in the linked post, that these are generally assuming that these uses do not destroy the market for the original works, and they come from the reality before all the art generation. The specific implementations everyone is talking about, especially in the art space, have now pretty definitively proven that we're in a new reality now.


Very good points. Thanks for your thoughts.


What would be the ramifications of one of those entities releasing their model as a torrent at the first sign of legal trouble?




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: