> blindly trusting they won't train on any of that
being allowed to train on any data that you can legally obtain ought to be a right for anyone.
After all, i am allowed to learn off anything i can legally read (and perhaps even illegally read). The only thing not allowed (rightly so) is to produce a copy with enough similarities that it can be replacing the original.
It's a bit different when "training on any data" means basically storing a lossily-compressed copy of that data, that could be spit out years later if the model decides to do so.
It's exactly the same problem as with humans, though.
It's part of why we sign NDAs, and why their duration is measured in years (and that's not even targeting the human retention - just duration after which information ages enough that its disclosure is not likely to negatively impact anyone who cares).
It's not exactly the same problem, in that you can parallelize usage of an LLM and copy it over to another computer, but cannot do the same things with a brain. Put it another way, humans do not have the processing power needed to answer hundreds of millions of queries per day, while LLMs do.
> being allowed to train on any data that you can legally obtain ought to be a right for anyone.
I have the opposit viewpoint to the extreme. They shouldn't be allowed to even read that data until they are very clear about what they will or not do with it.
Can they publish it? Can they store it? Can they use the information in it on prediction markets? Etc.
Humans reading texts historically come with little negative consequences, but machines reading and processing texts en masse is more dangerous and should be regulated.
> Humans reading texts historically come with little negative consequences
IDK, we do have laws against opening other people's mail. Those have been on the books for hundreds of years. Seems like someone figured out a while ago that certain unauthorized humans reading certain restricted text wouldn't be good.
> Humans reading texts historically come with little negative consequences, but machines reading and processing texts en masse is more dangerous and should be regulated.
> After all, i am allowed to learn off anything i can legally read (and perhaps even illegally read).
Are you a tool?
Because humans gets rights, tools don't.
Arguing that untrained or partially trained models should have have rights is a different argument to arguing that a trained model should get the same rights as a human.
At the risk of stating the obvious, there are a lot of legal rights that are human-specific (voting, holding office, filling lawsuits, etc.). It's not at all obvious why you think that you as a human being legally allowed to learn from something implies that it should be legal to train an LLM on.
> The only thing not allowed (rightly so) is to produce a copy with enough similarities that it can be replacing the original.
But LLMs are replacing the original, just in different words.
And what does 'legally obtain' mean in this context? Copyrighted content is usually licensed for specific purposes. So if a license is given from training your LLM, then by all means do! But what if the license is 'for personal use'... ?
Oh so if I use mickey mouse in a completely original production that doesn't replace the existing work by Walt Disney, you reckon they'll be fine with that?
> You are one person. The corporation is not. Scale matters
Correct, if you violate it too often to count, you have to pay around less than ~2.5ct per violation.
So the lesson here is: Create a company to do torrenting professionally, and resell its values for higher prices. Then get sued and pay a dime on the dollar you made.
Copyright (or any other such restriction on free use of information) creates power for owners by the simple fact that it turns information into something that can be owned.
being allowed to train on any data that you can legally obtain ought to be a right for anyone.
After all, i am allowed to learn off anything i can legally read (and perhaps even illegally read). The only thing not allowed (rightly so) is to produce a copy with enough similarities that it can be replacing the original.