This is just a nonsense line of reasoning. Training based on the solution to the problem (or the key insight behind the problem) is clearly a form of plagiarism.
What about my line of reasoning is nonsense? I made no claim either in support of or contrary to yours. Rather I pointed out that by this logic literally everything that an LLM spits out is plagiarism of the vast majority of the entire body of human literature in existence. Can you offer meaningful refutation of that observation of mine?
Many do indeed hold the position that all LLM output is uncopyrightable plagiarism. They're probably right, but there's an even stronger argument here:
Science papers of a phd level must contain:
1. one or more novel insights
2. a long list of citations to contextualize them and
3. some work to prove that the insights are in fact meaningful
---
In this context, consider a prompt based diffusion model which, when asked, will happily produce a few pictures of a horse in orbit. You then tell it "silly robot, horses can't breathe in space" to which it adds the necessary space suit in a follow up image.
That image is twice plagiarized:
1. the model did not come up with the original idea of putting a horse in space, nor with insight that horses need a space suit
2. the model failed to cite where it pulled the "horse" and "space" concepts from.
It merely did the work (3) to combine the concepts using the user provided insight.
---
The implied accusation here is that OpenAI used the insights from an existing prompt to train a new model that was able to one shot "a horse race in space" picture, and they were all wearing space suits.
This is still academic plagiarism, even if you disagree that all LLM outputs are.
I neither agree nor disagree that all LLM outputs are plagiarism. I merely objected that the line of argument engaged in was specious given the context.
As to your stronger argument. You only cite prior novel insights that you're actively building off of and that (approximately speaking) fall outside of the status quo. You don't for example cite leibniz or newton despite your paper making heavy use of calculus.
So is there any actual evidence that openai trained on the data in question? And further, did the openai proof directly build on someone else's novel insights as opposed to deriving everything from scratch? (I don't pretend to know but the vast majority of what I've seen so far in the comments here is what I'd characterize as brain-dead screeching. Certainly not the level of discussion I come to HN for.)
Separately, consider the implications of what you're arguing for there. Suppose your horse in a space suit picture were somehow valuable to society. Suppose that due to shortcomings of your tool you lacked the ability to readily and accurately identify the originators of the relevant concepts. Should you refrain from publishing this useful work due to the lack of citations? How are you supposed to handle this situation?
Remember that in this analogy everyone throughout society is on the same page that your tool consistently recycles other people's ideas while being technically incapable of producing reliable citations. The question is a simple trolley-esque problem - do you publish without proper citations for everyone's benefit and if so what are you supposed to say?
evidence that openai trained on the data: they would have denied it if they didn't train on it.
did the proof build on the insights:
the influence of an individual text in the training data is deeply weighted by quality, relevance, etc. a high quality proof in advanced mathematics written by a codex user is going to get boosted to the max.
the model is post-trained on prompt material. that is again going to boost it.
the prompt will boost this material specifically. perhaps they even rammed dense maths in particular into the model in post training.
anecdotally i have been able to get near-verbatim copies of original material out of models at inference. the type of work that buckmaster and alpoge fed into openai feels like the exact type of concept that would cause an "aha!" or "but what if?" in chain of thought. in fact i would bet that their work is in the logs.
the likes of astra and fable are thought to be up to 10T parameters in size. i consider it highly plausible that a semantic representation of the euler proof could be pulled out of the model weights in good shape.
The chats the professor had are not generic knowledge. And yes of course you still need to cite Newton and Leibniz depending on what result you want to mention. What’s allowed to be not cited are not status quo, the term you’re looking for is “folklore” results aka results that have been around so long that 1) nobody knows who came up with them or 2) everyone knows who came up with them.
The second point: if you say you can’t prove that OpenAI actually used it, it doesn’t mean that OpenAI did not use it. It’s hacker news not lawyers news here lol. And OpenAI can’t prove that they didn’t use it either. The whole point is that Levent felt he had reasonable suspicion to believe the AI did use the result, because he felt like without his input on an unpublished paper it was unlikely for AI to reach the same result. I haven’t read the paper so I don’t know where I stand on that.
On the last point, about your “for the greater good” argument. It’s higher maths lol. I don’t know about this field but I doubt it’ll be very useful for society. Maybe it’ll make one part 2x faster which makes some rocket cheaper to launch. Does the average person care? Debatable. I think it’s reasonable to hold published papers in proof based fields to a higher standard. Otherwise the current & future problems of ML engineer fields just expand to other fields. No thanks.
Finally, if you anonpost to the autistic Internet forum that everyone else is “brain dead screeching”, it really just says something about yourself lol.
Even if AI used the result, AI pushed it to the finish line while Levent and Tristan did not. But I understand the approach was different, the information leak was only that it was "doable".
That's exactly the argument of the people calling it plagiarism machines. No-one ever really did refute it there was just a bunch of settlements for elite institutions so they weren't left empty handed like the various small time creators/authors etc were.
I think the bigger issue here is this feels like some PR smoothing happening that after all the work that went into "it's safe to use for enterprises" now we have what looks like openAI using private user data to scoop novel research and the question of why couldn't they do it for an enterprise with much more money on the line.
> now we have what looks like openAI using private user data to scoop novel research
Is there any actual evidence of that? All I've seen so far are empty accusations because "it would be in their interests" or whatever. Personally I'm inclined to believe that they honor their terms until it's demonstrated otherwise.
What does it matter? We're supposed to not call it plagiarism anymore because it's inconvenient to call it the plagiarism machine? What's your actual argument? Otherwise it's completely irrelevant what an LLM does in other contexts or what we call it
Another problem I foresee for academia given the behaviour of AI companies is that even if they don't share their research with ChatGPT, as soon as they submit it for publication many reviewers likely will. Especially if the initial submission is rejected they then risk getting scooped. Possibly uploading preprints to arxiv could help.
Interesting that they have their own open-source OCaml toolchain for chip design. I thought received wisdom was that everyone in industry is still tied to horrendous vendor toolchains. Is this a realistic alternative for production-grade chip design?
I agree it's super annoying, but I can say for sure it's not limited to the US. Has also happened to me many times over here in Ireland/the UK. Usually by waiting staff that are a little inexperienced. Of course the people I'm in conversation with at the time are probably relieved to have a break from me rabbiting on...
One concern is that it will become more and more challenging to conduct cutting edge maths research without substantial resources only available at very rich institutions (to pay for state of the art AI assistants).
Out of interest, what would you estimate the proportion of new maths that is used by other fields to be? Do you think much of this new maths is potentially underutilised as it were?
reply