Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Humans do violate copyright if they use copyrighted passages directly in their work and pass it off as their own without any attribution, which is what copilot has been show to sometimes do, though not always. Copilot will sometimes offer chunks of code that can be found verbatim in open source code bases and passes it off to users without attribution. I agree it is ok to learn from copyrighted work and reproduce new different work from a human or a machine learning algorithm, but it isn't ok to pass along exact copies as your own without attribution. Microsoft will likely need to add checks to prevent copilot from offering verbatim copies of code going forward to try to avoid copyright violations here.


> Microsoft will likely need to add checks to prevent copilot from offering verbatim copies of code going forward to try to avoid copyright violations here.

A user once replied to one of my comment[0] about this with the following:

> It's not really an issue when you're a large software corporation; you already have mechanisms in place to check for license compliance in everything that ships, including F/OSS plagiarism checks [1].

IOW, from my understanding, they don't care. Big players do their own checks anyway, and small fish won't be creating problems because it's too convenient for them. Classic Microsoft (as I know from 90s).

The bigger thread can be seen in [2].

[0]: https://news.ycombinator.com/item?id=32534697

[1]: https://news.ycombinator.com/item?id=32539467

[2]: https://news.ycombinator.com/item?id=32533531


I'm curious about the endgame of copyright with respect to software. At some point, enough people will have written enough code that you can't write code anymore because some fragment of it violates a copyright. Where does the line get drawn? There's only so many ways to do certain algorithms, like DFS or BFS.


>you can't write code anymore because some fragment of it violates a copyright

Copyright (unlike patents in general) allows for independent creation. If I sit down to write a quicksort routine, it is going to look extremely similar to a zillion other quicksort routines out there.

The other question (IANAL) is whether writing a quicksort routine is even a creative act at this point.


At some point it'll have to come home to roost that code is a subset of discrete mathematics first, a literary/artistic work second. There is really no way around it.


> Microsoft will likely need to add checks to prevent copilot from offering verbatim copies of code going forward to try to avoid copyright violations here.

Or they could integrate a way to find the produced output back in the corpus if it's sufficiently close and provide a reference/attribution. Basically whatever tool a copyright lawyer would use to track down original work.

And that's just the engineering solution. The AI researcher solution would be to extend AI learning algorithms to attach attribution metadata to the learned data so that the output could already come annotated with information about the source.

But the latter is much harder to do, so maybe the engineering solution would suffice.


A Twitter thread linked yesterday showed that a keyword and the name of the original code's author in the Copilot prompt produced an almost exact copy of that developer's code. Copilot already does know the origin sometimes.

Edit: here's the related tweet: https://twitter.com/DocSparse/status/1581632706693079042


> Or they could integrate a way to find the produced output back in the corpus if it's sufficiently close and provide a reference/attribution. Basically whatever tool a copyright lawyer would use to track down original work.

That assumes that the licenses of your code and the original code are compatible which often isn't the case.


No, it doesn't assume that. Ensuring that they are compatible would be the next step. Either manually by the user or automatically by showing a fat warning or retracting the suggested code completion.


> which is what copilot has been show to sometimes do

In those cases it seems that humans are already copying code without also propagating licenses appropriately. LLMs are more likely to memorize things which occur a lot (and I'd bet rare things that are representative of some conceptual axis).

The main examples presented so far, Davis and Carmack, have the property of having been copied a lot. The generative model is only surfacing an existing pattern of ignoring attribution. Sort of like the code-gen version of generating bigotry if appropriately prompted.

I'll also note that this pattern of retrievable memorizing of copyrighted and sensitive material is present in GPT-3 too and not just for code. As the situation is equivalent, a lawsuit should address the concerns of non-programmers too.


That seems like a weak defense: "sure, we violated copyright but only because many other people do, too".

Kinda the same problem as YouTube. Lots of people copy movies on the high seas, but if you are as big as yt you cannot easily get away with it.


No, this is not meant as a defense. My point is that it's an issue that is already rampant and what Copilot (or any model) does is make it more readily visible.

This is not like youtube because Github is already hosting those violations and people are already inappropriately copying or including such code. It matters not whether the local inclusion was fetched by copilot or a human fetched it using more manual steps through search.


The difference is if you build tools that can be used to violate existing copyright laws your tool will get taken offline by the same corporations. Yet they are selling one that can be used to do so.


Yes, I agree there's a measure of double standards to this. It's why I feel it's important that AI does not remain in the control of just a handful of corporations. The decks are stacked against though, given how data and compute intensive SOTA is.

But in defense of copilot, code regurgitation is uncommon in routine use. An editor extension allowing search of github would be at least as easy to use to violate licenses but I do not think it'd be taken down since that would not be its core offering.

Copilot goes far beyond mere search and provides a useful service. GPT-3 can also be prompted into generating copyrighted works of writers but I do not see people talking as if that is its primary utility nor as much clamoring in these forums to end that service.


> An editor extension allowing search of github would be at least as easy to use to violate licenses

Well, try doing that for music or movies or proprietary leaked codebase.

If you think copilot is uniquely producing things that are not that different from humans then surely no one would have any problem with feeding it massive amounts of corporate programs?

I am not aware what writers are doing but there have been plenty of uproar regarding stable diffusion. I have a feeling that if any tools like this get built for musicians/film-makers, it will look vastly different from the current situation.


First, the music and to an extent movie industry enforcement of IP are uniquely pathological. But I am not talking about music or movies. I am contending that a simple search extension being much less capable than Copilot and so even more scopeable as aiding copyright violation would not be taken down.

> surely no one would have any problem with feeding it massive amounts of corporate programs?

There is a similar gymnastics done by human engineers today due to the issue of patents. I don't think this is a good trend to uphold.

> I am not aware what writers are doing but there have been plenty of uproar regarding stable diffusion

Yes but mostly in the art community. On HN there were plenty of arguments just the other day how it is not the same for art and programmers have a stronger case. I disagree but regardless, the case is exactly equivalent for GPT-3 and writers but it wasn't an issue generating about a thousand comments on respecting IP and ceasing deployment of LLMs until copilot.


> There is a similar gymnastics done by human engineers today due to the issue of patents. I don't think this is a good trend to uphold.

Are you arguing that copyright laws should be abolished? I have no problem with that as long as it's clearly defined, you can't not respect copyright of open source code but enforce it for proprietary code, just because the value is arguably non-monetary.


Without a sample case I can't say for certain, but couldn't it also be a defense that some code is generic enough that it shouldn't be copyrighted?


One thing I worry about is if the uncertainty around copyright violation cools down activity in open models while raising the price of commercial offerings. Commercial entities can afford devoting resources towards mitigating copyright violations such as eating the cost of maintaining a database of frequently copied code and identifying most likely origin combined with a large semantic database of code snippets.

An open equivalent might be wary of being accused of contributing to copyright violations since in that scenario, there is no way to force people to respect it.


> not just for code

This is quite important, actually, and I don't think enough people realize this. I am a photographer sometimes and it would be really cool if I could share my photos online under a copyright license that forbids their use in training AI.


The problem seems to be that someone could just copy your photo and repost it without that license and we're back to the same spot.




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: