>Makes sense to me. AI models are trained on (as large of a subset as possible of) the sum of human knowledge, so their outputs should belong to humanity as a whole.
No it doesn't, because that would mean any sort of secondary source shouldn't be eligible for copyright either, eg. encyclopedias, which are basically rehashing "the sum of human knowledge".
Interesting POV. Remind me - did those encyclopedias pay experts to write their contents, or did they just hoover up every bit of text they could find - regardless of owner - and toss it into a blender?
Sure, I'll remind you. The most popular encyclopedia of our time, Wikipedia, doesn't pay most of its contributors (I think they do have some administrative staff or something?). Nor do they pay journalists or authors for the articles and books that they cite.
Interesting, interesting. And this wikipedia - it gets its contents by hoovering up the web, then? Or do you maybe want to put on your thinking hat and consider the difference between voluntary contributions and theft?
Yes, Wikipedia gets its content largely by hovering up the web, without the consent of the authors. For example you can cite a New York Times article in Wikipedia, without getting the consent of the NYT author. Wikipedia also "hoovers up" (as you put it) offline sources like books. Again without consent!
Indeed, the need to get consent from the author before reading a published work Isn't A Thing in general, outside some very specific contractural scenarios.
Why are you claiming that citations are the same thing as theft without citation? There’s a mile of difference between citing a work and taking it, rewording it, and not crediting or compensating the original author.
The whole job of a Wikipedia author is to take source material and reword and summarize it into an article. They're essentially acting as human LLMs. Wikipedia even has very specific rules against "original research." You are supposed to be acting as a summarizer, not a researcher or author. Wikipedia also does not compensate the original authors of the source material.
Really the only difference between the Wikipedia author and the LLM is that the Wikipedia author will more frequently be asked to provide citations. But the LLM can also provide citations if asked. In neither case are the authors of what is being summarized compensated or asked for permission. In neither case is it theft.
> There’s a mile of difference between citing a work and taking it, rewording it, and not crediting or compensating the original author.
So what's wikipedia doing vs what LLMs do? So far as I can tell the only difference is in citations, but:
1. LLMs can be made to cite, eg. if you use google search's AI mode it'll happily provide citations. I doubt that would placate the AI haters though.
2. Outside of academia no one really cares about citations. There's no legal requirement to cite, nor do I think all the people complaining about AI "stealing" other peoples' work are going to be magically placated by the addition of a few citations. Moreover it's unclear whether the concept of citations makes sense in many contexts. If you ask a human programmer how to write fizzbuzz, they'll likely blurt out a solution without providing citations, much like an AI would. Same for most questions people are asking AI about, eg. "gimme a cake recipe", do you really need a citation back to some 18th century cook book?
For Wikipedia, people look up sources to create some original work. They do it while citing the original work but that's not the major point. That's clearly all lawful, you are allowed to do that.
For LLMs, people (or let a software do it, doesn't really matter) copy original work and use it for commercial purpose. If that's allowed fully depends on the license of the original work, but I didn't think anybody really beliefs that OpenAI and others check the license for every single original work they ever used. So it's basically the biggest copyright infringement ever.
Everything else is just framing coming from big companies.
So, the problem already starts while training the model.
Regarding its output, if it happens to output work that falls under copyright, the LLM company must make sure that it obeys the license connected to it (i.e. citing or not relaying the result to the user). Obviously, nobody does that and it's also not generally possible to do that anyway. So that would be second biggest copyright infringement ever that only works because it's hard to track when such an infringement happens.
So, if asked "could we use your work for our commercial software that might output something that would be still protected by your copyright, but nobody will be able if or when it happens and we won't check and won't tell the users" nobody would have given consent. So they went "duck it, we are talking about billions of dollars and AI is great etcetc., so let's just do it anyway"
I've explained this many times. LLMs don't "copy original work." ChatGPT doesn't contain copies of books inside it, any more than your brain is a copy of the various books you have read. The LLM model isn't physically big enough for that, it's like saying you fit 1000 gallons of water in a 1 gallon milk bottle. Can't be done. The model may be able to quote small snippets of works, just like you might remember various quotes from Shakespeare or someone. (That's not infringing either, by the way)
Wikipedia, and LLMs, can refuse to cite sources and still not infringe copyright. Citation simply isn't relevant to copyright. Not citing a work that you read previously is not a copyright infringement. Wikipedia or OpenAI being non-profit, or for-profit businesses, has nothing to do with copyright. Consent has nothing to do with copyright. Copyrighting something doesn't mean that you can require everyone who reads it to get your consent. You can require everyone who distributes it to get your consent, but once it's been distributed to someone, they can read it freely.
>For Wikipedia, people look up sources to create some original work. They do it while citing the original work but that's not the major point. That's clearly all lawful, you are allowed to do that.
>For LLMs, people (or let a software do it, doesn't really matter) copy original work and use it for commercial purpose. If that's allowed fully depends on the license of the original work, but I didn't think anybody really beliefs that OpenAI and others check the license for every single original work they ever used. So it's basically the biggest copyright infringement ever.
So what makes wikipedia (and other encyclopedias) legal but chatgpt not legal? By your own admission citation isn't "the major point". Wikipedia might get a pass because it's a non-profit, but every other encyclopedias operate on the same model.
I don't see how "experts" are relevant under OP's framework unless they're did the primary research themselves. Otherwise they're just regurgitating someone else's research. The "oh there's humans involved so that gets pass" excuse doesn't work either, because humans were also involved in training the AI.
Please reread the comment you are replying to. I will wait.
—-
Now that you have reread the initial comment, do you think that “experts” was the important part? Or do you think maybe it was the compensation for their work that matters?
>Please reread the comment you are replying to. I will wait.
If you're going to post thinly veiled implications that I didn't read your comment, you should be pretty damn sure that you make it look like you read my comment, which it doesn't seem like you did. The second of my comment said:
>The "oh there's humans involved so that gets pass" excuse doesn't work either, because humans were also involved in training the AI.
If you did read it, you sure did a poor job at rebutting it, leaving it unaddressed and preferring to waste words on writing snarky remarks instead.
> If you're going to post thinly veiled implications that I didn't read your comment
It was a statement, not an implication.
You still haven't replied to the point in the original comment about compensating the people who do this work, so I think it's quite obvious that you haven't read it.
>You still haven't replied to the point in the original comment about compensating the people who do this work, so I think it's quite obvious that you haven't read it.
Issac newton discovers the theory of gravity. He advanced the sum of human knowledge, so fair enough, he should get compensated.
Alice rehashes that and puts it into her encyclopedia, allowing others to learn the theory of gravity.
Bob writes an algorithm for training a chatbot that can produce responses rehashing the theory of gravity, also allowing others to learn the theory of gravity.
Why should Alice be compensated but not Bob? Neither discovered the theory of gravity, so it's not like by funding Alice we're helping discover quantum physics or whatever. It's also not obvious that Alice's work is more valuable. A chatbot interface is often better at teaching someone than a rehashed overview. Of course, you can try to fix this by declaring that human work is valuable and an AI model isn't, by fiat, but that's just a cope and a far cry from the original principle of "trained on [...] human knowledge, so their outputs should belong to humanity"
None of this matters for applying the law, because the law just says only human created works are eligible for copyright protection, but that's not the argument OP was trying to invoke.
Finally none of this actually matters because OP just bites the bullet and says that secondary sources shouldn't be eligible for copyright, period.
Like, now with the advent of AI, or even before? You might not have much love for encyclopedias, which were mostly replaced by wikipedia, but the "secondary sources don't get copyright protection" would also cover programming books, which roughly speaking are docs rewritten to a cohesive narrative.
No it doesn't, because that would mean any sort of secondary source shouldn't be eligible for copyright either, eg. encyclopedias, which are basically rehashing "the sum of human knowledge".