Well, as a user of LLMs I also find it annoying, the agents gather information on my behalf, they save me time and money. Are we splitting humanity? Those with info accessible to agents and those who only use their human sense directly?
This whole distinction is futile imho. And no, I'm not Sam Altman.
I mean training is one thing, you should honor people's licenses, but browsing and gathering information? Why force me to do it with my biological neural net?
My biggest fear with AI is that it basically devalues the production of knowledge. They’re a continued march away from actually showing people the original information, and the people doing that writing get less and less credit.
The companies themselves know they depend on this knowledge but have no plan to fund or replace the work done on sites like Wikipedia, stack overflow, or blogs like acoup.blog
Honestly, I'll take it a step further: what's wrong with training?
Having principles means applying them uniformly -- even to large entities or those you hate (it's fine if the principles themselves have size bounds in them, though -- versus them being implicitly glued on -- but then you need a universal justification for why that size. Which is possible and valid.)
I think that one should use the best algorithms, the best information, the best knowledge they have access to -- period. I don't like using gimped machines, I don't like making gimped machines, and I certainly don't like being sold them.
So I'm not going to turn around and say LLMs need to be gimped via arbitrary restrictions on their training data.
I agree. As soon as I understood the gist of how modern models do what they do, I found it unreasonable to apply any stricter standards to their "learning" than we do to humans. Humans learn by reading ideas and looking at images, then go on to have ideas and draw images based on that training.
We have definitions of plagiarism and copyright infringement that apply to people based on what they put out into the world, not how they trained themselves to get there[0]. Artists literally trace art and study specific examples in detail to learn, and that's not a bad thing. Humans can also accidentally plagiarize or make things identical to past works - it's easy to think a great guitar riff just came to you when it was really based on a song you heard years ago that stuck in a part of your mind but you don't even consciously remember the influence, for example.
It would certainly be nice for AI output to provide citations if it realizes it's using a significant chunk of an idea from its training that has a clear source (or many), though this is difficult in the same way it would be difficult for me to cite where I learned about the Towers of Hanoi. The AI frequently does web searches for specific resources to get ideas from these days, which are easy for it to cite.
However, it would have a chilling effect on progress as a whole if we all started jealously guarding our ideas so close to our chest that machines couldn't read them and only a select few humans who passed some kind of gate (or even paid us) were allowed to see them. We would be nowhere close to where we are if we had always had such a mindset.
---
[0] We use proven absence of viewing certain material as a legal shield against copyright infringement e.g. clean room engineering, but this isn't strictly necessary, and possibly even discouraged with modern precedent: https://reactos.org/forum/viewtopic.php?t=21740
This whole distinction is futile imho. And no, I'm not Sam Altman.
I mean training is one thing, you should honor people's licenses, but browsing and gathering information? Why force me to do it with my biological neural net?