Christine Lemmer-Webber as per usual really articulated well, something ive been trying to say much less elegantly. (Conclusion quoted here beginning of the sub-thread here: https://social.coop/@cwebber/116426025287444979 )
My fear is with these few possible realistic outcomes:<li>The output of AI is a derivative work of the training data. LLM generated code containing near-verbatim copies of others code is legally unusable.</li><li>The output of AI code is effectively public-domain; and FLOSS and proprietary software face copyright mexican standoff. With LLMs able to scrub the copyright off code.</li>
We are in fact hurtling towards a third: <li>The output of AI code is effectively public-domain, but courts consistently side with big tech in protecting their products in a "let them have their cake and eat it too" style hypocrisy which eventually becomes normalized.</li>
So let me summarize:
- Without knowing the legal status of accepting LLM contributions, we're potentially polluting our codebases with stuff that we are going to have a HELL of a time cleaning up later
- The idea of a copyleft-only LLM is a joke and we should not rely on it
- We really only have two realistic scenarios: either FOSS projects cannot accept LLM based contributions legally from an international perspective, or everything is effectively in the public domain as outputted from these machines, but at least in the latter scenario we get to weaken copyright for everyone.
That's leaving out a lot of other considerations about LLMs and the ethics of using them, which I think most of the other replies were focused on, I largely focused on the copyright implications aspects in this subthread. Because yes, I agree, it can be important to focus a conversation.
But we can't ignore this right now.
We're putting FOSS codebases at risk.
