Copyleft is a way of leveraging copyright against itself to ensure nobody else can make a proprietary version of the thing. It is not the same as Public Domain/lack of copyright.
LLM output doesn’t automatically get a copyright. But, if it is (part of) a work that includes “human creative effort” (prompts don’t count), the human(s) can hold a copyright on that work.
In addition, the output can still be a derivative work in violation of the copyrights of (some of) the training data, whether or not there are copyrights on that output. It would have to have sufficient similarity to some work in the training data, but that’s not too uncommon.
And, GPL and CC-SA works are known to be in the training data of most models, including Apertus.
All LLM code output is now copyleft because there was GPL stuff in the training data, LOL!
spoiler
(Actually it’s probably all just copyright infringement and not usable at all because of all the conflicting licenses, but a guy can dream…)
I think LLM code output is already copyleft because non-humans legally can not hold copyrights.
Copyleft is a way of leveraging copyright against itself to ensure nobody else can make a proprietary version of the thing. It is not the same as Public Domain/lack of copyright.
in my defense there are a lot of words to remember
LLM output doesn’t automatically get a copyright. But, if it is (part of) a work that includes “human creative effort” (prompts don’t count), the human(s) can hold a copyright on that work.
In addition, the output can still be a derivative work in violation of the copyrights of (some of) the training data, whether or not there are copyrights on that output. It would have to have sufficient similarity to some work in the training data, but that’s not too uncommon.
And, GPL and CC-SA works are known to be in the training data of most models, including Apertus.
:(