The use of large language models (“LLMs”) has exploded in recent years, including in the generation of source code. But even as their usage gains popularit...
An LLM, being a large database derived mechanically from its training data, seems to meet the definition of “derived work” as exists in most copyright legislation.
The output from an LLM, through the inference process, is also mechanically derived from that training data.
Lawmakers might try writing exceptions to please a stock market absolutely baying for hyperscaler growth. Despite that, it seems entirely viable that the copyright, held by humans on the training data fed to the LLM, continue to hold in the output derived from that input.
People saying it’s “obvious the output has no copyright”, or people who act as though it’s all free from any claim by the humans who wrote the works used as training data, are just showing their ignorance, it seems to me.
People learn from books and existing code. Does that means that anything we write is derived work?
An LLM knowledge matrix is a neural network, not a database. It does not contains an exact copy of any given text. It may remember it, but human brain can do the same.
You have to prove LLM code is derived work in the same way you prove human work is derived work: check the difference between the original work and the derived work.
No, i don’t think so, i said “if so”. A piece of code can’t own something, unlike a person. Would be a different matter, if the piece of code was a person, but alas.
An LLM, being a large database derived mechanically from its training data, seems to meet the definition of “derived work” as exists in most copyright legislation.
The output from an LLM, through the inference process, is also mechanically derived from that training data.
Lawmakers might try writing exceptions to please a stock market absolutely baying for hyperscaler growth. Despite that, it seems entirely viable that the copyright, held by humans on the training data fed to the LLM, continue to hold in the output derived from that input.
People saying it’s “obvious the output has no copyright”, or people who act as though it’s all free from any claim by the humans who wrote the works used as training data, are just showing their ignorance, it seems to me.
People learn from books and existing code. Does that means that anything we write is derived work?
An LLM knowledge matrix is a neural network, not a database. It does not contains an exact copy of any given text. It may remember it, but human brain can do the same.
You have to prove LLM code is derived work in the same way you prove human work is derived work: check the difference between the original work and the derived work.
If so, the copyright of the slop would be by the company running the LLM.
So self hosted inference means you own it then?
No, i don’t think so, i said “if so”. A piece of code can’t own something, unlike a person. Would be a different matter, if the piece of code was a person, but alas.