I am not a lawyer.
-
Don’t you hate it when you write a long post, proofread it, post it, see it in your feed, and immediately spot a typo in the first line?
(Fixed now, it said ‘firms’ when it should have said ‘forms’)
@david_chisnall I would love if it sometimes didn't happen.
-
Don’t you hate it when you write a long post, proofread it, post it, see it in your feed, and immediately spot a typo in the first line?
(Fixed now, it said ‘firms’ when it should have said ‘forms’)
@david_chisnall every single time I write some important-to-me post
-
@david_chisnall If I could only boost this multiple times! Also, any background on Watt subverting patent laws?
@bjn There’s a lot of steps but the short version is that he had some patents on key parts of steam engines and, every time they were about to expire, bribed politicians to extend them, much as Disney did with copyright in the USA. Some of these extended the terms of all patents, others special cases Watt’s patents.
This probably both slowed down the Industrial Revolution and ensured a specific concentration of wealth during the process.
It’s a case study I wish was required reading for MPs entering parliament because they keep repeating the same model, often through ignorance rather than corruption (though sometimes both).
EDIT: He also had a habit of using the patents to push competitors out of business and then hire buy their assets and cheaply and hire and underpay their assets. He was probably an inspiration for Edison and Gates.
-
I am not a lawyer. I have studied IP law in various forms guided by solicitors, barristers, and professors of law of my acquaintance since I was a teenager (including reading through big piles of case histories and commentaries) but not in a formal setting (mostly I learned that I would rather invent the things in the patents than draft the patents, so changed career direction). I’ve also spent a surprising amount of my career talking to copyright and patent lawyers (almost never trademark or trade-secret specialists). I have enough of a lay-person’s understanding of the topic that I was able to spot that a clause in the contract from my US publisher was unenforceable in the state that they claimed jurisdiction because it hinged on an aspect of copyright law that the USA delegates to states and which the state in question did not have relevant laws. Their lawyers subsequently confirmed and fixed this. But I a not lawyer and this is not legal advice.
I have two objections to the EFF’s position on ‘AI’ model training. One is technical, one is social.
The technical one first.
Imagine I rip a DVD and transcode it to MPEG-4 video. This is lossy recompression. The new copy is not identical to the original. It is a derived work. There was no transformative step. Specifically, losing fidelity of reproduction is not a transformative step.
Video CODECs take advantage of redundancy. Simple CODECs build predictive patterns for redundancy in a single frame (for example, is this all one colour or a gradient? Store just that fact not every pixel). More complex ones look at the previous frame and compare it to the current one and use that redundancy. Most modern ones do this in both directions. Effectively, they create a cube of voxels, where one dimension is time, and try to find redundancy in the cube.
For copyright law, this doesn’t matter. Lossy compression is not a transformative step, no matter how complex the compression.
A lot of compression schemes (rarely for video, mostly because it doesn’t make sense for video unless you have a lot) also support special cases for large quantities of redundancy across a data set. For example, if you wanted to compress English Wikipedia with ZSTD, you would use the dictionary mode. It will come up with a list of the words (or even common phrases such as ‘citation needed’) and Huffman encode them so that each page has a short encoding for referencing them. This makes each page smaller than it would be if you compressed it individually.
Compressing Wikipedia like this is not a transformative step.
Now, imagine that you create a video compression CODEC that does this. You buy a copy of every DVD or BluRay disk available and compress them together such that you have a large dictionary of all common compressed sequences. Given a prefix of a film, each film in the input set would be reproducible with some loss of quality (not necessarily the same level of quality). Similarly, if you provided an initial vector that was not something in the training set then you’d get out video that might be similar to one of the inputs, might be similar to many, or might not be obvious to a human is close to either.
This is the crux of the argument. Deep neural networks are functionally equivalent to lossy compression schemes. The inference or generation step in ‘generative AI’ is an initial vector and a random seed that decompresses the data that might be there. If nothing from the training (input) set exactly matches (or if the random seed moves away from that path) then you’ll get something new, possibly something that’s recognisable as a lossily compressed version of the input data.
Note, in particular, that a lot of compression schemes now do take advantage of neural networks. They are one of the most efficient known ways of generating a specialised lossy compression scheme over arbitrary data. The law typically doesn’t care what specific technology an action uses, only about the outcome. In this case, that doesn’t matter: exactly the same underlying technology, used in exactly the same way, covers both ‘AI’ and compression. If one is legal then so is the other because they are the same process.
The EFF’s argument hinges on the idea that this lossy compression is a transformative step. Not only is that an idea that is not supported in case or statute law, there is case law that makes it clear that lossy compression of a work is not transformative.
Their argument would be internally self consistent if it also argued that Netflix does not owe royalties on any of the third-party videos it streams (and that they can buy BluRays on Amazon, recompress them, and then stream them without paying royalties). But there is so much case and statue law that this is not the case that they didn’t make this claim.
Instead, they tried to claim that this is permitted if you call the system ‘AI’ even though it is settled law that it is not permitted if you do not call the system ‘AI’.
Second, the social aspect. Copyright law in the USA draws its legitimacy from this line in the Constitution:
To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries.
So does patent law. Patent law is more of a mess because the USA was founded just before James Watt bribed various MPs to subvert the patent system, but (due to later treaties) the modern US patent system inherits from that subversion (Watt, like Edison, was a bit of a dick).
This intent goes back to the invention of the printing press, where publishers made copies of books in large quantities without paying authors. This removed the incentive to write books. Allowing authors to control distribution rights put that incentive back. This short paragraph covers about a hundred years of the evolution of legal thought in this space, please forgive the many oversimplifications.
The EFF brief references this motivation but twists it. OpenAI and Anthropic are the equivalents of the post-Gutenberg printers. They are taking the work of creative individuals and producing output that competes directly with the products of those authors. This is precisely the situation that copyright law in the USA exists to prevent.
Yet the brief twists this to say that a lossily compressed duplication of original work is actually the kind of creativity that this was intended to cover.
This hinges on the idea that writing news, or other creative works, is a trivial commodity, whereas mechanically compressing them into a system that can lossily reproduce them is a key contribution to society.
Even if I agreed with their other arguments (I do not, I believe that they either misunderstand or deliberately misrepresent the technology and the relationship to other settled law), this is such a profoundly anti-human viewpoint. The idea that human creativity exists to feed poor-quality technological reproductions of that creativity is incompatible with any possible society that I would want to live in and I would struggle to find common ground with people who aim to create such a world.
@david_chisnall I'm not an expert, though I've also had my share of talking to IP lawyers.
What I know is that Intellectual Property doesn't necessarily exist the same way worldwide as it does in the US, e.g. in France, and that it's therefore impossible to reach worldwide conclusions by trying to generalize US IP laws.
Even further, legal systems themselves are very different between countries, so we can't even use the same reasonings on those different laws.
-
Don’t you hate it when you write a long post, proofread it, post it, see it in your feed, and immediately spot a typo in the first line?
(Fixed now, it said ‘firms’ when it should have said ‘forms’)
@david_chisnall Or worse, as you try to delete the offending letter causing the typo focus moves for reasons known only to the universe to another part of the page and the backspace key causes the web browser to navigate to the previous page.
Even worse when this happens as you have nearly completed your long and carefully crafted toot.
I am going to start to using Mousepad to write my toots then copy-pasting them into the compose pane.
-
@david_chisnall Or worse, as you try to delete the offending letter causing the typo focus moves for reasons known only to the universe to another part of the page and the backspace key causes the web browser to navigate to the previous page.
Even worse when this happens as you have nearly completed your long and carefully crafted toot.
I am going to start to using Mousepad to write my toots then copy-pasting them into the compose pane.
@david_chisnall If I ever get my hands on a time machine I will go back in time to find the person/people/team who thought that using the "Backspace" key to navigate back was a good idea and persuade them to change career before they make that fateful decision.
This feature is the bane of my internet life.
-
I am not a lawyer. I have studied IP law in various forms guided by solicitors, barristers, and professors of law of my acquaintance since I was a teenager (including reading through big piles of case histories and commentaries) but not in a formal setting (mostly I learned that I would rather invent the things in the patents than draft the patents, so changed career direction). I’ve also spent a surprising amount of my career talking to copyright and patent lawyers (almost never trademark or trade-secret specialists). I have enough of a lay-person’s understanding of the topic that I was able to spot that a clause in the contract from my US publisher was unenforceable in the state that they claimed jurisdiction because it hinged on an aspect of copyright law that the USA delegates to states and which the state in question did not have relevant laws. Their lawyers subsequently confirmed and fixed this. But I a not lawyer and this is not legal advice.
I have two objections to the EFF’s position on ‘AI’ model training. One is technical, one is social.
The technical one first.
Imagine I rip a DVD and transcode it to MPEG-4 video. This is lossy recompression. The new copy is not identical to the original. It is a derived work. There was no transformative step. Specifically, losing fidelity of reproduction is not a transformative step.
Video CODECs take advantage of redundancy. Simple CODECs build predictive patterns for redundancy in a single frame (for example, is this all one colour or a gradient? Store just that fact not every pixel). More complex ones look at the previous frame and compare it to the current one and use that redundancy. Most modern ones do this in both directions. Effectively, they create a cube of voxels, where one dimension is time, and try to find redundancy in the cube.
For copyright law, this doesn’t matter. Lossy compression is not a transformative step, no matter how complex the compression.
A lot of compression schemes (rarely for video, mostly because it doesn’t make sense for video unless you have a lot) also support special cases for large quantities of redundancy across a data set. For example, if you wanted to compress English Wikipedia with ZSTD, you would use the dictionary mode. It will come up with a list of the words (or even common phrases such as ‘citation needed’) and Huffman encode them so that each page has a short encoding for referencing them. This makes each page smaller than it would be if you compressed it individually.
Compressing Wikipedia like this is not a transformative step.
Now, imagine that you create a video compression CODEC that does this. You buy a copy of every DVD or BluRay disk available and compress them together such that you have a large dictionary of all common compressed sequences. Given a prefix of a film, each film in the input set would be reproducible with some loss of quality (not necessarily the same level of quality). Similarly, if you provided an initial vector that was not something in the training set then you’d get out video that might be similar to one of the inputs, might be similar to many, or might not be obvious to a human is close to either.
This is the crux of the argument. Deep neural networks are functionally equivalent to lossy compression schemes. The inference or generation step in ‘generative AI’ is an initial vector and a random seed that decompresses the data that might be there. If nothing from the training (input) set exactly matches (or if the random seed moves away from that path) then you’ll get something new, possibly something that’s recognisable as a lossily compressed version of the input data.
Note, in particular, that a lot of compression schemes now do take advantage of neural networks. They are one of the most efficient known ways of generating a specialised lossy compression scheme over arbitrary data. The law typically doesn’t care what specific technology an action uses, only about the outcome. In this case, that doesn’t matter: exactly the same underlying technology, used in exactly the same way, covers both ‘AI’ and compression. If one is legal then so is the other because they are the same process.
The EFF’s argument hinges on the idea that this lossy compression is a transformative step. Not only is that an idea that is not supported in case or statute law, there is case law that makes it clear that lossy compression of a work is not transformative.
Their argument would be internally self consistent if it also argued that Netflix does not owe royalties on any of the third-party videos it streams (and that they can buy BluRays on Amazon, recompress them, and then stream them without paying royalties). But there is so much case and statue law that this is not the case that they didn’t make this claim.
Instead, they tried to claim that this is permitted if you call the system ‘AI’ even though it is settled law that it is not permitted if you do not call the system ‘AI’.
Second, the social aspect. Copyright law in the USA draws its legitimacy from this line in the Constitution:
To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries.
So does patent law. Patent law is more of a mess because the USA was founded just before James Watt bribed various MPs to subvert the patent system, but (due to later treaties) the modern US patent system inherits from that subversion (Watt, like Edison, was a bit of a dick).
This intent goes back to the invention of the printing press, where publishers made copies of books in large quantities without paying authors. This removed the incentive to write books. Allowing authors to control distribution rights put that incentive back. This short paragraph covers about a hundred years of the evolution of legal thought in this space, please forgive the many oversimplifications.
The EFF brief references this motivation but twists it. OpenAI and Anthropic are the equivalents of the post-Gutenberg printers. They are taking the work of creative individuals and producing output that competes directly with the products of those authors. This is precisely the situation that copyright law in the USA exists to prevent.
Yet the brief twists this to say that a lossily compressed duplication of original work is actually the kind of creativity that this was intended to cover.
This hinges on the idea that writing news, or other creative works, is a trivial commodity, whereas mechanically compressing them into a system that can lossily reproduce them is a key contribution to society.
Even if I agreed with their other arguments (I do not, I believe that they either misunderstand or deliberately misrepresent the technology and the relationship to other settled law), this is such a profoundly anti-human viewpoint. The idea that human creativity exists to feed poor-quality technological reproductions of that creativity is incompatible with any possible society that I would want to live in and I would struggle to find common ground with people who aim to create such a world.
@david_chisnall I feel conflicted about the technical argument: lossy compression is not a binary but a spectrum. The amount of loss makes the result vary smoothly from "perfect reproduction" to "completely unrecognizable random corruption". IMO, law needs stable categories to determine outcomes, and when none are present it needs to create them, so I wouldn't be surprised if "gzip" landed on the opposite side than "Claude".
-
@david_chisnall If I ever get my hands on a time machine I will go back in time to find the person/people/team who thought that using the "Backspace" key to navigate back was a good idea and persuade them to change career before they make that fateful decision.
This feature is the bane of my internet life.
@the_wub For about ten years, Safari has preserved forms content when you hit back then forwards, so hitting back doesn’t delete the message. It’s saved me a bunch of times.
-
@david_chisnall I feel conflicted about the technical argument: lossy compression is not a binary but a spectrum. The amount of loss makes the result vary smoothly from "perfect reproduction" to "completely unrecognizable random corruption". IMO, law needs stable categories to determine outcomes, and when none are present it needs to create them, so I wouldn't be surprised if "gzip" landed on the opposite side than "Claude".
I’d be happy with an argument (which would be consistent with case law) that said:
If the outputs of a machine-learning system can be something that would meet the existing tests for being a derived work of one of the inputs, then the machine-learning system is a derived work of the training data.
That would be entirely consistent with existing law and would make OpenAI and Anthropic liable for a few trillion dollars of wilful copyright infringement.
-
@david_chisnall I feel conflicted about the technical argument: lossy compression is not a binary but a spectrum. The amount of loss makes the result vary smoothly from "perfect reproduction" to "completely unrecognizable random corruption". IMO, law needs stable categories to determine outcomes, and when none are present it needs to create them, so I wouldn't be surprised if "gzip" landed on the opposite side than "Claude".
@david_chisnall your social objection resonates much more with me (and it's the kind of advocacy I've been promoting so far among peers wrt. AI): no matter what reasonable arguments can be made about the nature of LLMs, their harms to society are factual. Much like with the printing press, the technology just clashed with how the world around it worked, and it damaged the very people it fed from. (cont)
-
@david_chisnall your social objection resonates much more with me (and it's the kind of advocacy I've been promoting so far among peers wrt. AI): no matter what reasonable arguments can be made about the nature of LLMs, their harms to society are factual. Much like with the printing press, the technology just clashed with how the world around it worked, and it damaged the very people it fed from. (cont)
@david_chisnall Hence why copyright was created: an artificial valve to let genuine creativity be profitable while leveraging the advantages of the new technology. In a similar vein, I think trying to shoehorn LLMs into our existing paradigm of "copyright" and "transformative work" is a fruitless endeavor, it just won't fit. We need either an amendment for copyright that answers to LLMs' impact on society and needs no justification, or the replacement of copyright with some different framework.
-
P pelle@veganism.social shared this topic