I said I *didn't* think Taylor Lorenz was an industry shill but... wow... now I'm starting to rethink that.
-
I said I *didn't* think Taylor Lorenz was an industry shill but... wow... now I'm starting to rethink that.
You know I might be wrong, but what is wrong with watermarks?
@futurebird @kim_harding Even before it becomes easy to detect the watermark itself, the noise that AI users make about it is an excellent detector for AI content.
-
The companies that run LLMs can also use this for PR to calm concerns from the public about the proliferation of such content.
From a CS perspective I can't think of any way to have a watermark that couldn't be easily defeated through additional processing. LLM dependent people currently take pains to remove em dashes so this new wrinkle has some of them in a panic.
But, why not just be honest and say you used and LLM? Why so bashful?
@futurebird I remember when Genius.com accused Google of plagiarizing their song lyrics and as evidence showed that the weird spacing and capitalization they had intentionally added to a song were reproduced 1:1 on Google's uncredited copy
so I think the idea is that people will be too lazy to try to defeat the "watermarking"? but that's clearly already not the case for LLM slop
-
The companies that run LLMs can also use this for PR to calm concerns from the public about the proliferation of such content.
From a CS perspective I can't think of any way to have a watermark that couldn't be easily defeated through additional processing. LLM dependent people currently take pains to remove em dashes so this new wrinkle has some of them in a panic.
But, why not just be honest and say you used and LLM? Why so bashful?
@futurebird Honestly it would make detecting AI plagerism/assigment writing easier.
-
Let me outline my understanding of "LLM Watermarking"
It is possible to embed special characters and patterns in the output of LLMs that would make text generated by these systems easier to reliably detect. This is mainly being done so that when LLMs scrape the web for new information they can avoid ingesting machine generated content.**
When you train an LLM on machine generated content it may lead to "model collapse."
**see next post for correction
@futurebird My understanding is that it's mostly happening because the EU requires it.
Pre-training datasets are pretty carefully curated and they have strong quality filters, so I don't think there is much of a risk of model collapse in practice.
It's also not so much about special characters or word choice as about the choices between equally good options (by the models' own probabilities). Every time you choose between two equally likely options you are essentially "encoding" one bit (i.e. a coinflip) of information. You can use these bits to create a watermark without losing quality.
-
@futurebird My understanding is that it's mostly happening because the EU requires it.
Pre-training datasets are pretty carefully curated and they have strong quality filters, so I don't think there is much of a risk of model collapse in practice.
It's also not so much about special characters or word choice as about the choices between equally good options (by the models' own probabilities). Every time you choose between two equally likely options you are essentially "encoding" one bit (i.e. a coinflip) of information. You can use these bits to create a watermark without losing quality.
@pbloem @futurebird Yes it seems very much driven by the EU AI act, presumably for general transparency rather than specifically about AI training.
-
@futurebird I remember when Genius.com accused Google of plagiarizing their song lyrics and as evidence showed that the weird spacing and capitalization they had intentionally added to a song were reproduced 1:1 on Google's uncredited copy
so I think the idea is that people will be too lazy to try to defeat the "watermarking"? but that's clearly already not the case for LLM slop
People using slop have self-selected to be too lazy to defeat anything.
But "just use our secret decoder to prove that someone used our secret coder" is a self-serving non-solution.
-
@futurebird I remember when Genius.com accused Google of plagiarizing their song lyrics and as evidence showed that the weird spacing and capitalization they had intentionally added to a song were reproduced 1:1 on Google's uncredited copy
so I think the idea is that people will be too lazy to try to defeat the "watermarking"? but that's clearly already not the case for LLM slop
@futurebird like if there's one discussion about watermarks I've seen on the internet ever, it's artists saying "the malicious person cropped out my signature"
the one thing I know about watermarks is that they are easily defeated by determined assholes
-
The companies that run LLMs can also use this for PR to calm concerns from the public about the proliferation of such content.
From a CS perspective I can't think of any way to have a watermark that couldn't be easily defeated through additional processing. LLM dependent people currently take pains to remove em dashes so this new wrinkle has some of them in a panic.
But, why not just be honest and say you used and LLM? Why so bashful?
@futurebird the way certain people are reacting to the watermarks—which have been in place for a little while already—is very telling.
-
@futurebird there's been more discourse around her before but i forget the context at the moment
-
Let me outline my understanding of "LLM Watermarking"
It is possible to embed special characters and patterns in the output of LLMs that would make text generated by these systems easier to reliably detect. This is mainly being done so that when LLMs scrape the web for new information they can avoid ingesting machine generated content.**
When you train an LLM on machine generated content it may lead to "model collapse."
**see next post for correction
** I think Peter is more correct. The primary driver making this happen is the EU. As an American I forget that this chaos COULD be regulated a little.
-
The companies that run LLMs can also use this for PR to calm concerns from the public about the proliferation of such content.
From a CS perspective I can't think of any way to have a watermark that couldn't be easily defeated through additional processing. LLM dependent people currently take pains to remove em dashes so this new wrinkle has some of them in a panic.
But, why not just be honest and say you used and LLM? Why so bashful?
@futurebird excellent explainer about watermarking in text: https://declaude.org/watermarking/
-
I said I *didn't* think Taylor Lorenz was an industry shill but... wow... now I'm starting to rethink that.
You know I might be wrong, but what is wrong with watermarks?
@futurebird Oh, she's been one of the bad ones for a long time...
-
The companies that run LLMs can also use this for PR to calm concerns from the public about the proliferation of such content.
From a CS perspective I can't think of any way to have a watermark that couldn't be easily defeated through additional processing. LLM dependent people currently take pains to remove em dashes so this new wrinkle has some of them in a panic.
But, why not just be honest and say you used and LLM? Why so bashful?
@futurebird the watermarking is also the result of some EU regulations I think.
Also some LLM enjoyers are a little bit delusional. They will generate a large document using LLMs wholesale and genuinely believe that whatever comes out is the product of their brilliant mind because they composed the prompt.
-
@futurebird excellent explainer about watermarking in text: https://declaude.org/watermarking/
It's just the AI industry pretending they are indispensible in one more way: Detecting their own bullshit.
Nobody asked them to do this.
Instead, the EU is demanding that their bullshit be clearly labeled as such. -
** I think Peter is more correct. The primary driver making this happen is the EU. As an American I forget that this chaos COULD be regulated a little.
@futurebird
The surname Bloem is very famous for having being adopted by jews in central europe when surnames became mandatory.
Every...fucking...time. -
@futurebird My understanding is that it's mostly happening because the EU requires it.
Pre-training datasets are pretty carefully curated and they have strong quality filters, so I don't think there is much of a risk of model collapse in practice.
It's also not so much about special characters or word choice as about the choices between equally good options (by the models' own probabilities). Every time you choose between two equally likely options you are essentially "encoding" one bit (i.e. a coinflip) of information. You can use these bits to create a watermark without losing quality.
@pbloem @futurebird No model collapse? Just societal collapse then.
-
@futurebird
The surname Bloem is very famous for having being adopted by jews in central europe when surnames became mandatory.
Every...fucking...time.wtf?
-
It's just the AI industry pretending they are indispensible in one more way: Detecting their own bullshit.
Nobody asked them to do this.
Instead, the EU is demanding that their bullshit be clearly labeled as such.@androcat @futurebird 100%
-
It's just the AI industry pretending they are indispensible in one more way: Detecting their own bullshit.
Nobody asked them to do this.
Instead, the EU is demanding that their bullshit be clearly labeled as such.We were thinking about what would happen if the whole AI industry vanished and someone mentioned "cyber security" ... but all of the problems they solve are the ones they created.
-
The companies that run LLMs can also use this for PR to calm concerns from the public about the proliferation of such content.
From a CS perspective I can't think of any way to have a watermark that couldn't be easily defeated through additional processing. LLM dependent people currently take pains to remove em dashes so this new wrinkle has some of them in a panic.
But, why not just be honest and say you used and LLM? Why so bashful?
But, why not just be honest and say you used and LLM? Why so bashful?
In my experience, it's because they know most people absolutely hate having slop sprayed at their face, so they'll go to great lengths to disguise it in an attempt to "get one over" on them.
It resembles the same type of shit with people who "test" someone else's allergies by sneaking stuff into their food: They're trying to "catch them in a lie" with the false assumption that they'll be fine with it if they don't know it's there.
And even just that type of behavior is extremely concerning for me, even if it doesn't involve allergens that can Actually Kill People. It shows a clear lack of respect for others, and deserves a swift football kick to the unmentionables and immediate expulsion.