Skip to content
  • Hjem
  • Seneste
  • Etiketter
  • Populære
  • Verden
  • Bruger
  • Grupper
Temaer
  • Light
  • Brite
  • Cerulean
  • Cosmo
  • Flatly
  • Journal
  • Litera
  • Lumen
  • Lux
  • Materia
  • Minty
  • Morph
  • Pulse
  • Sandstone
  • Simplex
  • Sketchy
  • Spacelab
  • United
  • Yeti
  • Zephyr
  • Dark
  • Cyborg
  • Darkly
  • Quartz
  • Slate
  • Solar
  • Superhero
  • Vapor

  • Default (No Skin)
  • No Skin
Kollaps
FARVEL BIG TECH
  1. Forside
  2. Ikke-kategoriseret
  3. I said I *didn't* think Taylor Lorenz was an industry shill but... wow... now I'm starting to rethink that.

I said I *didn't* think Taylor Lorenz was an industry shill but... wow... now I'm starting to rethink that.

Planlagt Fastgjort Låst Flyttet Ikke-kategoriseret
67 Indlæg 43 Posters 17 Visninger
  • Ældste til nyeste
  • Nyeste til ældste
  • Most Votes
Svar
  • Svar som emne
Login for at svare
Denne tråd er blevet slettet. Kun brugere med emne behandlings privilegier kan se den.
  • futurebird@sauropods.winF futurebird@sauropods.win

    I said I *didn't* think Taylor Lorenz was an industry shill but... wow... now I'm starting to rethink that.

    You know I might be wrong, but what is wrong with watermarks?

    mensrea@freeradical.zoneM This user is from outside of this forum
    mensrea@freeradical.zoneM This user is from outside of this forum
    mensrea@freeradical.zone
    wrote sidst redigeret af
    #3

    @futurebird there's been more discourse around her before but i forget the context at the moment

    venya@musicians.todayV gryphonmyers@mastodon.socialG 2 Replies Last reply
    0
    • futurebird@sauropods.winF futurebird@sauropods.win

      I said I *didn't* think Taylor Lorenz was an industry shill but... wow... now I'm starting to rethink that.

      You know I might be wrong, but what is wrong with watermarks?

      futurebird@sauropods.winF This user is from outside of this forum
      futurebird@sauropods.winF This user is from outside of this forum
      futurebird@sauropods.win
      wrote sidst redigeret af
      #4

      Let me outline my understanding of "LLM Watermarking"

      It is possible to embed special characters and patterns in the output of LLMs that would make text generated by these systems easier to reliably detect. This is mainly being done so that when LLMs scrape the web for new information they can avoid ingesting machine generated content.**

      When you train an LLM on machine generated content it may lead to "model collapse."

      **see next post for correction

      futurebird@sauropods.winF pbloem@mathstodon.xyzP aplundell@timeloop.cafeA 4 Replies Last reply
      0
      • futurebird@sauropods.winF futurebird@sauropods.win

        I said I *didn't* think Taylor Lorenz was an industry shill but... wow... now I'm starting to rethink that.

        You know I might be wrong, but what is wrong with watermarks?

        platypus@glammr.usP This user is from outside of this forum
        platypus@glammr.usP This user is from outside of this forum
        platypus@glammr.us
        wrote sidst redigeret af
        #5

        @futurebird so I find myself generally in the pro column, but I can see two reasons against:

        first, it does provide a certain level of tracking, not unlike that would just put in into our printers to identify exactly which printer something came from.

        I can also see it as a real boon for the companies so that they don’t choke on re-uploading their output and get worse, which is a known problem

        platypus@glammr.usP 1 Reply Last reply
        0
        • futurebird@sauropods.winF futurebird@sauropods.win

          I said I *didn't* think Taylor Lorenz was an industry shill but... wow... now I'm starting to rethink that.

          You know I might be wrong, but what is wrong with watermarks?

          inkyschwartz@mastodon.socialI This user is from outside of this forum
          inkyschwartz@mastodon.socialI This user is from outside of this forum
          inkyschwartz@mastodon.social
          wrote sidst redigeret af
          #6

          @futurebird Not sure on any of that but Watermarking seems to be a different issue.
          https://www.404media.co/anthropics-text-watermarking-proves-ai-companies-do-not-care-at-all-about-writing/

          1 Reply Last reply
          0
          • futurebird@sauropods.winF futurebird@sauropods.win

            I said I *didn't* think Taylor Lorenz was an industry shill but... wow... now I'm starting to rethink that.

            You know I might be wrong, but what is wrong with watermarks?

            H This user is from outside of this forum
            H This user is from outside of this forum
            hypolite@friendica.mrpetovan.com
            wrote sidst redigeret af
            #7
            @futurebird What’s really wrong with AI watermarks for text is that AI companies believe words are interchangeable for a given meaning. And that only them can theoretically verify the watermark, which feels like a conflict of interest?
            Q silverwizard@convenient.emailS 2 Replies Last reply
            0
            • futurebird@sauropods.winF futurebird@sauropods.win

              Let me outline my understanding of "LLM Watermarking"

              It is possible to embed special characters and patterns in the output of LLMs that would make text generated by these systems easier to reliably detect. This is mainly being done so that when LLMs scrape the web for new information they can avoid ingesting machine generated content.**

              When you train an LLM on machine generated content it may lead to "model collapse."

              **see next post for correction

              futurebird@sauropods.winF This user is from outside of this forum
              futurebird@sauropods.winF This user is from outside of this forum
              futurebird@sauropods.win
              wrote sidst redigeret af
              #8

              The companies that run LLMs can also use this for PR to calm concerns from the public about the proliferation of such content.

              From a CS perspective I can't think of any way to have a watermark that couldn't be easily defeated through additional processing. LLM dependent people currently take pains to remove em dashes so this new wrinkle has some of them in a panic.

              But, why not just be honest and say you used and LLM? Why so bashful?

              androcat@toot.catA ben@mastodon.lubar.meB inkyschwartz@mastodon.socialI alec@perkins.pubA benetherington@spacey.spaceB 9 Replies Last reply
              0
              • platypus@glammr.usP platypus@glammr.us

                @futurebird so I find myself generally in the pro column, but I can see two reasons against:

                first, it does provide a certain level of tracking, not unlike that would just put in into our printers to identify exactly which printer something came from.

                I can also see it as a real boon for the companies so that they don’t choke on re-uploading their output and get worse, which is a known problem

                platypus@glammr.usP This user is from outside of this forum
                platypus@glammr.usP This user is from outside of this forum
                platypus@glammr.us
                wrote sidst redigeret af
                #9

                @futurebird in the first case I am extremely anti-printer finger marking so I can be sympathetic to it… But since AI is used for widescale fraud or deceptive practices ot is not based on the government wanting to identify a person but based on supporting individuals trying to make sense of this extremely slopified world

                but the second, I think it’s the reason that the AI companies are willing to do it without a bigger fuss. They realized they can use it to eliminate bad training data.

                1 Reply Last reply
                0
                • futurebird@sauropods.winF futurebird@sauropods.win

                  The companies that run LLMs can also use this for PR to calm concerns from the public about the proliferation of such content.

                  From a CS perspective I can't think of any way to have a watermark that couldn't be easily defeated through additional processing. LLM dependent people currently take pains to remove em dashes so this new wrinkle has some of them in a panic.

                  But, why not just be honest and say you used and LLM? Why so bashful?

                  androcat@toot.catA This user is from outside of this forum
                  androcat@toot.catA This user is from outside of this forum
                  androcat@toot.cat
                  wrote sidst redigeret af
                  #10

                  @futurebird What I expect from AI companies is Malicious Compliance.
                  And Malicious Non-Compliance.

                  1 Reply Last reply
                  0
                  • futurebird@sauropods.winF futurebird@sauropods.win

                    I said I *didn't* think Taylor Lorenz was an industry shill but... wow... now I'm starting to rethink that.

                    You know I might be wrong, but what is wrong with watermarks?

                    mrtnsnp@mastodon.socialM This user is from outside of this forum
                    mrtnsnp@mastodon.socialM This user is from outside of this forum
                    mrtnsnp@mastodon.social
                    wrote sidst redigeret af
                    #11

                    @futurebird @kim_harding Even before it becomes easy to detect the watermark itself, the noise that AI users make about it is an excellent detector for AI content.

                    1 Reply Last reply
                    0
                    • futurebird@sauropods.winF futurebird@sauropods.win

                      The companies that run LLMs can also use this for PR to calm concerns from the public about the proliferation of such content.

                      From a CS perspective I can't think of any way to have a watermark that couldn't be easily defeated through additional processing. LLM dependent people currently take pains to remove em dashes so this new wrinkle has some of them in a panic.

                      But, why not just be honest and say you used and LLM? Why so bashful?

                      ben@mastodon.lubar.meB This user is from outside of this forum
                      ben@mastodon.lubar.meB This user is from outside of this forum
                      ben@mastodon.lubar.me
                      wrote sidst redigeret af
                      #12

                      @futurebird I remember when Genius.com accused Google of plagiarizing their song lyrics and as evidence showed that the weird spacing and capitalization they had intentionally added to a song were reproduced 1:1 on Google's uncredited copy

                      so I think the idea is that people will be too lazy to try to defeat the "watermarking"? but that's clearly already not the case for LLM slop

                      androcat@toot.catA ben@mastodon.lubar.meB 2 Replies Last reply
                      0
                      • futurebird@sauropods.winF futurebird@sauropods.win

                        The companies that run LLMs can also use this for PR to calm concerns from the public about the proliferation of such content.

                        From a CS perspective I can't think of any way to have a watermark that couldn't be easily defeated through additional processing. LLM dependent people currently take pains to remove em dashes so this new wrinkle has some of them in a panic.

                        But, why not just be honest and say you used and LLM? Why so bashful?

                        inkyschwartz@mastodon.socialI This user is from outside of this forum
                        inkyschwartz@mastodon.socialI This user is from outside of this forum
                        inkyschwartz@mastodon.social
                        wrote sidst redigeret af
                        #13

                        @futurebird Honestly it would make detecting AI plagerism/assigment writing easier.

                        1 Reply Last reply
                        0
                        • futurebird@sauropods.winF futurebird@sauropods.win

                          Let me outline my understanding of "LLM Watermarking"

                          It is possible to embed special characters and patterns in the output of LLMs that would make text generated by these systems easier to reliably detect. This is mainly being done so that when LLMs scrape the web for new information they can avoid ingesting machine generated content.**

                          When you train an LLM on machine generated content it may lead to "model collapse."

                          **see next post for correction

                          pbloem@mathstodon.xyzP This user is from outside of this forum
                          pbloem@mathstodon.xyzP This user is from outside of this forum
                          pbloem@mathstodon.xyz
                          wrote sidst redigeret af
                          #14

                          @futurebird My understanding is that it's mostly happening because the EU requires it.

                          Pre-training datasets are pretty carefully curated and they have strong quality filters, so I don't think there is much of a risk of model collapse in practice.

                          It's also not so much about special characters or word choice as about the choices between equally good options (by the models' own probabilities). Every time you choose between two equally likely options you are essentially "encoding" one bit (i.e. a coinflip) of information. You can use these bits to create a watermark without losing quality.

                          graham_knapp@hachyderm.ioG kbm0@mastodon.socialK jwcph@helvede.netJ 3 Replies Last reply
                          0
                          • pbloem@mathstodon.xyzP pbloem@mathstodon.xyz

                            @futurebird My understanding is that it's mostly happening because the EU requires it.

                            Pre-training datasets are pretty carefully curated and they have strong quality filters, so I don't think there is much of a risk of model collapse in practice.

                            It's also not so much about special characters or word choice as about the choices between equally good options (by the models' own probabilities). Every time you choose between two equally likely options you are essentially "encoding" one bit (i.e. a coinflip) of information. You can use these bits to create a watermark without losing quality.

                            graham_knapp@hachyderm.ioG This user is from outside of this forum
                            graham_knapp@hachyderm.ioG This user is from outside of this forum
                            graham_knapp@hachyderm.io
                            wrote sidst redigeret af
                            #15

                            @pbloem @futurebird Yes it seems very much driven by the EU AI act, presumably for general transparency rather than specifically about AI training.

                            1 Reply Last reply
                            0
                            • ben@mastodon.lubar.meB ben@mastodon.lubar.me

                              @futurebird I remember when Genius.com accused Google of plagiarizing their song lyrics and as evidence showed that the weird spacing and capitalization they had intentionally added to a song were reproduced 1:1 on Google's uncredited copy

                              so I think the idea is that people will be too lazy to try to defeat the "watermarking"? but that's clearly already not the case for LLM slop

                              androcat@toot.catA This user is from outside of this forum
                              androcat@toot.catA This user is from outside of this forum
                              androcat@toot.cat
                              wrote sidst redigeret af
                              #16

                              @ben

                              People using slop have self-selected to be too lazy to defeat anything.

                              But "just use our secret decoder to prove that someone used our secret coder" is a self-serving non-solution.

                              @futurebird

                              1 Reply Last reply
                              0
                              • ben@mastodon.lubar.meB ben@mastodon.lubar.me

                                @futurebird I remember when Genius.com accused Google of plagiarizing their song lyrics and as evidence showed that the weird spacing and capitalization they had intentionally added to a song were reproduced 1:1 on Google's uncredited copy

                                so I think the idea is that people will be too lazy to try to defeat the "watermarking"? but that's clearly already not the case for LLM slop

                                ben@mastodon.lubar.meB This user is from outside of this forum
                                ben@mastodon.lubar.meB This user is from outside of this forum
                                ben@mastodon.lubar.me
                                wrote sidst redigeret af
                                #17

                                @futurebird like if there's one discussion about watermarks I've seen on the internet ever, it's artists saying "the malicious person cropped out my signature"

                                the one thing I know about watermarks is that they are easily defeated by determined assholes

                                1 Reply Last reply
                                0
                                • futurebird@sauropods.winF futurebird@sauropods.win

                                  The companies that run LLMs can also use this for PR to calm concerns from the public about the proliferation of such content.

                                  From a CS perspective I can't think of any way to have a watermark that couldn't be easily defeated through additional processing. LLM dependent people currently take pains to remove em dashes so this new wrinkle has some of them in a panic.

                                  But, why not just be honest and say you used and LLM? Why so bashful?

                                  alec@perkins.pubA This user is from outside of this forum
                                  alec@perkins.pubA This user is from outside of this forum
                                  alec@perkins.pub
                                  wrote sidst redigeret af
                                  #18

                                  @futurebird the way certain people are reacting to the watermarks—which have been in place for a little while already—is very telling.

                                  1 Reply Last reply
                                  0
                                  • mensrea@freeradical.zoneM mensrea@freeradical.zone

                                    @futurebird there's been more discourse around her before but i forget the context at the moment

                                    venya@musicians.todayV This user is from outside of this forum
                                    venya@musicians.todayV This user is from outside of this forum
                                    venya@musicians.today
                                    wrote sidst redigeret af
                                    #19

                                    @mensrea

                                    @futurebird

                                    Same. I have an instinctive "meh" response to her but I do not remember why.

                                    drukac@mementomori.socialD 1 Reply Last reply
                                    0
                                    • futurebird@sauropods.winF futurebird@sauropods.win

                                      Let me outline my understanding of "LLM Watermarking"

                                      It is possible to embed special characters and patterns in the output of LLMs that would make text generated by these systems easier to reliably detect. This is mainly being done so that when LLMs scrape the web for new information they can avoid ingesting machine generated content.**

                                      When you train an LLM on machine generated content it may lead to "model collapse."

                                      **see next post for correction

                                      futurebird@sauropods.winF This user is from outside of this forum
                                      futurebird@sauropods.winF This user is from outside of this forum
                                      futurebird@sauropods.win
                                      wrote sidst redigeret af
                                      #20

                                      ** I think Peter is more correct. The primary driver making this happen is the EU. As an American I forget that this chaos COULD be regulated a little.

                                      https://mathstodon.xyz/@pbloem/117133380532237199

                                      zeb@merovingian.clubZ ericlawton@kolektiva.socialE 2 Replies Last reply
                                      0
                                      • futurebird@sauropods.winF futurebird@sauropods.win

                                        The companies that run LLMs can also use this for PR to calm concerns from the public about the proliferation of such content.

                                        From a CS perspective I can't think of any way to have a watermark that couldn't be easily defeated through additional processing. LLM dependent people currently take pains to remove em dashes so this new wrinkle has some of them in a panic.

                                        But, why not just be honest and say you used and LLM? Why so bashful?

                                        benetherington@spacey.spaceB This user is from outside of this forum
                                        benetherington@spacey.spaceB This user is from outside of this forum
                                        benetherington@spacey.space
                                        wrote sidst redigeret af
                                        #21

                                        @futurebird excellent explainer about watermarking in text: https://declaude.org/watermarking/

                                        androcat@toot.catA 1 Reply Last reply
                                        0
                                        • futurebird@sauropods.winF futurebird@sauropods.win

                                          I said I *didn't* think Taylor Lorenz was an industry shill but... wow... now I'm starting to rethink that.

                                          You know I might be wrong, but what is wrong with watermarks?

                                          dalias@hachyderm.ioD This user is from outside of this forum
                                          dalias@hachyderm.ioD This user is from outside of this forum
                                          dalias@hachyderm.io
                                          wrote sidst redigeret af
                                          #22

                                          @futurebird Oh, she's been one of the bad ones for a long time...

                                          1 Reply Last reply
                                          0
                                          Svar
                                          • Svar som emne
                                          Login for at svare
                                          • Ældste til nyeste
                                          • Nyeste til ældste
                                          • Most Votes


                                          • Log ind

                                          • Har du ikke en konto? Tilmeld

                                          • Login or register to search.
                                          Powered by NodeBB Contributors
                                          Graciously hosted by data.coop
                                          • First post
                                            Last post
                                          0
                                          • Hjem
                                          • Seneste
                                          • Etiketter
                                          • Populære
                                          • Verden
                                          • Bruger
                                          • Grupper