Skip to content
  • Hjem
  • Seneste
  • Etiketter
  • Populære
  • Verden
  • Bruger
  • Grupper
Temaer
  • Light
  • Brite
  • Cerulean
  • Cosmo
  • Flatly
  • Journal
  • Litera
  • Lumen
  • Lux
  • Materia
  • Minty
  • Morph
  • Pulse
  • Sandstone
  • Simplex
  • Sketchy
  • Spacelab
  • United
  • Yeti
  • Zephyr
  • Dark
  • Cyborg
  • Darkly
  • Quartz
  • Slate
  • Solar
  • Superhero
  • Vapor

  • Default (No Skin)
  • No Skin
Kollaps
FARVEL BIG TECH
  1. Forside
  2. Ikke-kategoriseret
  3. I said I *didn't* think Taylor Lorenz was an industry shill but... wow... now I'm starting to rethink that.

I said I *didn't* think Taylor Lorenz was an industry shill but... wow... now I'm starting to rethink that.

Planlagt Fastgjort Låst Flyttet Ikke-kategoriseret
67 Indlæg 43 Posters 17 Visninger
  • Ældste til nyeste
  • Nyeste til ældste
  • Most Votes
Svar
  • Svar som emne
Login for at svare
Denne tråd er blevet slettet. Kun brugere med emne behandlings privilegier kan se den.
  • mensrea@freeradical.zoneM mensrea@freeradical.zone

    @futurebird there's been more discourse around her before but i forget the context at the moment

    venya@musicians.todayV This user is from outside of this forum
    venya@musicians.todayV This user is from outside of this forum
    venya@musicians.today
    wrote sidst redigeret af
    #19

    @mensrea

    @futurebird

    Same. I have an instinctive "meh" response to her but I do not remember why.

    drukac@mementomori.socialD 1 Reply Last reply
    0
    • futurebird@sauropods.winF futurebird@sauropods.win

      Let me outline my understanding of "LLM Watermarking"

      It is possible to embed special characters and patterns in the output of LLMs that would make text generated by these systems easier to reliably detect. This is mainly being done so that when LLMs scrape the web for new information they can avoid ingesting machine generated content.**

      When you train an LLM on machine generated content it may lead to "model collapse."

      **see next post for correction

      futurebird@sauropods.winF This user is from outside of this forum
      futurebird@sauropods.winF This user is from outside of this forum
      futurebird@sauropods.win
      wrote sidst redigeret af
      #20

      ** I think Peter is more correct. The primary driver making this happen is the EU. As an American I forget that this chaos COULD be regulated a little.

      https://mathstodon.xyz/@pbloem/117133380532237199

      zeb@merovingian.clubZ ericlawton@kolektiva.socialE 2 Replies Last reply
      0
      • futurebird@sauropods.winF futurebird@sauropods.win

        The companies that run LLMs can also use this for PR to calm concerns from the public about the proliferation of such content.

        From a CS perspective I can't think of any way to have a watermark that couldn't be easily defeated through additional processing. LLM dependent people currently take pains to remove em dashes so this new wrinkle has some of them in a panic.

        But, why not just be honest and say you used and LLM? Why so bashful?

        benetherington@spacey.spaceB This user is from outside of this forum
        benetherington@spacey.spaceB This user is from outside of this forum
        benetherington@spacey.space
        wrote sidst redigeret af
        #21

        @futurebird excellent explainer about watermarking in text: https://declaude.org/watermarking/

        androcat@toot.catA 1 Reply Last reply
        0
        • futurebird@sauropods.winF futurebird@sauropods.win

          I said I *didn't* think Taylor Lorenz was an industry shill but... wow... now I'm starting to rethink that.

          You know I might be wrong, but what is wrong with watermarks?

          dalias@hachyderm.ioD This user is from outside of this forum
          dalias@hachyderm.ioD This user is from outside of this forum
          dalias@hachyderm.io
          wrote sidst redigeret af
          #22

          @futurebird Oh, she's been one of the bad ones for a long time...

          1 Reply Last reply
          0
          • futurebird@sauropods.winF futurebird@sauropods.win

            The companies that run LLMs can also use this for PR to calm concerns from the public about the proliferation of such content.

            From a CS perspective I can't think of any way to have a watermark that couldn't be easily defeated through additional processing. LLM dependent people currently take pains to remove em dashes so this new wrinkle has some of them in a panic.

            But, why not just be honest and say you used and LLM? Why so bashful?

            loathsome_dongeater@toots.matapacos.dogL This user is from outside of this forum
            loathsome_dongeater@toots.matapacos.dogL This user is from outside of this forum
            loathsome_dongeater@toots.matapacos.dog
            wrote sidst redigeret af
            #23

            @futurebird the watermarking is also the result of some EU regulations I think.

            Also some LLM enjoyers are a little bit delusional. They will generate a large document using LLMs wholesale and genuinely believe that whatever comes out is the product of their brilliant mind because they composed the prompt.

            toerror@mastodon.gamedev.placeT 1 Reply Last reply
            0
            • benetherington@spacey.spaceB benetherington@spacey.space

              @futurebird excellent explainer about watermarking in text: https://declaude.org/watermarking/

              androcat@toot.catA This user is from outside of this forum
              androcat@toot.catA This user is from outside of this forum
              androcat@toot.cat
              wrote sidst redigeret af
              #24

              @benetherington

              It's just the AI industry pretending they are indispensible in one more way: Detecting their own bullshit.

              Nobody asked them to do this.
              Instead, the EU is demanding that their bullshit be clearly labeled as such.

              @futurebird

              benetherington@spacey.spaceB futurebird@sauropods.winF 2 Replies Last reply
              0
              • futurebird@sauropods.winF futurebird@sauropods.win

                ** I think Peter is more correct. The primary driver making this happen is the EU. As an American I forget that this chaos COULD be regulated a little.

                https://mathstodon.xyz/@pbloem/117133380532237199

                zeb@merovingian.clubZ This user is from outside of this forum
                zeb@merovingian.clubZ This user is from outside of this forum
                zeb@merovingian.club
                wrote sidst redigeret af
                #25

                @futurebird
                The surname Bloem is very famous for having being adopted by jews in central europe when surnames became mandatory.
                Every...fucking...time.

                futurebird@sauropods.winF 1 Reply Last reply
                0
                • pbloem@mathstodon.xyzP pbloem@mathstodon.xyz

                  @futurebird My understanding is that it's mostly happening because the EU requires it.

                  Pre-training datasets are pretty carefully curated and they have strong quality filters, so I don't think there is much of a risk of model collapse in practice.

                  It's also not so much about special characters or word choice as about the choices between equally good options (by the models' own probabilities). Every time you choose between two equally likely options you are essentially "encoding" one bit (i.e. a coinflip) of information. You can use these bits to create a watermark without losing quality.

                  kbm0@mastodon.socialK This user is from outside of this forum
                  kbm0@mastodon.socialK This user is from outside of this forum
                  kbm0@mastodon.social
                  wrote sidst redigeret af
                  #26

                  @pbloem @futurebird No model collapse? Just societal collapse then.

                  1 Reply Last reply
                  0
                  • zeb@merovingian.clubZ zeb@merovingian.club

                    @futurebird
                    The surname Bloem is very famous for having being adopted by jews in central europe when surnames became mandatory.
                    Every...fucking...time.

                    futurebird@sauropods.winF This user is from outside of this forum
                    futurebird@sauropods.winF This user is from outside of this forum
                    futurebird@sauropods.win
                    wrote sidst redigeret af
                    #27

                    @Zeb

                    wtf?

                    1 Reply Last reply
                    0
                    • androcat@toot.catA androcat@toot.cat

                      @benetherington

                      It's just the AI industry pretending they are indispensible in one more way: Detecting their own bullshit.

                      Nobody asked them to do this.
                      Instead, the EU is demanding that their bullshit be clearly labeled as such.

                      @futurebird

                      benetherington@spacey.spaceB This user is from outside of this forum
                      benetherington@spacey.spaceB This user is from outside of this forum
                      benetherington@spacey.space
                      wrote sidst redigeret af
                      #28

                      @androcat @futurebird 100%

                      1 Reply Last reply
                      0
                      • androcat@toot.catA androcat@toot.cat

                        @benetherington

                        It's just the AI industry pretending they are indispensible in one more way: Detecting their own bullshit.

                        Nobody asked them to do this.
                        Instead, the EU is demanding that their bullshit be clearly labeled as such.

                        @futurebird

                        futurebird@sauropods.winF This user is from outside of this forum
                        futurebird@sauropods.winF This user is from outside of this forum
                        futurebird@sauropods.win
                        wrote sidst redigeret af
                        #29

                        @androcat @benetherington

                        We were thinking about what would happen if the whole AI industry vanished and someone mentioned "cyber security" ... but all of the problems they solve are the ones they created.

                        androcat@toot.catA 1 Reply Last reply
                        0
                        • futurebird@sauropods.winF futurebird@sauropods.win

                          The companies that run LLMs can also use this for PR to calm concerns from the public about the proliferation of such content.

                          From a CS perspective I can't think of any way to have a watermark that couldn't be easily defeated through additional processing. LLM dependent people currently take pains to remove em dashes so this new wrinkle has some of them in a panic.

                          But, why not just be honest and say you used and LLM? Why so bashful?

                          becomethewaifu@tech.lgbtB This user is from outside of this forum
                          becomethewaifu@tech.lgbtB This user is from outside of this forum
                          becomethewaifu@tech.lgbt
                          wrote sidst redigeret af
                          #30

                          @futurebird

                          But, why not just be honest and say you used and LLM? Why so bashful?

                          In my experience, it's because they know most people absolutely hate having slop sprayed at their face, so they'll go to great lengths to disguise it in an attempt to "get one over" on them.

                          It resembles the same type of shit with people who "test" someone else's allergies by sneaking stuff into their food: They're trying to "catch them in a lie" with the false assumption that they'll be fine with it if they don't know it's there.

                          And even just that type of behavior is extremely concerning for me, even if it doesn't involve allergens that can Actually Kill People. It shows a clear lack of respect for others, and deserves a swift football kick to the unmentionables and immediate expulsion.

                          futurebird@sauropods.winF abstract_mage_annastasia@tech.lgbtA stumpythemutt@social.linux.pizzaS 3 Replies Last reply
                          0
                          • futurebird@sauropods.winF futurebird@sauropods.win

                            @androcat @benetherington

                            We were thinking about what would happen if the whole AI industry vanished and someone mentioned "cyber security" ... but all of the problems they solve are the ones they created.

                            androcat@toot.catA This user is from outside of this forum
                            androcat@toot.catA This user is from outside of this forum
                            androcat@toot.cat
                            wrote sidst redigeret af
                            #31

                            @futurebird

                            AI industry: "this is our greatest success"
                            The greatest success: Finding [potential bugs] i.e. bits of code, to an extent that is a problem in itself, drowning maintainers in reports.

                            @benetherington

                            1 Reply Last reply
                            0
                            • loathsome_dongeater@toots.matapacos.dogL loathsome_dongeater@toots.matapacos.dog

                              @futurebird the watermarking is also the result of some EU regulations I think.

                              Also some LLM enjoyers are a little bit delusional. They will generate a large document using LLMs wholesale and genuinely believe that whatever comes out is the product of their brilliant mind because they composed the prompt.

                              toerror@mastodon.gamedev.placeT This user is from outside of this forum
                              toerror@mastodon.gamedev.placeT This user is from outside of this forum
                              toerror@mastodon.gamedev.place
                              wrote sidst redigeret af
                              #32

                              @loathsome_dongeater @futurebird There are a lot of people who secretly hate the activity of making things, who have to make things as a chore related to pursuing a career; it's a real gift to them, as is all the aggressive rhetoric in support of the systems normalised by the vendors and their huge bags of PR money.

                              1 Reply Last reply
                              0
                              • becomethewaifu@tech.lgbtB becomethewaifu@tech.lgbt

                                @futurebird

                                But, why not just be honest and say you used and LLM? Why so bashful?

                                In my experience, it's because they know most people absolutely hate having slop sprayed at their face, so they'll go to great lengths to disguise it in an attempt to "get one over" on them.

                                It resembles the same type of shit with people who "test" someone else's allergies by sneaking stuff into their food: They're trying to "catch them in a lie" with the false assumption that they'll be fine with it if they don't know it's there.

                                And even just that type of behavior is extremely concerning for me, even if it doesn't involve allergens that can Actually Kill People. It shows a clear lack of respect for others, and deserves a swift football kick to the unmentionables and immediate expulsion.

                                futurebird@sauropods.winF This user is from outside of this forum
                                futurebird@sauropods.winF This user is from outside of this forum
                                futurebird@sauropods.win
                                wrote sidst redigeret af
                                #33

                                @becomethewaifu

                                Isn't that amazing? Even people who really love boosting AI feel hurt when someone feeds them AI generated content.

                                It's embarrassing to admit to liking or not noticing AI generated content. There was a music video I really enjoyed a few months back and I think it might have AI generated music or AI assisted animation. The creator hasn't been very transparent and everyone who liked it is kind of worried and unhappy.

                                When the same account posted something new I ignored it.

                                futurebird@sauropods.winF brett_e_carlock@mastodon.onlineB 2 Replies Last reply
                                0
                                • futurebird@sauropods.winF futurebird@sauropods.win

                                  @becomethewaifu

                                  Isn't that amazing? Even people who really love boosting AI feel hurt when someone feeds them AI generated content.

                                  It's embarrassing to admit to liking or not noticing AI generated content. There was a music video I really enjoyed a few months back and I think it might have AI generated music or AI assisted animation. The creator hasn't been very transparent and everyone who liked it is kind of worried and unhappy.

                                  When the same account posted something new I ignored it.

                                  futurebird@sauropods.winF This user is from outside of this forum
                                  futurebird@sauropods.winF This user is from outside of this forum
                                  futurebird@sauropods.win
                                  wrote sidst redigeret af
                                  #34

                                  @becomethewaifu

                                  This is the song/creator

                                  I even thought about taking this post down because I don't really know. But, I decided not to and to simply wait to see what was really going on.

                                  I still don't know. But, due to generated content there is a cloud hanging over many creative ventures, fairly or not.

                                  https://sauropods.win/@futurebird/116587364549479132

                                  fandasin@social.linux.pizzaF michaeltbacon@social.coopM 2 Replies Last reply
                                  0
                                  • pbloem@mathstodon.xyzP pbloem@mathstodon.xyz

                                    @futurebird My understanding is that it's mostly happening because the EU requires it.

                                    Pre-training datasets are pretty carefully curated and they have strong quality filters, so I don't think there is much of a risk of model collapse in practice.

                                    It's also not so much about special characters or word choice as about the choices between equally good options (by the models' own probabilities). Every time you choose between two equally likely options you are essentially "encoding" one bit (i.e. a coinflip) of information. You can use these bits to create a watermark without losing quality.

                                    jwcph@helvede.netJ This user is from outside of this forum
                                    jwcph@helvede.netJ This user is from outside of this forum
                                    jwcph@helvede.net
                                    wrote sidst redigeret af
                                    #35

                                    @pbloem @futurebird Yeah, none of that. If datasets were curated we wouldn't get glue on pizza & vendors are already fretting about possible model collapse. Also, "watermarking", which is really fingerprinting, can't possibly work as advertised & even if it did it's trivial to defeat & per OpenAI is only correct a measly 99.9% of the time, which would mean thousands upon thousands of errors every day...

                                    pbloem@mathstodon.xyzP 1 Reply Last reply
                                    0
                                    • futurebird@sauropods.winF futurebird@sauropods.win

                                      The companies that run LLMs can also use this for PR to calm concerns from the public about the proliferation of such content.

                                      From a CS perspective I can't think of any way to have a watermark that couldn't be easily defeated through additional processing. LLM dependent people currently take pains to remove em dashes so this new wrinkle has some of them in a panic.

                                      But, why not just be honest and say you used and LLM? Why so bashful?

                                      david_chisnall@infosec.exchangeD This user is from outside of this forum
                                      david_chisnall@infosec.exchangeD This user is from outside of this forum
                                      david_chisnall@infosec.exchange
                                      wrote sidst redigeret af
                                      #36

                                      @futurebird

                                      From a CS perspective I can't think of any way to have a watermark that couldn't be easily defeated through additional processing

                                      The way that the current ones work is that they tweak the probabilities to generate specific patterns in the output. Where there's an equal probability of two words following another in the raw weights, they'll tune it so that there's a higher probability of one than the other.

                                      If you know the weights and know the biassing factor, you can look at each word pair and see what the probability would be of the model generating that.

                                      This means that the watermark is smeared all over the output. And it's not a binary thing though, each pair of word contributes something to the probability of matching the watermark and looking at the whole thing will give you a probability at the end.

                                      Changing every other word should give you close to a 0% of matching the watermark but, at that point, why bother using the LLM at all? If you're going to change half of the words, you may as well just write them yourself. And that's a problem for people who want to share low-effort slop and pretend to be creative.

                                      ajn142@infosec.exchangeA rysiek@mstdn.socialR 2 Replies Last reply
                                      0
                                      • venya@musicians.todayV venya@musicians.today

                                        @mensrea

                                        @futurebird

                                        Same. I have an instinctive "meh" response to her but I do not remember why.

                                        drukac@mementomori.socialD This user is from outside of this forum
                                        drukac@mementomori.socialD This user is from outside of this forum
                                        drukac@mementomori.social
                                        wrote sidst redigeret af
                                        #37

                                        @venya @mensrea @futurebird

                                        She platformed Chaya Raichik with an interview that really didn't challenge the bigot. I also saw a personal story about Lorenz using dodgy practices at the NYT to build her portfolio by stealing from the sources and beat of others.

                                        1 Reply Last reply
                                        0
                                        • david_chisnall@infosec.exchangeD david_chisnall@infosec.exchange

                                          @futurebird

                                          From a CS perspective I can't think of any way to have a watermark that couldn't be easily defeated through additional processing

                                          The way that the current ones work is that they tweak the probabilities to generate specific patterns in the output. Where there's an equal probability of two words following another in the raw weights, they'll tune it so that there's a higher probability of one than the other.

                                          If you know the weights and know the biassing factor, you can look at each word pair and see what the probability would be of the model generating that.

                                          This means that the watermark is smeared all over the output. And it's not a binary thing though, each pair of word contributes something to the probability of matching the watermark and looking at the whole thing will give you a probability at the end.

                                          Changing every other word should give you close to a 0% of matching the watermark but, at that point, why bother using the LLM at all? If you're going to change half of the words, you may as well just write them yourself. And that's a problem for people who want to share low-effort slop and pretend to be creative.

                                          ajn142@infosec.exchangeA This user is from outside of this forum
                                          ajn142@infosec.exchangeA This user is from outside of this forum
                                          ajn142@infosec.exchange
                                          wrote sidst redigeret af
                                          #38

                                          @david_chisnall @futurebird @404mediaco has a pretty good article OpEd on it as well.

                                          https://www.404media.co/anthropics-text-watermarking-proves-ai-companies-do-not-care-at-all-about-writing/

                                          1 Reply Last reply
                                          0
                                          Svar
                                          • Svar som emne
                                          Login for at svare
                                          • Ældste til nyeste
                                          • Nyeste til ældste
                                          • Most Votes


                                          • Log ind

                                          • Har du ikke en konto? Tilmeld

                                          • Login or register to search.
                                          Powered by NodeBB Contributors
                                          Graciously hosted by data.coop
                                          • First post
                                            Last post
                                          0
                                          • Hjem
                                          • Seneste
                                          • Etiketter
                                          • Populære
                                          • Verden
                                          • Bruger
                                          • Grupper