Skip to content
  • Hjem
  • Seneste
  • Etiketter
  • Populære
  • Verden
  • Bruger
  • Grupper
Temaer
  • Light
  • Brite
  • Cerulean
  • Cosmo
  • Flatly
  • Journal
  • Litera
  • Lumen
  • Lux
  • Materia
  • Minty
  • Morph
  • Pulse
  • Sandstone
  • Simplex
  • Sketchy
  • Spacelab
  • United
  • Yeti
  • Zephyr
  • Dark
  • Cyborg
  • Darkly
  • Quartz
  • Slate
  • Solar
  • Superhero
  • Vapor

  • Default (No Skin)
  • No Skin
Kollaps
FARVEL BIG TECH
  1. Forside
  2. Ikke-kategoriseret
  3. Myth: "Public libraries and book stores have to dispose of books all the time so at least if these companies are scanning the books for AI they will be preserved for history."

Myth: "Public libraries and book stores have to dispose of books all the time so at least if these companies are scanning the books for AI they will be preserved for history."

Planlagt Fastgjort Låst Flyttet Ikke-kategoriseret
43 Indlæg 17 Posters 0 Visninger
  • Ældste til nyeste
  • Nyeste til ældste
  • Most Votes
Svar
  • Svar som emne
Login for at svare
Denne tråd er blevet slettet. Kun brugere med emne behandlings privilegier kan se den.
  • xan@xantronix.socialX xan@xantronix.social

    @futurebird apropos to absolutely nothing do you ever have to restrain yourself from using ant puns in daily speech

    futurebird@sauropods.winF This user is from outside of this forum
    futurebird@sauropods.winF This user is from outside of this forum
    futurebird@sauropods.win
    wrote sidst redigeret af
    #7

    @xan

    no... I just use them. what is the point of holding back?

    😆

    flamecat@bark.lgbtF 1 Reply Last reply
    0
    • futurebird@sauropods.winF futurebird@sauropods.win

      Myth: "Public libraries and book stores have to dispose of books all the time so at least if these companies are scanning the books for AI they will be preserved for history."

      I keep hearing people say this and it's based on a huge misunderstanding of what training data is and how it works.

      When books are scanned to train AI it's not like with "Internet Archive" they aren't preserved. It's not like you could ask the LLM to tell you what was on page 46 of the book and it could look that up.

      1/

      dotsie@mastodon.socialD This user is from outside of this forum
      dotsie@mastodon.socialD This user is from outside of this forum
      dotsie@mastodon.social
      wrote sidst redigeret af
      #8

      @futurebird I truly think if a considered effort about this got out without the typical anti-AI stuff attached (not that I’m saying that’s an incorrect point of view), the general public at large if they were informed about how there is no preservation happening and we have to basically buy information back that was already ours… they won’t like that.

      I think this narrow area, which is critically important to our civilization is something that actually could be changed with amended laws.

      dotsie@mastodon.socialD 1 Reply Last reply
      0
      • sebastian@social.itu.dkS sebastian@social.itu.dk shared this topic
      • futurebird@sauropods.winF futurebird@sauropods.win

        Myth: "Public libraries and book stores have to dispose of books all the time so at least if these companies are scanning the books for AI they will be preserved for history."

        I keep hearing people say this and it's based on a huge misunderstanding of what training data is and how it works.

        When books are scanned to train AI it's not like with "Internet Archive" they aren't preserved. It's not like you could ask the LLM to tell you what was on page 46 of the book and it could look that up.

        1/

        faithfulljohn@mastodon.scotF This user is from outside of this forum
        faithfulljohn@mastodon.scotF This user is from outside of this forum
        faithfulljohn@mastodon.scot
        wrote sidst redigeret af
        #9

        @futurebird 💯 this 😭🤬

        futurebird@sauropods.winF 1 Reply Last reply
        0
        • dotsie@mastodon.socialD dotsie@mastodon.social

          @futurebird I truly think if a considered effort about this got out without the typical anti-AI stuff attached (not that I’m saying that’s an incorrect point of view), the general public at large if they were informed about how there is no preservation happening and we have to basically buy information back that was already ours… they won’t like that.

          I think this narrow area, which is critically important to our civilization is something that actually could be changed with amended laws.

          dotsie@mastodon.socialD This user is from outside of this forum
          dotsie@mastodon.socialD This user is from outside of this forum
          dotsie@mastodon.social
          wrote sidst redigeret af
          #10

          @futurebird like for the most part I feel that it’s just a common sense thing. We shouldn’t be destroying knowledge as a means of intaking it due to a copyright rule. These companies should be required to preserve their own library or something and then if the company dissolves it should go into a public trust. This specific thing is a solvable problem outside of, you know, probably millions of books already being destroyed.

          I just get the sense that this is a flock-like “duh” situation.

          1 Reply Last reply
          0
          • faithfulljohn@mastodon.scotF faithfulljohn@mastodon.scot

            @futurebird 💯 this 😭🤬

            futurebird@sauropods.winF This user is from outside of this forum
            futurebird@sauropods.winF This user is from outside of this forum
            futurebird@sauropods.win
            wrote sidst redigeret af
            #11

            @FaithfullJohn

            The mental image of greedy tech bros feeding precious rare books into the open flaming maw of some monstrous contraption is... kinda valid.

            That is what is happening.

            We could take action to change this. Just as you may be arrested for killing bald eagles, or setting a forest fire the law can protect things that no one person owns. Some of those things are the most valuable things that we have.

            faithfulljohn@mastodon.scotF 1 Reply Last reply
            0
            • futurebird@sauropods.winF futurebird@sauropods.win

              @FaithfullJohn

              The mental image of greedy tech bros feeding precious rare books into the open flaming maw of some monstrous contraption is... kinda valid.

              That is what is happening.

              We could take action to change this. Just as you may be arrested for killing bald eagles, or setting a forest fire the law can protect things that no one person owns. Some of those things are the most valuable things that we have.

              faithfulljohn@mastodon.scotF This user is from outside of this forum
              faithfulljohn@mastodon.scotF This user is from outside of this forum
              faithfulljohn@mastodon.scot
              wrote sidst redigeret af
              #12

              @futurebird

              1 Reply Last reply
              0
              • xan@xantronix.socialX xan@xantronix.social

                @futurebird apropos to absolutely nothing do you ever have to restrain yourself from using ant puns in daily speech

                apostateenglishman@mastodon.worldA This user is from outside of this forum
                apostateenglishman@mastodon.worldA This user is from outside of this forum
                apostateenglishman@mastodon.world
                wrote sidst redigeret af
                #13

                @xan There is hardly any topic for which @futurebird doesn't have an ant-based analogy or metaphor - and on the rare occasions she doesn't, she'll just shamelessly shoehorn in a mention of ants, anyway.

                But that's why we love her. 🥰

                xan@xantronix.socialX 1 Reply Last reply
                0
                • apostateenglishman@mastodon.worldA apostateenglishman@mastodon.world

                  @xan There is hardly any topic for which @futurebird doesn't have an ant-based analogy or metaphor - and on the rare occasions she doesn't, she'll just shamelessly shoehorn in a mention of ants, anyway.

                  But that's why we love her. 🥰

                  xan@xantronix.socialX This user is from outside of this forum
                  xan@xantronix.socialX This user is from outside of this forum
                  xan@xantronix.social
                  wrote sidst redigeret af
                  #14

                  @ApostateEnglishman In meatspace, being as cat-aligned as I am, it is physically impossible for me to not curl my fingers into a paw shape in anticipation of uttering the word "pawsibly" in most interactions. I imagine the case is similar for @futurebird, otherwise I will eat my hat

                  1 Reply Last reply
                  0
                  • futurebird@sauropods.winF futurebird@sauropods.win

                    Is this because it trained on these sites? Probably. But the LLM isn't designed to determine what part of the training data was significant for a particular output. Instead, after writing a response, like an undergraduate student who forgot to cite sources, it looks through the web for the sources that are the most similar to whatever it wrote.

                    If you know about research you know this is ass backwards.

                    But it's the best these kinds of systems can do.

                    3/

                    janbogar@mastodonczech.czJ This user is from outside of this forum
                    janbogar@mastodonczech.czJ This user is from outside of this forum
                    janbogar@mastodonczech.cz
                    wrote sidst redigeret af
                    #15

                    @futurebird I don't think this is true. Chatbot can search web first, put the stuff in its context window (i.e. read it) and generate answer based on that. There is no reason to do it the other way round.

                    It's closer to RAG than generation from memmory. Both Chatgpt and Claude, afaik, can do this.

                    LLM can still misinterpret what it read it or make up nonexistent sources, but I think it's designed better than as you describe it.

                    futurebird@sauropods.winF raymaccarthy@mastodon.ieR 2 Replies Last reply
                    0
                    • janbogar@mastodonczech.czJ janbogar@mastodonczech.cz

                      @futurebird I don't think this is true. Chatbot can search web first, put the stuff in its context window (i.e. read it) and generate answer based on that. There is no reason to do it the other way round.

                      It's closer to RAG than generation from memmory. Both Chatgpt and Claude, afaik, can do this.

                      LLM can still misinterpret what it read it or make up nonexistent sources, but I think it's designed better than as you describe it.

                      futurebird@sauropods.winF This user is from outside of this forum
                      futurebird@sauropods.winF This user is from outside of this forum
                      futurebird@sauropods.win
                      wrote sidst redigeret af
                      #16

                      @janbogar

                      That could also work. But the fundamental issue is you don't really have a way to know what collection of 100s of documents created the text that you are reading.

                      And that's what I'd really like to see.

                      The rarebook that was fed to the machine may never come into this.

                      1 Reply Last reply
                      0
                      • futurebird@sauropods.winF futurebird@sauropods.win

                        Myth: "Public libraries and book stores have to dispose of books all the time so at least if these companies are scanning the books for AI they will be preserved for history."

                        I keep hearing people say this and it's based on a huge misunderstanding of what training data is and how it works.

                        When books are scanned to train AI it's not like with "Internet Archive" they aren't preserved. It's not like you could ask the LLM to tell you what was on page 46 of the book and it could look that up.

                        1/

                        gbsills@social.vivaldi.netG This user is from outside of this forum
                        gbsills@social.vivaldi.netG This user is from outside of this forum
                        gbsills@social.vivaldi.net
                        wrote sidst redigeret af
                        #17

                        @futurebird The problem isn't with the books being disposed off. The problem is with the authors copyright being violated because a machine is using a media created for individual humans. This is clearly a mechanical way to skirt the copyright law.

                        1 Reply Last reply
                        0
                        • futurebird@sauropods.winF futurebird@sauropods.win

                          Myth: "Public libraries and book stores have to dispose of books all the time so at least if these companies are scanning the books for AI they will be preserved for history."

                          I keep hearing people say this and it's based on a huge misunderstanding of what training data is and how it works.

                          When books are scanned to train AI it's not like with "Internet Archive" they aren't preserved. It's not like you could ask the LLM to tell you what was on page 46 of the book and it could look that up.

                          1/

                          markdw@mstdn.socialM This user is from outside of this forum
                          markdw@mstdn.socialM This user is from outside of this forum
                          markdw@mstdn.social
                          wrote sidst redigeret af
                          #18

                          @futurebird would it make sense to remove certain or some pages of books that are being sent to an AI feeding frenzy?

                          futurebird@sauropods.winF 1 Reply Last reply
                          0
                          • markdw@mstdn.socialM markdw@mstdn.social

                            @futurebird would it make sense to remove certain or some pages of books that are being sent to an AI feeding frenzy?

                            futurebird@sauropods.winF This user is from outside of this forum
                            futurebird@sauropods.winF This user is from outside of this forum
                            futurebird@sauropods.win
                            wrote sidst redigeret af
                            #19

                            @MarkDW

                            LLMs work on big numbers and big data. To have an impact with any strategy that targets the training set you need to a spammer.

                            For example when and LLM reproduces entire pages of books verbatim that is often because those pages appeared in the training set multiple times. Probably 100s of times.

                            This is why "training set poisoning" strategies probably won't be as effective as people hope unless we're working on big scales.

                            futurebird@sauropods.winF sundew@beige.partyS 2 Replies Last reply
                            0
                            • futurebird@sauropods.winF futurebird@sauropods.win

                              @MarkDW

                              LLMs work on big numbers and big data. To have an impact with any strategy that targets the training set you need to a spammer.

                              For example when and LLM reproduces entire pages of books verbatim that is often because those pages appeared in the training set multiple times. Probably 100s of times.

                              This is why "training set poisoning" strategies probably won't be as effective as people hope unless we're working on big scales.

                              futurebird@sauropods.winF This user is from outside of this forum
                              futurebird@sauropods.winF This user is from outside of this forum
                              futurebird@sauropods.win
                              wrote sidst redigeret af
                              #20

                              @MarkDW

                              I'm really interested in the customization and curation of training sets.

                              I think there is a lot of potential to make LLMs more useful with bespoke training sets and a totally different kind of output that isn't about chatting but rather about helping to organize and make the document set more useful.

                              raymaccarthy@mastodon.ieR 1 Reply Last reply
                              0
                              • futurebird@sauropods.winF futurebird@sauropods.win

                                @MarkDW

                                LLMs work on big numbers and big data. To have an impact with any strategy that targets the training set you need to a spammer.

                                For example when and LLM reproduces entire pages of books verbatim that is often because those pages appeared in the training set multiple times. Probably 100s of times.

                                This is why "training set poisoning" strategies probably won't be as effective as people hope unless we're working on big scales.

                                sundew@beige.partyS This user is from outside of this forum
                                sundew@beige.partyS This user is from outside of this forum
                                sundew@beige.party
                                wrote sidst redigeret af
                                #21

                                @futurebird @MarkDW Can you explain why you think a large volume is required? I have seen papers to the contrary, eg:
                                https://arxiv.org/html/2510.07192v1

                                1 Reply Last reply
                                0
                                • janbogar@mastodonczech.czJ janbogar@mastodonczech.cz

                                  @futurebird I don't think this is true. Chatbot can search web first, put the stuff in its context window (i.e. read it) and generate answer based on that. There is no reason to do it the other way round.

                                  It's closer to RAG than generation from memmory. Both Chatgpt and Claude, afaik, can do this.

                                  LLM can still misinterpret what it read it or make up nonexistent sources, but I think it's designed better than as you describe it.

                                  raymaccarthy@mastodon.ieR This user is from outside of this forum
                                  raymaccarthy@mastodon.ieR This user is from outside of this forum
                                  raymaccarthy@mastodon.ie
                                  wrote sidst redigeret af
                                  #22

                                  @janbogar @futurebird
                                  Chatbot/ LLM /Generative AI is the worst Web search or document search ever invented.

                                  Just use a real search engine.

                                  futurebird@sauropods.winF 1 Reply Last reply
                                  0
                                  • futurebird@sauropods.winF futurebird@sauropods.win

                                    @MarkDW

                                    I'm really interested in the customization and curation of training sets.

                                    I think there is a lot of potential to make LLMs more useful with bespoke training sets and a totally different kind of output that isn't about chatting but rather about helping to organize and make the document set more useful.

                                    raymaccarthy@mastodon.ieR This user is from outside of this forum
                                    raymaccarthy@mastodon.ieR This user is from outside of this forum
                                    raymaccarthy@mastodon.ie
                                    wrote sidst redigeret af
                                    #23

                                    @futurebird @MarkDW
                                    You don't need an LLM to do that better!

                                    1 Reply Last reply
                                    0
                                    • futurebird@sauropods.winF futurebird@sauropods.win

                                      It is likely that these companies *do* keep copies for future training. But they don't want to share this data or make it searchable by the public.

                                      If someone destroys an old rare book to scan it the public should get a copy of the scan. This is about protecting our culture, history, and heritage.

                                      It's about preserving research and science.

                                      5/5

                                      cford@toot.thoughtworks.comC This user is from outside of this forum
                                      cford@toot.thoughtworks.comC This user is from outside of this forum
                                      cford@toot.thoughtworks.com
                                      wrote sidst redigeret af
                                      #24

                                      @futurebird Indeed. As I understand it, they specifically avoid keeping copies to avoid copyright issues.

                                      raven667@hachyderm.ioR 1 Reply Last reply
                                      0
                                      • jwcph@helvede.netJ jwcph@helvede.net shared this topic
                                      • raymaccarthy@mastodon.ieR raymaccarthy@mastodon.ie

                                        @janbogar @futurebird
                                        Chatbot/ LLM /Generative AI is the worst Web search or document search ever invented.

                                        Just use a real search engine.

                                        futurebird@sauropods.winF This user is from outside of this forum
                                        futurebird@sauropods.winF This user is from outside of this forum
                                        futurebird@sauropods.win
                                        wrote sidst redigeret af
                                        #25

                                        @raymaccarthy @janbogar

                                        I was saying this two years ago however I don't think it's true anymore.

                                        LLMs currently do a better job than web search engines filtering spam and returning the most relevant documents. This is in part because search engines have been made worse. But, we can't just tell people "use real search" when their experiences won't make that advice seem effective.

                                        I agree that these systems should not be search engines.

                                        But that is what they are becoming.

                                        futurebird@sauropods.winF raymaccarthy@mastodon.ieR alec@perkins.pubA 3 Replies Last reply
                                        0
                                        • futurebird@sauropods.winF futurebird@sauropods.win

                                          @raymaccarthy @janbogar

                                          I was saying this two years ago however I don't think it's true anymore.

                                          LLMs currently do a better job than web search engines filtering spam and returning the most relevant documents. This is in part because search engines have been made worse. But, we can't just tell people "use real search" when their experiences won't make that advice seem effective.

                                          I agree that these systems should not be search engines.

                                          But that is what they are becoming.

                                          futurebird@sauropods.winF This user is from outside of this forum
                                          futurebird@sauropods.winF This user is from outside of this forum
                                          futurebird@sauropods.win
                                          wrote sidst redigeret af
                                          #26

                                          @raymaccarthy @janbogar

                                          It's obvious that the makers of consumer LLMs want their systems to become the new primary interface for the web. When I look at what other people I know are doing in their office work, at colleges, they are using LLMs as search engines. Simply to avoid all of the junk and spam most search engines return. Many of them are bashful about it and don't like "AI" in general.

                                          futurebird@sauropods.winF xarvos@outerheaven.clubX 2 Replies Last reply
                                          0
                                          Svar
                                          • Svar som emne
                                          Login for at svare
                                          • Ældste til nyeste
                                          • Nyeste til ældste
                                          • Most Votes


                                          • Log ind

                                          • Har du ikke en konto? Tilmeld

                                          • Login or register to search.
                                          Powered by NodeBB Contributors
                                          Graciously hosted by data.coop
                                          • First post
                                            Last post
                                          0
                                          • Hjem
                                          • Seneste
                                          • Etiketter
                                          • Populære
                                          • Verden
                                          • Bruger
                                          • Grupper