Skip to content
  • Hjem
  • Seneste
  • Etiketter
  • Populære
  • Verden
  • Bruger
  • Grupper
Temaer
  • Light
  • Brite
  • Cerulean
  • Cosmo
  • Flatly
  • Journal
  • Litera
  • Lumen
  • Lux
  • Materia
  • Minty
  • Morph
  • Pulse
  • Sandstone
  • Simplex
  • Sketchy
  • Spacelab
  • United
  • Yeti
  • Zephyr
  • Dark
  • Cyborg
  • Darkly
  • Quartz
  • Slate
  • Solar
  • Superhero
  • Vapor

  • Default (No Skin)
  • No Skin
Kollaps
FARVEL BIG TECH
  1. Forside
  2. Ikke-kategoriseret
  3. Myth: "Public libraries and book stores have to dispose of books all the time so at least if these companies are scanning the books for AI they will be preserved for history."

Myth: "Public libraries and book stores have to dispose of books all the time so at least if these companies are scanning the books for AI they will be preserved for history."

Planlagt Fastgjort Låst Flyttet Ikke-kategoriseret
43 Indlæg 17 Posters 0 Visninger
  • Ældste til nyeste
  • Nyeste til ældste
  • Most Votes
Svar
  • Svar som emne
Login for at svare
Denne tråd er blevet slettet. Kun brugere med emne behandlings privilegier kan se den.
  • futurebird@sauropods.winF futurebird@sauropods.win

    Myth: "Public libraries and book stores have to dispose of books all the time so at least if these companies are scanning the books for AI they will be preserved for history."

    I keep hearing people say this and it's based on a huge misunderstanding of what training data is and how it works.

    When books are scanned to train AI it's not like with "Internet Archive" they aren't preserved. It's not like you could ask the LLM to tell you what was on page 46 of the book and it could look that up.

    1/

    futurebird@sauropods.winF This user is from outside of this forum
    futurebird@sauropods.winF This user is from outside of this forum
    futurebird@sauropods.win
    wrote sidst redigeret af
    #2

    Recently LLMs have gotten better about providing source links with the information they share. But wait until I tell you how those source links are generated! It's not what you might think.

    Let's say that I ask chatGPT "Which Carpenter ants live in NYC?"

    chatGPT will answer this question and provide excerpts from wikipedia, antWeb and antWiki to support parts of this answer.

    2/

    futurebird@sauropods.winF xan@xantronix.socialX 2 Replies Last reply
    0
    • futurebird@sauropods.winF futurebird@sauropods.win

      Recently LLMs have gotten better about providing source links with the information they share. But wait until I tell you how those source links are generated! It's not what you might think.

      Let's say that I ask chatGPT "Which Carpenter ants live in NYC?"

      chatGPT will answer this question and provide excerpts from wikipedia, antWeb and antWiki to support parts of this answer.

      2/

      futurebird@sauropods.winF This user is from outside of this forum
      futurebird@sauropods.winF This user is from outside of this forum
      futurebird@sauropods.win
      wrote sidst redigeret af
      #3

      Is this because it trained on these sites? Probably. But the LLM isn't designed to determine what part of the training data was significant for a particular output. Instead, after writing a response, like an undergraduate student who forgot to cite sources, it looks through the web for the sources that are the most similar to whatever it wrote.

      If you know about research you know this is ass backwards.

      But it's the best these kinds of systems can do.

      3/

      futurebird@sauropods.winF janbogar@mastodonczech.czJ 2 Replies Last reply
      1
      0
      • futurebird@sauropods.winF futurebird@sauropods.win

        Recently LLMs have gotten better about providing source links with the information they share. But wait until I tell you how those source links are generated! It's not what you might think.

        Let's say that I ask chatGPT "Which Carpenter ants live in NYC?"

        chatGPT will answer this question and provide excerpts from wikipedia, antWeb and antWiki to support parts of this answer.

        2/

        xan@xantronix.socialX This user is from outside of this forum
        xan@xantronix.socialX This user is from outside of this forum
        xan@xantronix.social
        wrote sidst redigeret af
        #4

        @futurebird apropos to absolutely nothing do you ever have to restrain yourself from using ant puns in daily speech

        futurebird@sauropods.winF apostateenglishman@mastodon.worldA 2 Replies Last reply
        0
        • futurebird@sauropods.winF futurebird@sauropods.win

          Is this because it trained on these sites? Probably. But the LLM isn't designed to determine what part of the training data was significant for a particular output. Instead, after writing a response, like an undergraduate student who forgot to cite sources, it looks through the web for the sources that are the most similar to whatever it wrote.

          If you know about research you know this is ass backwards.

          But it's the best these kinds of systems can do.

          3/

          futurebird@sauropods.winF This user is from outside of this forum
          futurebird@sauropods.winF This user is from outside of this forum
          futurebird@sauropods.win
          wrote sidst redigeret af
          #5

          In reality the sentence structure and word choice of any LLM response are dependent on millions of documents.

          This is why it will sometimes make a statement and provide a supporting link that says the exact opposite. The statements are based on massive statistical trends in the training data. The fake "sources" are added after.

          Unless a company chooses to keep a copy of the training data it is gone forever.

          4/

          futurebird@sauropods.winF 1 Reply Last reply
          0
          • futurebird@sauropods.winF futurebird@sauropods.win

            In reality the sentence structure and word choice of any LLM response are dependent on millions of documents.

            This is why it will sometimes make a statement and provide a supporting link that says the exact opposite. The statements are based on massive statistical trends in the training data. The fake "sources" are added after.

            Unless a company chooses to keep a copy of the training data it is gone forever.

            4/

            futurebird@sauropods.winF This user is from outside of this forum
            futurebird@sauropods.winF This user is from outside of this forum
            futurebird@sauropods.win
            wrote sidst redigeret af
            #6

            It is likely that these companies *do* keep copies for future training. But they don't want to share this data or make it searchable by the public.

            If someone destroys an old rare book to scan it the public should get a copy of the scan. This is about protecting our culture, history, and heritage.

            It's about preserving research and science.

            5/5

            cford@toot.thoughtworks.comC 1 Reply Last reply
            1
            0
            • xan@xantronix.socialX xan@xantronix.social

              @futurebird apropos to absolutely nothing do you ever have to restrain yourself from using ant puns in daily speech

              futurebird@sauropods.winF This user is from outside of this forum
              futurebird@sauropods.winF This user is from outside of this forum
              futurebird@sauropods.win
              wrote sidst redigeret af
              #7

              @xan

              no... I just use them. what is the point of holding back?

              😆

              flamecat@bark.lgbtF 1 Reply Last reply
              0
              • futurebird@sauropods.winF futurebird@sauropods.win

                Myth: "Public libraries and book stores have to dispose of books all the time so at least if these companies are scanning the books for AI they will be preserved for history."

                I keep hearing people say this and it's based on a huge misunderstanding of what training data is and how it works.

                When books are scanned to train AI it's not like with "Internet Archive" they aren't preserved. It's not like you could ask the LLM to tell you what was on page 46 of the book and it could look that up.

                1/

                dotsie@mastodon.socialD This user is from outside of this forum
                dotsie@mastodon.socialD This user is from outside of this forum
                dotsie@mastodon.social
                wrote sidst redigeret af
                #8

                @futurebird I truly think if a considered effort about this got out without the typical anti-AI stuff attached (not that I’m saying that’s an incorrect point of view), the general public at large if they were informed about how there is no preservation happening and we have to basically buy information back that was already ours… they won’t like that.

                I think this narrow area, which is critically important to our civilization is something that actually could be changed with amended laws.

                dotsie@mastodon.socialD 1 Reply Last reply
                0
                • sebastian@social.itu.dkS sebastian@social.itu.dk shared this topic
                • futurebird@sauropods.winF futurebird@sauropods.win

                  Myth: "Public libraries and book stores have to dispose of books all the time so at least if these companies are scanning the books for AI they will be preserved for history."

                  I keep hearing people say this and it's based on a huge misunderstanding of what training data is and how it works.

                  When books are scanned to train AI it's not like with "Internet Archive" they aren't preserved. It's not like you could ask the LLM to tell you what was on page 46 of the book and it could look that up.

                  1/

                  faithfulljohn@mastodon.scotF This user is from outside of this forum
                  faithfulljohn@mastodon.scotF This user is from outside of this forum
                  faithfulljohn@mastodon.scot
                  wrote sidst redigeret af
                  #9

                  @futurebird 💯 this 😭🤬

                  futurebird@sauropods.winF 1 Reply Last reply
                  0
                  • dotsie@mastodon.socialD dotsie@mastodon.social

                    @futurebird I truly think if a considered effort about this got out without the typical anti-AI stuff attached (not that I’m saying that’s an incorrect point of view), the general public at large if they were informed about how there is no preservation happening and we have to basically buy information back that was already ours… they won’t like that.

                    I think this narrow area, which is critically important to our civilization is something that actually could be changed with amended laws.

                    dotsie@mastodon.socialD This user is from outside of this forum
                    dotsie@mastodon.socialD This user is from outside of this forum
                    dotsie@mastodon.social
                    wrote sidst redigeret af
                    #10

                    @futurebird like for the most part I feel that it’s just a common sense thing. We shouldn’t be destroying knowledge as a means of intaking it due to a copyright rule. These companies should be required to preserve their own library or something and then if the company dissolves it should go into a public trust. This specific thing is a solvable problem outside of, you know, probably millions of books already being destroyed.

                    I just get the sense that this is a flock-like “duh” situation.

                    1 Reply Last reply
                    0
                    • faithfulljohn@mastodon.scotF faithfulljohn@mastodon.scot

                      @futurebird 💯 this 😭🤬

                      futurebird@sauropods.winF This user is from outside of this forum
                      futurebird@sauropods.winF This user is from outside of this forum
                      futurebird@sauropods.win
                      wrote sidst redigeret af
                      #11

                      @FaithfullJohn

                      The mental image of greedy tech bros feeding precious rare books into the open flaming maw of some monstrous contraption is... kinda valid.

                      That is what is happening.

                      We could take action to change this. Just as you may be arrested for killing bald eagles, or setting a forest fire the law can protect things that no one person owns. Some of those things are the most valuable things that we have.

                      faithfulljohn@mastodon.scotF 1 Reply Last reply
                      0
                      • futurebird@sauropods.winF futurebird@sauropods.win

                        @FaithfullJohn

                        The mental image of greedy tech bros feeding precious rare books into the open flaming maw of some monstrous contraption is... kinda valid.

                        That is what is happening.

                        We could take action to change this. Just as you may be arrested for killing bald eagles, or setting a forest fire the law can protect things that no one person owns. Some of those things are the most valuable things that we have.

                        faithfulljohn@mastodon.scotF This user is from outside of this forum
                        faithfulljohn@mastodon.scotF This user is from outside of this forum
                        faithfulljohn@mastodon.scot
                        wrote sidst redigeret af
                        #12

                        @futurebird

                        1 Reply Last reply
                        0
                        • xan@xantronix.socialX xan@xantronix.social

                          @futurebird apropos to absolutely nothing do you ever have to restrain yourself from using ant puns in daily speech

                          apostateenglishman@mastodon.worldA This user is from outside of this forum
                          apostateenglishman@mastodon.worldA This user is from outside of this forum
                          apostateenglishman@mastodon.world
                          wrote sidst redigeret af
                          #13

                          @xan There is hardly any topic for which @futurebird doesn't have an ant-based analogy or metaphor - and on the rare occasions she doesn't, she'll just shamelessly shoehorn in a mention of ants, anyway.

                          But that's why we love her. 🥰

                          xan@xantronix.socialX 1 Reply Last reply
                          0
                          • apostateenglishman@mastodon.worldA apostateenglishman@mastodon.world

                            @xan There is hardly any topic for which @futurebird doesn't have an ant-based analogy or metaphor - and on the rare occasions she doesn't, she'll just shamelessly shoehorn in a mention of ants, anyway.

                            But that's why we love her. 🥰

                            xan@xantronix.socialX This user is from outside of this forum
                            xan@xantronix.socialX This user is from outside of this forum
                            xan@xantronix.social
                            wrote sidst redigeret af
                            #14

                            @ApostateEnglishman In meatspace, being as cat-aligned as I am, it is physically impossible for me to not curl my fingers into a paw shape in anticipation of uttering the word "pawsibly" in most interactions. I imagine the case is similar for @futurebird, otherwise I will eat my hat

                            1 Reply Last reply
                            0
                            • futurebird@sauropods.winF futurebird@sauropods.win

                              Is this because it trained on these sites? Probably. But the LLM isn't designed to determine what part of the training data was significant for a particular output. Instead, after writing a response, like an undergraduate student who forgot to cite sources, it looks through the web for the sources that are the most similar to whatever it wrote.

                              If you know about research you know this is ass backwards.

                              But it's the best these kinds of systems can do.

                              3/

                              janbogar@mastodonczech.czJ This user is from outside of this forum
                              janbogar@mastodonczech.czJ This user is from outside of this forum
                              janbogar@mastodonczech.cz
                              wrote sidst redigeret af
                              #15

                              @futurebird I don't think this is true. Chatbot can search web first, put the stuff in its context window (i.e. read it) and generate answer based on that. There is no reason to do it the other way round.

                              It's closer to RAG than generation from memmory. Both Chatgpt and Claude, afaik, can do this.

                              LLM can still misinterpret what it read it or make up nonexistent sources, but I think it's designed better than as you describe it.

                              futurebird@sauropods.winF raymaccarthy@mastodon.ieR 2 Replies Last reply
                              0
                              • janbogar@mastodonczech.czJ janbogar@mastodonczech.cz

                                @futurebird I don't think this is true. Chatbot can search web first, put the stuff in its context window (i.e. read it) and generate answer based on that. There is no reason to do it the other way round.

                                It's closer to RAG than generation from memmory. Both Chatgpt and Claude, afaik, can do this.

                                LLM can still misinterpret what it read it or make up nonexistent sources, but I think it's designed better than as you describe it.

                                futurebird@sauropods.winF This user is from outside of this forum
                                futurebird@sauropods.winF This user is from outside of this forum
                                futurebird@sauropods.win
                                wrote sidst redigeret af
                                #16

                                @janbogar

                                That could also work. But the fundamental issue is you don't really have a way to know what collection of 100s of documents created the text that you are reading.

                                And that's what I'd really like to see.

                                The rarebook that was fed to the machine may never come into this.

                                1 Reply Last reply
                                0
                                • futurebird@sauropods.winF futurebird@sauropods.win

                                  Myth: "Public libraries and book stores have to dispose of books all the time so at least if these companies are scanning the books for AI they will be preserved for history."

                                  I keep hearing people say this and it's based on a huge misunderstanding of what training data is and how it works.

                                  When books are scanned to train AI it's not like with "Internet Archive" they aren't preserved. It's not like you could ask the LLM to tell you what was on page 46 of the book and it could look that up.

                                  1/

                                  gbsills@social.vivaldi.netG This user is from outside of this forum
                                  gbsills@social.vivaldi.netG This user is from outside of this forum
                                  gbsills@social.vivaldi.net
                                  wrote sidst redigeret af
                                  #17

                                  @futurebird The problem isn't with the books being disposed off. The problem is with the authors copyright being violated because a machine is using a media created for individual humans. This is clearly a mechanical way to skirt the copyright law.

                                  1 Reply Last reply
                                  0
                                  • futurebird@sauropods.winF futurebird@sauropods.win

                                    Myth: "Public libraries and book stores have to dispose of books all the time so at least if these companies are scanning the books for AI they will be preserved for history."

                                    I keep hearing people say this and it's based on a huge misunderstanding of what training data is and how it works.

                                    When books are scanned to train AI it's not like with "Internet Archive" they aren't preserved. It's not like you could ask the LLM to tell you what was on page 46 of the book and it could look that up.

                                    1/

                                    markdw@mstdn.socialM This user is from outside of this forum
                                    markdw@mstdn.socialM This user is from outside of this forum
                                    markdw@mstdn.social
                                    wrote sidst redigeret af
                                    #18

                                    @futurebird would it make sense to remove certain or some pages of books that are being sent to an AI feeding frenzy?

                                    futurebird@sauropods.winF 1 Reply Last reply
                                    0
                                    • markdw@mstdn.socialM markdw@mstdn.social

                                      @futurebird would it make sense to remove certain or some pages of books that are being sent to an AI feeding frenzy?

                                      futurebird@sauropods.winF This user is from outside of this forum
                                      futurebird@sauropods.winF This user is from outside of this forum
                                      futurebird@sauropods.win
                                      wrote sidst redigeret af
                                      #19

                                      @MarkDW

                                      LLMs work on big numbers and big data. To have an impact with any strategy that targets the training set you need to a spammer.

                                      For example when and LLM reproduces entire pages of books verbatim that is often because those pages appeared in the training set multiple times. Probably 100s of times.

                                      This is why "training set poisoning" strategies probably won't be as effective as people hope unless we're working on big scales.

                                      futurebird@sauropods.winF sundew@beige.partyS 2 Replies Last reply
                                      0
                                      • futurebird@sauropods.winF futurebird@sauropods.win

                                        @MarkDW

                                        LLMs work on big numbers and big data. To have an impact with any strategy that targets the training set you need to a spammer.

                                        For example when and LLM reproduces entire pages of books verbatim that is often because those pages appeared in the training set multiple times. Probably 100s of times.

                                        This is why "training set poisoning" strategies probably won't be as effective as people hope unless we're working on big scales.

                                        futurebird@sauropods.winF This user is from outside of this forum
                                        futurebird@sauropods.winF This user is from outside of this forum
                                        futurebird@sauropods.win
                                        wrote sidst redigeret af
                                        #20

                                        @MarkDW

                                        I'm really interested in the customization and curation of training sets.

                                        I think there is a lot of potential to make LLMs more useful with bespoke training sets and a totally different kind of output that isn't about chatting but rather about helping to organize and make the document set more useful.

                                        raymaccarthy@mastodon.ieR 1 Reply Last reply
                                        0
                                        • futurebird@sauropods.winF futurebird@sauropods.win

                                          @MarkDW

                                          LLMs work on big numbers and big data. To have an impact with any strategy that targets the training set you need to a spammer.

                                          For example when and LLM reproduces entire pages of books verbatim that is often because those pages appeared in the training set multiple times. Probably 100s of times.

                                          This is why "training set poisoning" strategies probably won't be as effective as people hope unless we're working on big scales.

                                          sundew@beige.partyS This user is from outside of this forum
                                          sundew@beige.partyS This user is from outside of this forum
                                          sundew@beige.party
                                          wrote sidst redigeret af
                                          #21

                                          @futurebird @MarkDW Can you explain why you think a large volume is required? I have seen papers to the contrary, eg:
                                          https://arxiv.org/html/2510.07192v1

                                          1 Reply Last reply
                                          0
                                          Svar
                                          • Svar som emne
                                          Login for at svare
                                          • Ældste til nyeste
                                          • Nyeste til ældste
                                          • Most Votes


                                          • Log ind

                                          • Har du ikke en konto? Tilmeld

                                          • Login or register to search.
                                          Powered by NodeBB Contributors
                                          Graciously hosted by data.coop
                                          • First post
                                            Last post
                                          0
                                          • Hjem
                                          • Seneste
                                          • Etiketter
                                          • Populære
                                          • Verden
                                          • Bruger
                                          • Grupper