Skip to content
  • Hjem
  • Seneste
  • Etiketter
  • Populære
  • Verden
  • Bruger
  • Grupper
Temaer
  • Light
  • Brite
  • Cerulean
  • Cosmo
  • Flatly
  • Journal
  • Litera
  • Lumen
  • Lux
  • Materia
  • Minty
  • Morph
  • Pulse
  • Sandstone
  • Simplex
  • Sketchy
  • Spacelab
  • United
  • Yeti
  • Zephyr
  • Dark
  • Cyborg
  • Darkly
  • Quartz
  • Slate
  • Solar
  • Superhero
  • Vapor

  • Default (No Skin)
  • No Skin
Kollaps
FARVEL BIG TECH
  1. Forside
  2. Ikke-kategoriseret
  3. Myth: "Public libraries and book stores have to dispose of books all the time so at least if these companies are scanning the books for AI they will be preserved for history."

Myth: "Public libraries and book stores have to dispose of books all the time so at least if these companies are scanning the books for AI they will be preserved for history."

Planlagt Fastgjort Låst Flyttet Ikke-kategoriseret
43 Indlæg 17 Posters 0 Visninger
  • Ældste til nyeste
  • Nyeste til ældste
  • Most Votes
Svar
  • Svar som emne
Login for at svare
Denne tråd er blevet slettet. Kun brugere med emne behandlings privilegier kan se den.
  • faithfulljohn@mastodon.scotF faithfulljohn@mastodon.scot

    @futurebird 💯 this 😭🤬

    futurebird@sauropods.winF This user is from outside of this forum
    futurebird@sauropods.winF This user is from outside of this forum
    futurebird@sauropods.win
    wrote sidst redigeret af
    #11

    @FaithfullJohn

    The mental image of greedy tech bros feeding precious rare books into the open flaming maw of some monstrous contraption is... kinda valid.

    That is what is happening.

    We could take action to change this. Just as you may be arrested for killing bald eagles, or setting a forest fire the law can protect things that no one person owns. Some of those things are the most valuable things that we have.

    faithfulljohn@mastodon.scotF 1 Reply Last reply
    0
    • futurebird@sauropods.winF futurebird@sauropods.win

      @FaithfullJohn

      The mental image of greedy tech bros feeding precious rare books into the open flaming maw of some monstrous contraption is... kinda valid.

      That is what is happening.

      We could take action to change this. Just as you may be arrested for killing bald eagles, or setting a forest fire the law can protect things that no one person owns. Some of those things are the most valuable things that we have.

      faithfulljohn@mastodon.scotF This user is from outside of this forum
      faithfulljohn@mastodon.scotF This user is from outside of this forum
      faithfulljohn@mastodon.scot
      wrote sidst redigeret af
      #12

      @futurebird

      1 Reply Last reply
      0
      • xan@xantronix.socialX xan@xantronix.social

        @futurebird apropos to absolutely nothing do you ever have to restrain yourself from using ant puns in daily speech

        apostateenglishman@mastodon.worldA This user is from outside of this forum
        apostateenglishman@mastodon.worldA This user is from outside of this forum
        apostateenglishman@mastodon.world
        wrote sidst redigeret af
        #13

        @xan There is hardly any topic for which @futurebird doesn't have an ant-based analogy or metaphor - and on the rare occasions she doesn't, she'll just shamelessly shoehorn in a mention of ants, anyway.

        But that's why we love her. 🥰

        xan@xantronix.socialX 1 Reply Last reply
        0
        • apostateenglishman@mastodon.worldA apostateenglishman@mastodon.world

          @xan There is hardly any topic for which @futurebird doesn't have an ant-based analogy or metaphor - and on the rare occasions she doesn't, she'll just shamelessly shoehorn in a mention of ants, anyway.

          But that's why we love her. 🥰

          xan@xantronix.socialX This user is from outside of this forum
          xan@xantronix.socialX This user is from outside of this forum
          xan@xantronix.social
          wrote sidst redigeret af
          #14

          @ApostateEnglishman In meatspace, being as cat-aligned as I am, it is physically impossible for me to not curl my fingers into a paw shape in anticipation of uttering the word "pawsibly" in most interactions. I imagine the case is similar for @futurebird, otherwise I will eat my hat

          1 Reply Last reply
          0
          • futurebird@sauropods.winF futurebird@sauropods.win

            Is this because it trained on these sites? Probably. But the LLM isn't designed to determine what part of the training data was significant for a particular output. Instead, after writing a response, like an undergraduate student who forgot to cite sources, it looks through the web for the sources that are the most similar to whatever it wrote.

            If you know about research you know this is ass backwards.

            But it's the best these kinds of systems can do.

            3/

            janbogar@mastodonczech.czJ This user is from outside of this forum
            janbogar@mastodonczech.czJ This user is from outside of this forum
            janbogar@mastodonczech.cz
            wrote sidst redigeret af
            #15

            @futurebird I don't think this is true. Chatbot can search web first, put the stuff in its context window (i.e. read it) and generate answer based on that. There is no reason to do it the other way round.

            It's closer to RAG than generation from memmory. Both Chatgpt and Claude, afaik, can do this.

            LLM can still misinterpret what it read it or make up nonexistent sources, but I think it's designed better than as you describe it.

            futurebird@sauropods.winF raymaccarthy@mastodon.ieR 2 Replies Last reply
            0
            • janbogar@mastodonczech.czJ janbogar@mastodonczech.cz

              @futurebird I don't think this is true. Chatbot can search web first, put the stuff in its context window (i.e. read it) and generate answer based on that. There is no reason to do it the other way round.

              It's closer to RAG than generation from memmory. Both Chatgpt and Claude, afaik, can do this.

              LLM can still misinterpret what it read it or make up nonexistent sources, but I think it's designed better than as you describe it.

              futurebird@sauropods.winF This user is from outside of this forum
              futurebird@sauropods.winF This user is from outside of this forum
              futurebird@sauropods.win
              wrote sidst redigeret af
              #16

              @janbogar

              That could also work. But the fundamental issue is you don't really have a way to know what collection of 100s of documents created the text that you are reading.

              And that's what I'd really like to see.

              The rarebook that was fed to the machine may never come into this.

              1 Reply Last reply
              0
              • futurebird@sauropods.winF futurebird@sauropods.win

                Myth: "Public libraries and book stores have to dispose of books all the time so at least if these companies are scanning the books for AI they will be preserved for history."

                I keep hearing people say this and it's based on a huge misunderstanding of what training data is and how it works.

                When books are scanned to train AI it's not like with "Internet Archive" they aren't preserved. It's not like you could ask the LLM to tell you what was on page 46 of the book and it could look that up.

                1/

                gbsills@social.vivaldi.netG This user is from outside of this forum
                gbsills@social.vivaldi.netG This user is from outside of this forum
                gbsills@social.vivaldi.net
                wrote sidst redigeret af
                #17

                @futurebird The problem isn't with the books being disposed off. The problem is with the authors copyright being violated because a machine is using a media created for individual humans. This is clearly a mechanical way to skirt the copyright law.

                1 Reply Last reply
                0
                • futurebird@sauropods.winF futurebird@sauropods.win

                  Myth: "Public libraries and book stores have to dispose of books all the time so at least if these companies are scanning the books for AI they will be preserved for history."

                  I keep hearing people say this and it's based on a huge misunderstanding of what training data is and how it works.

                  When books are scanned to train AI it's not like with "Internet Archive" they aren't preserved. It's not like you could ask the LLM to tell you what was on page 46 of the book and it could look that up.

                  1/

                  markdw@mstdn.socialM This user is from outside of this forum
                  markdw@mstdn.socialM This user is from outside of this forum
                  markdw@mstdn.social
                  wrote sidst redigeret af
                  #18

                  @futurebird would it make sense to remove certain or some pages of books that are being sent to an AI feeding frenzy?

                  futurebird@sauropods.winF 1 Reply Last reply
                  0
                  • markdw@mstdn.socialM markdw@mstdn.social

                    @futurebird would it make sense to remove certain or some pages of books that are being sent to an AI feeding frenzy?

                    futurebird@sauropods.winF This user is from outside of this forum
                    futurebird@sauropods.winF This user is from outside of this forum
                    futurebird@sauropods.win
                    wrote sidst redigeret af
                    #19

                    @MarkDW

                    LLMs work on big numbers and big data. To have an impact with any strategy that targets the training set you need to a spammer.

                    For example when and LLM reproduces entire pages of books verbatim that is often because those pages appeared in the training set multiple times. Probably 100s of times.

                    This is why "training set poisoning" strategies probably won't be as effective as people hope unless we're working on big scales.

                    futurebird@sauropods.winF sundew@beige.partyS 2 Replies Last reply
                    0
                    • futurebird@sauropods.winF futurebird@sauropods.win

                      @MarkDW

                      LLMs work on big numbers and big data. To have an impact with any strategy that targets the training set you need to a spammer.

                      For example when and LLM reproduces entire pages of books verbatim that is often because those pages appeared in the training set multiple times. Probably 100s of times.

                      This is why "training set poisoning" strategies probably won't be as effective as people hope unless we're working on big scales.

                      futurebird@sauropods.winF This user is from outside of this forum
                      futurebird@sauropods.winF This user is from outside of this forum
                      futurebird@sauropods.win
                      wrote sidst redigeret af
                      #20

                      @MarkDW

                      I'm really interested in the customization and curation of training sets.

                      I think there is a lot of potential to make LLMs more useful with bespoke training sets and a totally different kind of output that isn't about chatting but rather about helping to organize and make the document set more useful.

                      raymaccarthy@mastodon.ieR 1 Reply Last reply
                      0
                      • futurebird@sauropods.winF futurebird@sauropods.win

                        @MarkDW

                        LLMs work on big numbers and big data. To have an impact with any strategy that targets the training set you need to a spammer.

                        For example when and LLM reproduces entire pages of books verbatim that is often because those pages appeared in the training set multiple times. Probably 100s of times.

                        This is why "training set poisoning" strategies probably won't be as effective as people hope unless we're working on big scales.

                        sundew@beige.partyS This user is from outside of this forum
                        sundew@beige.partyS This user is from outside of this forum
                        sundew@beige.party
                        wrote sidst redigeret af
                        #21

                        @futurebird @MarkDW Can you explain why you think a large volume is required? I have seen papers to the contrary, eg:
                        https://arxiv.org/html/2510.07192v1

                        1 Reply Last reply
                        0
                        • janbogar@mastodonczech.czJ janbogar@mastodonczech.cz

                          @futurebird I don't think this is true. Chatbot can search web first, put the stuff in its context window (i.e. read it) and generate answer based on that. There is no reason to do it the other way round.

                          It's closer to RAG than generation from memmory. Both Chatgpt and Claude, afaik, can do this.

                          LLM can still misinterpret what it read it or make up nonexistent sources, but I think it's designed better than as you describe it.

                          raymaccarthy@mastodon.ieR This user is from outside of this forum
                          raymaccarthy@mastodon.ieR This user is from outside of this forum
                          raymaccarthy@mastodon.ie
                          wrote sidst redigeret af
                          #22

                          @janbogar @futurebird
                          Chatbot/ LLM /Generative AI is the worst Web search or document search ever invented.

                          Just use a real search engine.

                          futurebird@sauropods.winF 1 Reply Last reply
                          0
                          • futurebird@sauropods.winF futurebird@sauropods.win

                            @MarkDW

                            I'm really interested in the customization and curation of training sets.

                            I think there is a lot of potential to make LLMs more useful with bespoke training sets and a totally different kind of output that isn't about chatting but rather about helping to organize and make the document set more useful.

                            raymaccarthy@mastodon.ieR This user is from outside of this forum
                            raymaccarthy@mastodon.ieR This user is from outside of this forum
                            raymaccarthy@mastodon.ie
                            wrote sidst redigeret af
                            #23

                            @futurebird @MarkDW
                            You don't need an LLM to do that better!

                            1 Reply Last reply
                            0
                            • futurebird@sauropods.winF futurebird@sauropods.win

                              It is likely that these companies *do* keep copies for future training. But they don't want to share this data or make it searchable by the public.

                              If someone destroys an old rare book to scan it the public should get a copy of the scan. This is about protecting our culture, history, and heritage.

                              It's about preserving research and science.

                              5/5

                              cford@toot.thoughtworks.comC This user is from outside of this forum
                              cford@toot.thoughtworks.comC This user is from outside of this forum
                              cford@toot.thoughtworks.com
                              wrote sidst redigeret af
                              #24

                              @futurebird Indeed. As I understand it, they specifically avoid keeping copies to avoid copyright issues.

                              raven667@hachyderm.ioR 1 Reply Last reply
                              0
                              • jwcph@helvede.netJ jwcph@helvede.net shared this topic
                              • raymaccarthy@mastodon.ieR raymaccarthy@mastodon.ie

                                @janbogar @futurebird
                                Chatbot/ LLM /Generative AI is the worst Web search or document search ever invented.

                                Just use a real search engine.

                                futurebird@sauropods.winF This user is from outside of this forum
                                futurebird@sauropods.winF This user is from outside of this forum
                                futurebird@sauropods.win
                                wrote sidst redigeret af
                                #25

                                @raymaccarthy @janbogar

                                I was saying this two years ago however I don't think it's true anymore.

                                LLMs currently do a better job than web search engines filtering spam and returning the most relevant documents. This is in part because search engines have been made worse. But, we can't just tell people "use real search" when their experiences won't make that advice seem effective.

                                I agree that these systems should not be search engines.

                                But that is what they are becoming.

                                futurebird@sauropods.winF raymaccarthy@mastodon.ieR alec@perkins.pubA 3 Replies Last reply
                                0
                                • futurebird@sauropods.winF futurebird@sauropods.win

                                  @raymaccarthy @janbogar

                                  I was saying this two years ago however I don't think it's true anymore.

                                  LLMs currently do a better job than web search engines filtering spam and returning the most relevant documents. This is in part because search engines have been made worse. But, we can't just tell people "use real search" when their experiences won't make that advice seem effective.

                                  I agree that these systems should not be search engines.

                                  But that is what they are becoming.

                                  futurebird@sauropods.winF This user is from outside of this forum
                                  futurebird@sauropods.winF This user is from outside of this forum
                                  futurebird@sauropods.win
                                  wrote sidst redigeret af
                                  #26

                                  @raymaccarthy @janbogar

                                  It's obvious that the makers of consumer LLMs want their systems to become the new primary interface for the web. When I look at what other people I know are doing in their office work, at colleges, they are using LLMs as search engines. Simply to avoid all of the junk and spam most search engines return. Many of them are bashful about it and don't like "AI" in general.

                                  futurebird@sauropods.winF xarvos@outerheaven.clubX 2 Replies Last reply
                                  0
                                  • futurebird@sauropods.winF futurebird@sauropods.win

                                    @raymaccarthy @janbogar

                                    I was saying this two years ago however I don't think it's true anymore.

                                    LLMs currently do a better job than web search engines filtering spam and returning the most relevant documents. This is in part because search engines have been made worse. But, we can't just tell people "use real search" when their experiences won't make that advice seem effective.

                                    I agree that these systems should not be search engines.

                                    But that is what they are becoming.

                                    raymaccarthy@mastodon.ieR This user is from outside of this forum
                                    raymaccarthy@mastodon.ieR This user is from outside of this forum
                                    raymaccarthy@mastodon.ie
                                    wrote sidst redigeret af
                                    #27

                                    @futurebird @janbogar
                                    "LLMs currently do a better job than web search engines filtering spam and returning the most relevant documents."

                                    I've a nice bridge over the Shannon you might like to buy!

                                    futurebird@sauropods.winF janbogar@mastodonczech.czJ 2 Replies Last reply
                                    0
                                    • futurebird@sauropods.winF futurebird@sauropods.win

                                      @raymaccarthy @janbogar

                                      It's obvious that the makers of consumer LLMs want their systems to become the new primary interface for the web. When I look at what other people I know are doing in their office work, at colleges, they are using LLMs as search engines. Simply to avoid all of the junk and spam most search engines return. Many of them are bashful about it and don't like "AI" in general.

                                      futurebird@sauropods.winF This user is from outside of this forum
                                      futurebird@sauropods.winF This user is from outside of this forum
                                      futurebird@sauropods.win
                                      wrote sidst redigeret af
                                      #28

                                      @raymaccarthy @janbogar

                                      This is another stage of enclosure and control of information. Obviously some things will not be included in these systems. But, ever since Google started messing with their search algorithm and including ads we were already on this road.

                                      This technology is very much overhyped but I think it's critical to recognize how it is really being used and where it's really experiencing growth and becoming popular.

                                      futurebird@sauropods.winF raymaccarthy@mastodon.ieR 2 Replies Last reply
                                      0
                                      • futurebird@sauropods.winF futurebird@sauropods.win

                                        @raymaccarthy @janbogar

                                        This is another stage of enclosure and control of information. Obviously some things will not be included in these systems. But, ever since Google started messing with their search algorithm and including ads we were already on this road.

                                        This technology is very much overhyped but I think it's critical to recognize how it is really being used and where it's really experiencing growth and becoming popular.

                                        futurebird@sauropods.winF This user is from outside of this forum
                                        futurebird@sauropods.winF This user is from outside of this forum
                                        futurebird@sauropods.win
                                        wrote sidst redigeret af
                                        #29

                                        @raymaccarthy @janbogar

                                        The real success of "AI" has mostly been in:

                                        * spamming
                                        * bad therapy
                                        * search

                                        Search has been broken and bad for years. Many people have sounded the alarm. Here is the "solution" and it is worse than the problem. It's a disaster.

                                        chiraag@mastodon.onlineC 1 Reply Last reply
                                        0
                                        • raymaccarthy@mastodon.ieR raymaccarthy@mastodon.ie

                                          @futurebird @janbogar
                                          "LLMs currently do a better job than web search engines filtering spam and returning the most relevant documents."

                                          I've a nice bridge over the Shannon you might like to buy!

                                          futurebird@sauropods.winF This user is from outside of this forum
                                          futurebird@sauropods.winF This user is from outside of this forum
                                          futurebird@sauropods.win
                                          wrote sidst redigeret af
                                          #30

                                          @raymaccarthy @janbogar

                                          This is just true. It was not true a year or two ago but it is now.

                                          I found out because my husband like to use the LLM to search for articles. He's a PR guy and often needs to find out which paper contained a particular article.

                                          Google, yahoo, duckduck go are a maze of paywalls and fake sites. The LLM just gives the answer to the question with a link so you can verify it is correct.

                                          I am horrified.

                                          raymaccarthy@mastodon.ieR 1 Reply Last reply
                                          0
                                          Svar
                                          • Svar som emne
                                          Login for at svare
                                          • Ældste til nyeste
                                          • Nyeste til ældste
                                          • Most Votes


                                          • Log ind

                                          • Har du ikke en konto? Tilmeld

                                          • Login or register to search.
                                          Powered by NodeBB Contributors
                                          Graciously hosted by data.coop
                                          • First post
                                            Last post
                                          0
                                          • Hjem
                                          • Seneste
                                          • Etiketter
                                          • Populære
                                          • Verden
                                          • Bruger
                                          • Grupper