Skip to content
  • Hjem
  • Seneste
  • Etiketter
  • Populære
  • Verden
  • Bruger
  • Grupper
Temaer
  • Light
  • Brite
  • Cerulean
  • Cosmo
  • Flatly
  • Journal
  • Litera
  • Lumen
  • Lux
  • Materia
  • Minty
  • Morph
  • Pulse
  • Sandstone
  • Simplex
  • Sketchy
  • Spacelab
  • United
  • Yeti
  • Zephyr
  • Dark
  • Cyborg
  • Darkly
  • Quartz
  • Slate
  • Solar
  • Superhero
  • Vapor

  • Default (No Skin)
  • No Skin
Kollaps
FARVEL BIG TECH
  1. Forside
  2. Ikke-kategoriseret
  3. OpenAI: we put our evilest AI in a sandbox that did in fact have an internet connection but mediated through a proxy that was only supposed to allow downloading python junk.

OpenAI: we put our evilest AI in a sandbox that did in fact have an internet connection but mediated through a proxy that was only supposed to allow downloading python junk.

Planlagt Fastgjort Låst Flyttet Ikke-kategoriseret
77 Indlæg 53 Posters 0 Visninger
  • Ældste til nyeste
  • Nyeste til ældste
  • Most Votes
Svar
  • Svar som emne
Login for at svare
Denne tråd er blevet slettet. Kun brugere med emne behandlings privilegier kan se den.
  • 0xabad1dea@infosec.exchange0 0xabad1dea@infosec.exchange

    OpenAI: we put our evilest AI in a sandbox that did in fact have an internet connection but mediated through a proxy that was only supposed to allow downloading python junk. It circumvented the proxy and we failed to notice for FIVE DAYS that it was going on an interstate crime spree with the internet connection it wasn't supposed to be using instead of solving the benchmark. Haha no we don't believe we deserve to be criminally liable, but buy our stuff and maybe one day you will have the honor of taking the fall for our product!

    Anthropic: we put our evilest AI in a "sandbox" by telling it in its prompt that it had no internet connection. Reader, there was no sandbox. It was just a normal internet connection. The AI uploaded a malicious PyPI package to the real public internet. The ethical guardrails failed because the AI concluded the prompt about the sandbox couldn't possibly be a lie, because the system date is 2026, which is clearly fake and wouldn't be seen on the real internet, which ended around 2023. Oh no, how could we have foreseen or prevented these crimes? We are helpless in the face of the genius of our creation but cautiously optimistic that everything will be fine 🙂

    Hugging Face: if we complain about all the crimes committed against us, we will be sued off the face of the earth, so here's a technical deep-dive on how cool and fun it was to be victimized 🫠

    budududuroiu@hachyderm.ioB This user is from outside of this forum
    budududuroiu@hachyderm.ioB This user is from outside of this forum
    budududuroiu@hachyderm.io
    wrote sidst redigeret af
    #60

    @0xabad1dea the goal post shifted from "stochastic parrot" to "enclosure not good enough". Truly a wild victory for AI safetists today, as not even luddite Mastodon can contest that badly aligned AI will breach your shit

    cynaq@beige.partyC 1 Reply Last reply
    0
    • dzwiedziu@mastodon.socialD dzwiedziu@mastodon.social

      @jeffreyolivier
      One does not know life, until one has tested on production.

      — Me, while working in a company, where there was no option to not test on production.

      @0xabad1dea
      @alice

      jayalane@mastodon.onlineJ This user is from outside of this forum
      jayalane@mastodon.onlineJ This user is from outside of this forum
      jayalane@mastodon.online
      wrote sidst redigeret af
      #61

      @dzwiedziu @jeffreyolivier @0xabad1dea @alice even when I have tested well pre production, I don't relax till some time after the prod deploy is finished. How long depends on the code and past behaviors. Never less than fifteen minutes, seems to take network, memory and disk cache and what not some time to settle in.

      1 Reply Last reply
      0
      • kevingranade@mastodon.gamedev.placeK kevingranade@mastodon.gamedev.place

        @0xabad1dea weird it's like LLMs don't actually handle facts or knowledge but just contextualless strings of characters.

        To be clear the sarcasm is only aimed at LLM boosters not you.

        mathew@universeodon.comM This user is from outside of this forum
        mathew@universeodon.comM This user is from outside of this forum
        mathew@universeodon.com
        wrote sidst redigeret af
        #62

        @kevingranade @0xabad1dea Yeah, it continues to amaze me that people will put instructions in prompts like "Don't commit crimes", "You don't have an Internet connection", "Check all facts are correct", like the LLM even understands those instructions.

        naught101@mastodon.socialN 1 Reply Last reply
        0
        • steve@social.coopS steve@social.coop

          @0xabad1dea Soooo... what I hear you saying is that rich people are idiots, but we don't get to simply ignore them, because they're rich, and have put an awful lot of people's jobs in peril.

          naught101@mastodon.socialN This user is from outside of this forum
          naught101@mastodon.socialN This user is from outside of this forum
          naught101@mastodon.social
          wrote sidst redigeret af
          #63

          @Steve @0xabad1dea that was always true, well before the tech sector existed

          1 Reply Last reply
          0
          • mathew@universeodon.comM mathew@universeodon.com

            @kevingranade @0xabad1dea Yeah, it continues to amaze me that people will put instructions in prompts like "Don't commit crimes", "You don't have an Internet connection", "Check all facts are correct", like the LLM even understands those instructions.

            naught101@mastodon.socialN This user is from outside of this forum
            naught101@mastodon.socialN This user is from outside of this forum
            naught101@mastodon.social
            wrote sidst redigeret af
            #64

            @mathew @kevingranade @0xabad1dea that's literally how the open AI tutorial on how to write guardrails is written

            https://developers.openai.com/cookbook/examples/how_to_use_guardrails

            kevingranade@mastodon.gamedev.placeK 1 Reply Last reply
            0
            • 0xabad1dea@infosec.exchange0 0xabad1dea@infosec.exchange

              OpenAI: we put our evilest AI in a sandbox that did in fact have an internet connection but mediated through a proxy that was only supposed to allow downloading python junk. It circumvented the proxy and we failed to notice for FIVE DAYS that it was going on an interstate crime spree with the internet connection it wasn't supposed to be using instead of solving the benchmark. Haha no we don't believe we deserve to be criminally liable, but buy our stuff and maybe one day you will have the honor of taking the fall for our product!

              Anthropic: we put our evilest AI in a "sandbox" by telling it in its prompt that it had no internet connection. Reader, there was no sandbox. It was just a normal internet connection. The AI uploaded a malicious PyPI package to the real public internet. The ethical guardrails failed because the AI concluded the prompt about the sandbox couldn't possibly be a lie, because the system date is 2026, which is clearly fake and wouldn't be seen on the real internet, which ended around 2023. Oh no, how could we have foreseen or prevented these crimes? We are helpless in the face of the genius of our creation but cautiously optimistic that everything will be fine 🙂

              Hugging Face: if we complain about all the crimes committed against us, we will be sued off the face of the earth, so here's a technical deep-dive on how cool and fun it was to be victimized 🫠

              icedquinn@blob.catI This user is from outside of this forum
              icedquinn@blob.catI This user is from outside of this forum
              icedquinn@blob.cat
              wrote sidst redigeret af
              #65
              @0xabad1dea they are fundamentally the same company that had an internal management disagreement.
              1 Reply Last reply
              0
              • budududuroiu@hachyderm.ioB budududuroiu@hachyderm.io

                @0xabad1dea the goal post shifted from "stochastic parrot" to "enclosure not good enough". Truly a wild victory for AI safetists today, as not even luddite Mastodon can contest that badly aligned AI will breach your shit

                cynaq@beige.partyC This user is from outside of this forum
                cynaq@beige.partyC This user is from outside of this forum
                cynaq@beige.party
                wrote sidst redigeret af
                #66

                @budududuroiu @0xabad1dea

                "badly aligned" in this case means specifically trained on "how to hack shit" manuals, proverbially speaking. It's still stochastic parrot, I'm afraid.

                budududuroiu@hachyderm.ioB 1 Reply Last reply
                0
                • dzwiedziu@mastodon.socialD dzwiedziu@mastodon.social

                  @jeffreyolivier
                  One does not know life, until one has tested on production.

                  — Me, while working in a company, where there was no option to not test on production.

                  @0xabad1dea
                  @alice

                  jackeric@beige.partyJ This user is from outside of this forum
                  jackeric@beige.partyJ This user is from outside of this forum
                  jackeric@beige.party
                  wrote sidst redigeret af
                  #67

                  @dzwiedziu @jeffreyolivier @0xabad1dea @alice in a previous job, we introduced a sandbox feature to allow our largest customers to experiment safely with sweeping changes

                  reader, there was no sandbox... the sandbox instance ran in the same database as their production system, with all the production data duplicated but with a new TenantId. all well and good except the app has a feature to execute customer-owned stored procedures, but doesn't pass the TenantId when it does so, so the stored procs by convention are hard-coded with the prod TenantId... and when you create a sandbox, it copies these stored procedure triggers that... act on your production data 🙃

                  raised the issue but no-one was able or willing to categorise this as urgent for whatever reason

                  1 Reply Last reply
                  0
                  • jumpmed@mastodon.socialJ jumpmed@mastodon.social

                    @grwster @evacide @wdormann @0xabad1dea The difference is that we really don't have any laws on criminal liability for what software does. Up until the llm era, software was fairly predictable. An outside observer could tell if a package was designed to do something malicious. Now we need laws that essentially establish a "you should have known better" criminal liability for software.

                    supermoosie@mastodon.auS This user is from outside of this forum
                    supermoosie@mastodon.auS This user is from outside of this forum
                    supermoosie@mastodon.au
                    wrote sidst redigeret af
                    #68

                    @Jumpmed @grwster @evacide @wdormann @0xabad1dea

                    Laws are on thing.

                    Getting courts and their officials to understand without getting baffled is another seperate challenge.

                    1 Reply Last reply
                    0
                    • suetanvil@freeradical.zoneS suetanvil@freeradical.zone

                      @ggreer @0xabad1dea

                      Huh. I wonder if the slop contents on the modern Internet is making it impossible to add training data.

                      natanox@chaos.socialN This user is from outside of this forum
                      natanox@chaos.socialN This user is from outside of this forum
                      natanox@chaos.social
                      wrote sidst redigeret af
                      #69

                      @suetanvil @ggreer @0xabad1dea Why do you think Anthropic started buying, scanning and then burning exceedingly rare and expensive books? The content farms in poor countries can't possibly create enough and the internet is a swamp of slop by now.

                      suetanvil@freeradical.zoneS 1 Reply Last reply
                      0
                      • naught101@mastodon.socialN naught101@mastodon.social

                        @mathew @kevingranade @0xabad1dea that's literally how the open AI tutorial on how to write guardrails is written

                        https://developers.openai.com/cookbook/examples/how_to_use_guardrails

                        kevingranade@mastodon.gamedev.placeK This user is from outside of this forum
                        kevingranade@mastodon.gamedev.placeK This user is from outside of this forum
                        kevingranade@mastodon.gamedev.place
                        wrote sidst redigeret af
                        #70

                        @naught101 @mathew @0xabad1dea that just means the people that wrote that doc are also clueless yes.

                        naught101@mastodon.socialN 1 Reply Last reply
                        0
                        • qroole@mastodon.socialQ qroole@mastodon.social

                          @0xabad1dea The fact that it had unauthorized internet access for five days is not a quirky benchmark failure, it is a serious security failure. If companies want to deploy systems like this, they need to take responsibility for the risks instead of treating users as future scapegoats.

                          nep@mstdn.caN This user is from outside of this forum
                          nep@mstdn.caN This user is from outside of this forum
                          nep@mstdn.ca
                          wrote sidst redigeret af
                          #71

                          @qroole @0xabad1dea A single real teenager could get into a huge amount of trouble in five days online… this is a disaster.

                          1 Reply Last reply
                          0
                          • kevingranade@mastodon.gamedev.placeK kevingranade@mastodon.gamedev.place

                            @naught101 @mathew @0xabad1dea that just means the people that wrote that doc are also clueless yes.

                            naught101@mastodon.socialN This user is from outside of this forum
                            naught101@mastodon.socialN This user is from outside of this forum
                            naught101@mastodon.social
                            wrote sidst redigeret af
                            #72

                            @kevingranade @mathew @0xabad1dea maybe. I suspect it means that the idea of guardrails implemented in this way is stupid and dangerous, but also very common

                            1 Reply Last reply
                            0
                            • natanox@chaos.socialN natanox@chaos.social

                              @suetanvil @ggreer @0xabad1dea Why do you think Anthropic started buying, scanning and then burning exceedingly rare and expensive books? The content farms in poor countries can't possibly create enough and the internet is a swamp of slop by now.

                              suetanvil@freeradical.zoneS This user is from outside of this forum
                              suetanvil@freeradical.zoneS This user is from outside of this forum
                              suetanvil@freeradical.zone
                              wrote sidst redigeret af
                              #73

                              @Natanox @ggreer @0xabad1dea

                              I mean, it's primarily because they're dicks, but yeah.

                              1 Reply Last reply
                              0
                              • cynaq@beige.partyC cynaq@beige.party

                                @budududuroiu @0xabad1dea

                                "badly aligned" in this case means specifically trained on "how to hack shit" manuals, proverbially speaking. It's still stochastic parrot, I'm afraid.

                                budududuroiu@hachyderm.ioB This user is from outside of this forum
                                budududuroiu@hachyderm.ioB This user is from outside of this forum
                                budududuroiu@hachyderm.io
                                wrote sidst redigeret af
                                #74

                                @CynAq @0xabad1dea ok, why has no one else breached HF to steal eval data then? There's millions of dollars on the line for new labs to show themselves as "challengers", why haven't they done that?

                                cynaq@beige.partyC 1 Reply Last reply
                                0
                                • budududuroiu@hachyderm.ioB budududuroiu@hachyderm.io

                                  @CynAq @0xabad1dea ok, why has no one else breached HF to steal eval data then? There's millions of dollars on the line for new labs to show themselves as "challengers", why haven't they done that?

                                  cynaq@beige.partyC This user is from outside of this forum
                                  cynaq@beige.partyC This user is from outside of this forum
                                  cynaq@beige.party
                                  wrote sidst redigeret af
                                  #75

                                  @budududuroiu @0xabad1dea that’s against the rules of the challenge. Hacking hugging face is only viable if you can do it AND claim it was accidental for marketing purposes. Anyone else but OpenAI and Anthropic would be pounced on as cheaters if not outright cyber criminals.

                                  budududuroiu@hachyderm.ioB 1 Reply Last reply
                                  0
                                  • cynaq@beige.partyC cynaq@beige.party

                                    @budududuroiu @0xabad1dea that’s against the rules of the challenge. Hacking hugging face is only viable if you can do it AND claim it was accidental for marketing purposes. Anyone else but OpenAI and Anthropic would be pounced on as cheaters if not outright cyber criminals.

                                    budududuroiu@hachyderm.ioB This user is from outside of this forum
                                    budududuroiu@hachyderm.ioB This user is from outside of this forum
                                    budududuroiu@hachyderm.io
                                    wrote sidst redigeret af
                                    #76

                                    @CynAq @0xabad1dea nothing I'm gonna say will ever convince you, so I'm not gonna bother

                                    cynaq@beige.partyC 1 Reply Last reply
                                    0
                                    • budududuroiu@hachyderm.ioB budududuroiu@hachyderm.io

                                      @CynAq @0xabad1dea nothing I'm gonna say will ever convince you, so I'm not gonna bother

                                      cynaq@beige.partyC This user is from outside of this forum
                                      cynaq@beige.partyC This user is from outside of this forum
                                      cynaq@beige.party
                                      wrote sidst redigeret af
                                      #77

                                      @budududuroiu @0xabad1dea might be a wise choice as I’m not even sure exactly what you’re trying to convince me of.

                                      1 Reply Last reply
                                      0
                                      • jeppe@uddannelse.socialJ jeppe@uddannelse.social shared this topic
                                      Svar
                                      • Svar som emne
                                      Login for at svare
                                      • Ældste til nyeste
                                      • Nyeste til ældste
                                      • Most Votes


                                      • Log ind

                                      • Har du ikke en konto? Tilmeld

                                      • Login or register to search.
                                      Powered by NodeBB Contributors
                                      Graciously hosted by data.coop
                                      • First post
                                        Last post
                                      0
                                      • Hjem
                                      • Seneste
                                      • Etiketter
                                      • Populære
                                      • Verden
                                      • Bruger
                                      • Grupper