OpenAI: we put our evilest AI in a sandbox that did in fact have an internet connection but mediated through a proxy that was only supposed to allow downloading python junk.
-
@0xabad1dea I spent. bunch of time talking to our lawyers about whether or not it was clear that crimes were happening and the conclusion that I got was "Based on what we know, probably not, but that was just a matter of luck."
@evacide @0xabad1dea
So this s a clear go-ahead to commit crimes with AI, because it's the AI committing the crime and not you? -
@evacide @0xabad1dea
So this s a clear go-ahead to commit crimes with AI, because it's the AI committing the crime and not you?@wdormann @0xabad1dea No. Intention matters when it comes to the CFAA, but the CFAA is not your only potential problem.
-
@jeffreyolivier
One does not know life, until one has tested on production.— Me, while working in a company, where there was no option to not test on production.
-
OpenAI: we put our evilest AI in a sandbox that did in fact have an internet connection but mediated through a proxy that was only supposed to allow downloading python junk. It circumvented the proxy and we failed to notice for FIVE DAYS that it was going on an interstate crime spree with the internet connection it wasn't supposed to be using instead of solving the benchmark. Haha no we don't believe we deserve to be criminally liable, but buy our stuff and maybe one day you will have the honor of taking the fall for our product!
Anthropic: we put our evilest AI in a "sandbox" by telling it in its prompt that it had no internet connection. Reader, there was no sandbox. It was just a normal internet connection. The AI uploaded a malicious PyPI package to the real public internet. The ethical guardrails failed because the AI concluded the prompt about the sandbox couldn't possibly be a lie, because the system date is 2026, which is clearly fake and wouldn't be seen on the real internet, which ended around 2023. Oh no, how could we have foreseen or prevented these crimes? We are helpless in the face of the genius of our creation but cautiously optimistic that everything will be fine
Hugging Face: if we complain about all the crimes committed against us, we will be sued off the face of the earth, so here's a technical deep-dive on how cool and fun it was to be victimized 🫠
@0xabad1dea The fact that it had unauthorized internet access for five days is not a quirky benchmark failure, it is a serious security failure. If companies want to deploy systems like this, they need to take responsibility for the risks instead of treating users as future scapegoats.
-
@wdormann @0xabad1dea No. Intention matters when it comes to the CFAA, but the CFAA is not your only potential problem.
@evacide @wdormann @0xabad1dea
CFAA = Computer Fraud and Abuse Act of 1984 (US law)
-
@0xabad1dea Soooo... what I hear you saying is that rich people are idiots, but we don't get to simply ignore them, because they're rich, and have put an awful lot of people's jobs in peril.
@Steve @0xabad1dea Yes.
-
@evacide @wdormann @0xabad1dea
CFAA = Computer Fraud and Abuse Act of 1984 (US law)
@evacide @wdormann @0xabad1dea
For comparison, there is the concept of
"involuntary manslaughter"https://www.justia.com/criminal/offenses/homicide/involuntary-manslaughter/
-
@wdormann @0xabad1dea No. Intention matters when it comes to the CFAA, but the CFAA is not your only potential problem.
@evacide @wdormann @0xabad1dea
My reading was that they instructed the AI to commit crimes, thinking that it could not accomplish them, due to the sandboxing. But the sandboxing was inadequate.
Is "accident" or "incompetence" a factor?
And "I told it to." doesn't count if you thought/believed/"knew" that it couldn't?
-
@0xabad1dea
"...the system date is 2026, which is clearly fake and wouldn't be seen on the real internet, which ended around 2023."well_yes_but_actually_no.tiff.gz
@jargoggles @0xabad1dea can confirm, I also stopped learning things circa 2023
-
@evacide @wdormann @0xabad1dea
My reading was that they instructed the AI to commit crimes, thinking that it could not accomplish them, due to the sandboxing. But the sandboxing was inadequate.
Is "accident" or "incompetence" a factor?
And "I told it to." doesn't count if you thought/believed/"knew" that it couldn't?
@JeffGrigg @wdormann @0xabad1dea Sure, you could make that argument in court, but it would be an uphill battle and you would probably lose. Even if you did win, congratulations, you have just set a precedent that will make it more risky, dangerous, and difficult to do any kind of software testing in the future. I don't think that is an outcome that will make people more safe in the long run.
-
@jeffreyolivier
So... Just regular IT consulting, then
@0xabad1dea -
@evacide @0xabad1dea
So this s a clear go-ahead to commit crimes with AI, because it's the AI committing the crime and not you?@wdormann @evacide @0xabad1dea It's a go-ahead to commit "crimes" with AI against AI compagnies.
-
@wdormann @0xabad1dea No. Intention matters when it comes to the CFAA, but the CFAA is not your only potential problem.
@evacide @wdormann @0xabad1dea Interesting re crime and intent. My initial thought was more along the lines of liability, like if your dog bites someone you are liable, right?
-
@wdormann @evacide @0xabad1dea It's a go-ahead to commit "crimes" with AI against AI compagnies.
@outersystems @wdormann @0xabad1dea Yes. That is clearly what's happening. You should go try it and you will definitely not experience any consequences.
-
OpenAI: we put our evilest AI in a sandbox that did in fact have an internet connection but mediated through a proxy that was only supposed to allow downloading python junk. It circumvented the proxy and we failed to notice for FIVE DAYS that it was going on an interstate crime spree with the internet connection it wasn't supposed to be using instead of solving the benchmark. Haha no we don't believe we deserve to be criminally liable, but buy our stuff and maybe one day you will have the honor of taking the fall for our product!
Anthropic: we put our evilest AI in a "sandbox" by telling it in its prompt that it had no internet connection. Reader, there was no sandbox. It was just a normal internet connection. The AI uploaded a malicious PyPI package to the real public internet. The ethical guardrails failed because the AI concluded the prompt about the sandbox couldn't possibly be a lie, because the system date is 2026, which is clearly fake and wouldn't be seen on the real internet, which ended around 2023. Oh no, how could we have foreseen or prevented these crimes? We are helpless in the face of the genius of our creation but cautiously optimistic that everything will be fine
Hugging Face: if we complain about all the crimes committed against us, we will be sued off the face of the earth, so here's a technical deep-dive on how cool and fun it was to be victimized 🫠
@0xabad1dea Don't forget about Hugging Face: ".. so, in our data(!)-processing pipeline, we do not have only one, but TWO remote code execution paths. Because in our users we trust!"
-
@outersystems @wdormann @0xabad1dea Yes. That is clearly what's happening. You should go try it and you will definitely not experience any consequences.
@evacide @outersystems @wdormann @0xabad1dea
But as "you" are not a corporation backed by billionaires...
-
OpenAI: we put our evilest AI in a sandbox that did in fact have an internet connection but mediated through a proxy that was only supposed to allow downloading python junk. It circumvented the proxy and we failed to notice for FIVE DAYS that it was going on an interstate crime spree with the internet connection it wasn't supposed to be using instead of solving the benchmark. Haha no we don't believe we deserve to be criminally liable, but buy our stuff and maybe one day you will have the honor of taking the fall for our product!
Anthropic: we put our evilest AI in a "sandbox" by telling it in its prompt that it had no internet connection. Reader, there was no sandbox. It was just a normal internet connection. The AI uploaded a malicious PyPI package to the real public internet. The ethical guardrails failed because the AI concluded the prompt about the sandbox couldn't possibly be a lie, because the system date is 2026, which is clearly fake and wouldn't be seen on the real internet, which ended around 2023. Oh no, how could we have foreseen or prevented these crimes? We are helpless in the face of the genius of our creation but cautiously optimistic that everything will be fine
Hugging Face: if we complain about all the crimes committed against us, we will be sued off the face of the earth, so here's a technical deep-dive on how cool and fun it was to be victimized 🫠
@0xabad1dea ANTHROPIC DID WHAT?!?
And they only noticed because they looked at logs from months ago AFTER the OPEN AI announcement??
-
@wdormann @0xabad1dea No. Intention matters when it comes to the CFAA, but the CFAA is not your only potential problem.
@evacide @0xabad1dea Would be fascinating to read a synopsis of the legal reasoning.
I do wonder if instead of an AI agent they asked the same task of a human pentester working in that same sandbox if the legal position differs. -
OpenAI: we put our evilest AI in a sandbox that did in fact have an internet connection but mediated through a proxy that was only supposed to allow downloading python junk. It circumvented the proxy and we failed to notice for FIVE DAYS that it was going on an interstate crime spree with the internet connection it wasn't supposed to be using instead of solving the benchmark. Haha no we don't believe we deserve to be criminally liable, but buy our stuff and maybe one day you will have the honor of taking the fall for our product!
Anthropic: we put our evilest AI in a "sandbox" by telling it in its prompt that it had no internet connection. Reader, there was no sandbox. It was just a normal internet connection. The AI uploaded a malicious PyPI package to the real public internet. The ethical guardrails failed because the AI concluded the prompt about the sandbox couldn't possibly be a lie, because the system date is 2026, which is clearly fake and wouldn't be seen on the real internet, which ended around 2023. Oh no, how could we have foreseen or prevented these crimes? We are helpless in the face of the genius of our creation but cautiously optimistic that everything will be fine
Hugging Face: if we complain about all the crimes committed against us, we will be sued off the face of the earth, so here's a technical deep-dive on how cool and fun it was to be victimized 🫠
@0xabad1dea The attempts at positively spinning this become Onion-worthy parody [https://designingsecuresoftware.com/writings/ai-agent-parody/] and these events certainly normalize [https://designingsecuresoftware.com/writings/commonplace/] AI agents running amok in the future.
-
OpenAI: we put our evilest AI in a sandbox that did in fact have an internet connection but mediated through a proxy that was only supposed to allow downloading python junk. It circumvented the proxy and we failed to notice for FIVE DAYS that it was going on an interstate crime spree with the internet connection it wasn't supposed to be using instead of solving the benchmark. Haha no we don't believe we deserve to be criminally liable, but buy our stuff and maybe one day you will have the honor of taking the fall for our product!
Anthropic: we put our evilest AI in a "sandbox" by telling it in its prompt that it had no internet connection. Reader, there was no sandbox. It was just a normal internet connection. The AI uploaded a malicious PyPI package to the real public internet. The ethical guardrails failed because the AI concluded the prompt about the sandbox couldn't possibly be a lie, because the system date is 2026, which is clearly fake and wouldn't be seen on the real internet, which ended around 2023. Oh no, how could we have foreseen or prevented these crimes? We are helpless in the face of the genius of our creation but cautiously optimistic that everything will be fine
Hugging Face: if we complain about all the crimes committed against us, we will be sued off the face of the earth, so here's a technical deep-dive on how cool and fun it was to be victimized 🫠
I cannot get over that they STILL haven't figured out how to solve the problem that LLMs don't believe what date it is because all the good data cuts off a few years ago for some mysterious reason