OpenAI: we put our evilest AI in a sandbox that did in fact have an internet connection but mediated through a proxy that was only supposed to allow downloading python junk.
-
@grwster @evacide @wdormann @0xabad1dea The difference is that we really don't have any laws on criminal liability for what software does. Up until the llm era, software was fairly predictable. An outside observer could tell if a package was designed to do something malicious. Now we need laws that essentially establish a "you should have known better" criminal liability for software.
@Jumpmed @grwster @evacide @wdormann @0xabad1dea
Laws are on thing.
Getting courts and their officials to understand without getting baffled is another seperate challenge.
-
Huh. I wonder if the slop contents on the modern Internet is making it impossible to add training data.
@suetanvil @ggreer @0xabad1dea Why do you think Anthropic started buying, scanning and then burning exceedingly rare and expensive books? The content farms in poor countries can't possibly create enough and the internet is a swamp of slop by now.
-
@mathew @kevingranade @0xabad1dea that's literally how the open AI tutorial on how to write guardrails is written
https://developers.openai.com/cookbook/examples/how_to_use_guardrails
@naught101 @mathew @0xabad1dea that just means the people that wrote that doc are also clueless yes.
-
@0xabad1dea The fact that it had unauthorized internet access for five days is not a quirky benchmark failure, it is a serious security failure. If companies want to deploy systems like this, they need to take responsibility for the risks instead of treating users as future scapegoats.
@qroole @0xabad1dea A single real teenager could get into a huge amount of trouble in five days online… this is a disaster.
-
@naught101 @mathew @0xabad1dea that just means the people that wrote that doc are also clueless yes.
@kevingranade @mathew @0xabad1dea maybe. I suspect it means that the idea of guardrails implemented in this way is stupid and dangerous, but also very common
-
@suetanvil @ggreer @0xabad1dea Why do you think Anthropic started buying, scanning and then burning exceedingly rare and expensive books? The content farms in poor countries can't possibly create enough and the internet is a swamp of slop by now.
I mean, it's primarily because they're dicks, but yeah.
-
"badly aligned" in this case means specifically trained on "how to hack shit" manuals, proverbially speaking. It's still stochastic parrot, I'm afraid.
@CynAq @0xabad1dea ok, why has no one else breached HF to steal eval data then? There's millions of dollars on the line for new labs to show themselves as "challengers", why haven't they done that?
-
@CynAq @0xabad1dea ok, why has no one else breached HF to steal eval data then? There's millions of dollars on the line for new labs to show themselves as "challengers", why haven't they done that?
@budududuroiu @0xabad1dea that’s against the rules of the challenge. Hacking hugging face is only viable if you can do it AND claim it was accidental for marketing purposes. Anyone else but OpenAI and Anthropic would be pounced on as cheaters if not outright cyber criminals.
-
@budududuroiu @0xabad1dea that’s against the rules of the challenge. Hacking hugging face is only viable if you can do it AND claim it was accidental for marketing purposes. Anyone else but OpenAI and Anthropic would be pounced on as cheaters if not outright cyber criminals.
@CynAq @0xabad1dea nothing I'm gonna say will ever convince you, so I'm not gonna bother
-
@CynAq @0xabad1dea nothing I'm gonna say will ever convince you, so I'm not gonna bother
@budududuroiu @0xabad1dea might be a wise choice as I’m not even sure exactly what you’re trying to convince me of.
-
J jeppe@uddannelse.social shared this topic