No, OpenAI's new magic models did not autonomously hack Huggingface.
-
@tante
Hadn't heard of "Hugging Face" before. Running out of evil things in Tolkien to name companies & systems for, I guess so here's the face hugger from Alien to be an AI overlord. Lovely.@Gurre @tante No to defend anyone but more to take that concern from you (I would also hate if they start using Alien universe names). The company is named after the hugging face emoji https://en.wikipedia.org/wiki/Hugging_Face
-
@davidgerard @tante The whole story was implausible from the second sentence: Where Huggingface claims that their “AI” detected the attack.
@erlenmayr@chaos.social @davidgerard@circumstances.run @tante@tldr.nettime.org The wild thing about these PR/marketing stunts is that if the "AI" really could do what they're pretending it can do, it'd be doing it all day every day. We'd drown in the announcements. At some point they'd become boring.
This one falls flat on its huggingface right out of the gate.
-
@hagen @tante I don't know, I'd want details on "To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy." and "the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path" before dismissing this...
'How did it get the "stolen credentials"?'
-
@tante How Tagesschau is reporting about it can only be described as refusal to practice journalism. 7 out of 9 paragraphs say "Open AI says" and the other two quote nameless "Experts" and Anthropic's product. And tomorrow I have to discuss with someone again about the possibility of consciousness in token generator software https://www.tagesschau.de/wirtschaft/unternehmen/openai-ki-hackerangriff-100.html
-
@michelin well they worked together with Huggingface (on PR and telling them about a bug they may or may not have found). So unless Huggingface sues them who would attack OpenAI?
-
@Gurre @tante No to defend anyone but more to take that concern from you (I would also hate if they start using Alien universe names). The company is named after the hugging face emoji https://en.wikipedia.org/wiki/Hugging_Face
Their Public Relations department can say that all they want.
But we know what it really stands for.
-
Their Public Relations department can say that all they want.
But we know what it really stands for.
Open AI's statement,
"These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities."
makes me think of all the actions of all the "corporate types" in the Alien franchise too.
And that's *never* a good thing.

-
@tante How Tagesschau is reporting about it can only be described as refusal to practice journalism. 7 out of 9 paragraphs say "Open AI says" and the other two quote nameless "Experts" and Anthropic's product. And tomorrow I have to discuss with someone again about the possibility of consciousness in token generator software https://www.tagesschau.de/wirtschaft/unternehmen/openai-ki-hackerangriff-100.html
@miles_leif @tante this is how 99% of the press is dealing with AI since the beginning of the hype, a hype they contribute to feed.
-
No, OpenAI's new magic models did not autonomously hack Huggingface.
Per OpenAI's PR blog post (https://openai.com/index/hugging-face-model-evaluation-security-incident/

"This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity."
So they prompted their model running without guardrails to hack some shit. OpenAI told it to do that, the model didn't do shit autonomously (no LLM ever does anything autonomously, it's always prompted).
The whole story is PR. We know that from Anthopic's Mythos: "Look we have this super secret new model and it's so powerful. We are scared ourselves. And you will soon be able to rent it!"
-
No, OpenAI's new magic models did not autonomously hack Huggingface.
Per OpenAI's PR blog post (https://openai.com/index/hugging-face-model-evaluation-security-incident/

"This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity."
So they prompted their model running without guardrails to hack some shit. OpenAI told it to do that, the model didn't do shit autonomously (no LLM ever does anything autonomously, it's always prompted).
The whole story is PR. We know that from Anthopic's Mythos: "Look we have this super secret new model and it's so powerful. We are scared ourselves. And you will soon be able to rent it!"
yesterday NPR had a long story on the new hot battery powered pickup truck from the startup
https://www.slate.auto/enand the NPR piece was full of the most hilarious nonsense, all provided to NPR by Slate's excellent PR team
shrug
-
No, OpenAI's new magic models did not autonomously hack Huggingface.
Per OpenAI's PR blog post (https://openai.com/index/hugging-face-model-evaluation-security-incident/

"This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity."
So they prompted their model running without guardrails to hack some shit. OpenAI told it to do that, the model didn't do shit autonomously (no LLM ever does anything autonomously, it's always prompted).
The whole story is PR. We know that from Anthopic's Mythos: "Look we have this super secret new model and it's so powerful. We are scared ourselves. And you will soon be able to rent it!"
@tante "OpenAI told it to do that, the model didn't do shit autonomously [...]"
Exactly, what one should have expected (and how it turned out before every time). -
@tante How Tagesschau is reporting about it can only be described as refusal to practice journalism. 7 out of 9 paragraphs say "Open AI says" and the other two quote nameless "Experts" and Anthropic's product. And tomorrow I have to discuss with someone again about the possibility of consciousness in token generator software https://www.tagesschau.de/wirtschaft/unternehmen/openai-ki-hackerangriff-100.html
@miles_leif @tante Not the only one unfortunately, sigh
-
@tante a friend suggests it's co-marketing with Huggingface, if you look at the timeline of announcements
Yeh, they're both chatbot companies.
-
'How did it get the "stolen credentials"?'
@JeffGrigg @SonOfSunTzu @hagen @tante
password.txt -
@SonOfSunTzu @tante that’s what mythos did as well though? you point at something and say go and if you’re willing to pay for compute it goes. it literally did what it was meant to do – and was explicitly told to do. what’s the news here?
@hagen @tante assuming a genuine question ... AIUI Mythos carried out tests in a controlled and authorised way, directly answering the tests it was set.
OpenAI's "collection of models" hacked out of a sandbox, and from there hacked into a different organisation, when given a set of tests it was meant to solve inside the sandbox.
Kind of as if a student sitting an exam broke into the exam board's office to steal the answers - so it's different in terms of alignment, and legal responsibility...
-
-
No, OpenAI's new magic models did not autonomously hack Huggingface.
Per OpenAI's PR blog post (https://openai.com/index/hugging-face-model-evaluation-security-incident/

"This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity."
So they prompted their model running without guardrails to hack some shit. OpenAI told it to do that, the model didn't do shit autonomously (no LLM ever does anything autonomously, it's always prompted).
The whole story is PR. We know that from Anthopic's Mythos: "Look we have this super secret new model and it's so powerful. We are scared ourselves. And you will soon be able to rent it!"
"no LLM ever does anything autonomously, it's always prompted)."
Subgoals. You give it a complex task and it tries to solve it by breaking it down in subgoals... That it invents all itself...
Sure this COULD be a PR stunt. But then a lot of employees from 2 different companies would have to be in on the conspiracy... Not so likely
-
"no LLM ever does anything autonomously, it's always prompted)."
Subgoals. You give it a complex task and it tries to solve it by breaking it down in subgoals... That it invents all itself...
Sure this COULD be a PR stunt. But then a lot of employees from 2 different companies would have to be in on the conspiracy... Not so likely
-
No, OpenAI's new magic models did not autonomously hack Huggingface.
Per OpenAI's PR blog post (https://openai.com/index/hugging-face-model-evaluation-security-incident/

"This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity."
So they prompted their model running without guardrails to hack some shit. OpenAI told it to do that, the model didn't do shit autonomously (no LLM ever does anything autonomously, it's always prompted).
The whole story is PR. We know that from Anthopic's Mythos: "Look we have this super secret new model and it's so powerful. We are scared ourselves. And you will soon be able to rent it!"
@tante Anything to keep the billions flowing. What a disgrace.
-
No, OpenAI's new magic models did not autonomously hack Huggingface.
Per OpenAI's PR blog post (https://openai.com/index/hugging-face-model-evaluation-security-incident/

"This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity."
So they prompted their model running without guardrails to hack some shit. OpenAI told it to do that, the model didn't do shit autonomously (no LLM ever does anything autonomously, it's always prompted).
The whole story is PR. We know that from Anthopic's Mythos: "Look we have this super secret new model and it's so powerful. We are scared ourselves. And you will soon be able to rent it!"
@tante Indeed. As I put it yesterday, OpenAI’s agent “broke containment” the same way my cuckoo clock will break containment if I put it in front of a button.

