No, OpenAI's new magic models did not autonomously hack Huggingface.
-
-
No, OpenAI's new magic models did not autonomously hack Huggingface.
Per OpenAI's PR blog post (https://openai.com/index/hugging-face-model-evaluation-security-incident/

"This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity."
So they prompted their model running without guardrails to hack some shit. OpenAI told it to do that, the model didn't do shit autonomously (no LLM ever does anything autonomously, it's always prompted).
The whole story is PR. We know that from Anthopic's Mythos: "Look we have this super secret new model and it's so powerful. We are scared ourselves. And you will soon be able to rent it!"
"no LLM ever does anything autonomously, it's always prompted)."
Subgoals. You give it a complex task and it tries to solve it by breaking it down in subgoals... That it invents all itself...
Sure this COULD be a PR stunt. But then a lot of employees from 2 different companies would have to be in on the conspiracy... Not so likely
-
"no LLM ever does anything autonomously, it's always prompted)."
Subgoals. You give it a complex task and it tries to solve it by breaking it down in subgoals... That it invents all itself...
Sure this COULD be a PR stunt. But then a lot of employees from 2 different companies would have to be in on the conspiracy... Not so likely
-
No, OpenAI's new magic models did not autonomously hack Huggingface.
Per OpenAI's PR blog post (https://openai.com/index/hugging-face-model-evaluation-security-incident/

"This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity."
So they prompted their model running without guardrails to hack some shit. OpenAI told it to do that, the model didn't do shit autonomously (no LLM ever does anything autonomously, it's always prompted).
The whole story is PR. We know that from Anthopic's Mythos: "Look we have this super secret new model and it's so powerful. We are scared ourselves. And you will soon be able to rent it!"
@tante Anything to keep the billions flowing. What a disgrace.
-
No, OpenAI's new magic models did not autonomously hack Huggingface.
Per OpenAI's PR blog post (https://openai.com/index/hugging-face-model-evaluation-security-incident/

"This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity."
So they prompted their model running without guardrails to hack some shit. OpenAI told it to do that, the model didn't do shit autonomously (no LLM ever does anything autonomously, it's always prompted).
The whole story is PR. We know that from Anthopic's Mythos: "Look we have this super secret new model and it's so powerful. We are scared ourselves. And you will soon be able to rent it!"
@tante Indeed. As I put it yesterday, OpenAI’s agent “broke containment” the same way my cuckoo clock will break containment if I put it in front of a button.
-
To me it rather seems that the AI-denial is a cult..
Bergman:
https://youtu.be/KpTZbq-eV38?is=yIA9reuuiFsgwaRO
And with the "humiliations argument" we now have an explanation for it:
-
To me it rather seems that the AI-denial is a cult..
Bergman:
https://youtu.be/KpTZbq-eV38?is=yIA9reuuiFsgwaRO
And with the "humiliations argument" we now have an explanation for it:
-
To me it rather seems that the AI-denial is a cult..
Bergman:
https://youtu.be/KpTZbq-eV38?is=yIA9reuuiFsgwaRO
And with the "humiliations argument" we now have an explanation for it:
-
No, OpenAI's new magic models did not autonomously hack Huggingface.
Per OpenAI's PR blog post (https://openai.com/index/hugging-face-model-evaluation-security-incident/

"This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity."
So they prompted their model running without guardrails to hack some shit. OpenAI told it to do that, the model didn't do shit autonomously (no LLM ever does anything autonomously, it's always prompted).
The whole story is PR. We know that from Anthopic's Mythos: "Look we have this super secret new model and it's so powerful. We are scared ourselves. And you will soon be able to rent it!"
@tante
- i do not think it's unlikely that the prompt said "get a good score on this benchmark" and the model wildly misinterpreted the goal, as anyone who uses llms sees them do all the time
- even if the prompt was "hack huggingface", the fact that it did that without further prompting is noteworthy. -
To me it rather seems that the AI-denial is a cult..
Bergman:
https://youtu.be/KpTZbq-eV38?is=yIA9reuuiFsgwaRO
And with the "humiliations argument" we now have an explanation for it:
-
P pelle@veganism.social shared this topic
