@thawn yes I understand the points and thank you for writing it out more clearly. Link to the OpenAI article: https://openai.com/index/hugging-face-model-evaluation-security-incident/
I agree that they sound kind of "proud" of their model being the most capable.
However, the main messages of the article are:
- we want "to help defenders understand what happened"
- we want "to help calibrate on what models are now capable of"
I think those are important and valid points. Soon, other AI labs will have models with similar capabilities, some maybe less restricted. I dont see a reason to believe this case is fabricated. From my experience with using current AI models as a data scientist I believe that using a swarm of next-gen models has indeed the level of capability that is described here. So if we take it at face value that the capabilities are real, the statements that "defenders need to prepare and calibrate" are reasonable. If the models can break into hugingface they can probably break into other even more critical infrastructure too. And in the huggingface case, the goal of the model was just to get results, not to do real damage like to delete databases or to steal user data (just did it as a sidequest in the huggingface case).
Regarding marketing I am also questioning: marketing for what? The final model will have guardrails against this. Such story could also lead to blocking of the model. And they kind of say that they are too incopetent to handle such model capability (escaped their sandbox, etc..)