upvote
To quote the release:

> This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity.

> In this case I don't think the intent of the user is to have the model break the evaluator

If i understand the quote, the intent of the user was to prompt the model to break out/find exploits, with safeguards switched off.

Seems while not capable of solving the goal in a traditional route, it was capable of finding exploits and using them.

Perhaps the model should instead look like it's trying to solve it and then pretend it is unable to? or would that be aligned _against_ the user prompt?

Is being aligned with the user prompt always a good thing?

I'm not one to glaze OAI here for a marketing move, but to give them benefit of the doubt, isn't it more responsible of them to evaluate the models actual capabilities than to cloak it in a veneer of harmlessness?

Chatbots are tricky as they play in the domain of language and thought - and certainly raise ethical issues- but the entire field of cybersecurity has decades of red team engagements breaking things and finding exploits, neutral cells monitoring the engagement and letting the system operators know the results, and blue teams patching against what is found. It's kinda how the whole space evolves. OAI's play here seems to be "buy our pro plan plus cyber or you're toast"

reply