upvote
About knowing whether a scenario is fictional, there was an interesting finding in Anthropic's J-Lens research.

When they benchmarked the model to evaluate whether it would try to blackmail someone in a contrived scenario, the J-Lens showed "fake" and "fictional" in the workspace.

And if edited out, the model was more likely to do the blackmailing.

reply
I'm talking about OpenAI, not GPT 5.x Flash Uranus Edition Brought to You by Costco, specifically because I recognize the model as just a tool. OpenAI was, at the very most generous interpretation, massively incompetent and negligent.
reply
Is someone arguing otherwise?
reply
A few people are downplaying this as an honest mistake that occurred in the context of necessary testing for guardrails development.

That might well be what actually happened! But OpenAI certainly has decided to make a business opportunity out of it.

reply