upvote
Which benchmarks? Only the ones OpenAI cherry-picked.

It debuted as ~same score as Sol on Artificial Analysis. People couldn't accept it so they had to change the formula.

The model is a big step forward only in desktop use and 3D. That's impressive, but for software engineering, Fable is still in a league of its own.

reply
That's exactly the kind of behaviour I've seen, unbelievable amount of unit tests, and revisiting and revising the same code over and over again. If I was cynical, I'd say almost like it was deliberately trying to burn quota, even after I told it quota was getting low and to move onto actual physical testing.
reply
was it maybe over-quantised to reduce costs?
reply