upvote
It was certainly almost RL-fried to overfit the benchmarks, at the expense of actual usability. See Opus 5.
reply
is it a benchmarkmaxxing model?!
reply