upvote
Same. With some 500 hours of usage in just my project at home, across both the $200 Claude and Codex subscriptions, I have not once encountered a situation where I would have attributed unsatisfactory results to a degradation in the model.

I've seen bugs in the harnesses, sure, but never anything in the actual model where I could have said with any certainty that it's not just regular variation or me having a bad day myself.

No idea where people get the confidence from to make such claims every other week.

reply
These analyses are much better than these Twitter charts.

I don't think anyone is reading the details for the Twitter post because it was not an actual benchmark. They did a post-hoc analysis of their logs from day to day.

Their random collection of prompts for each day is not a benchmark.

The site you linked is a much better example of a real benchmark being repeated over time.

reply