I've seen bugs in the harnesses, sure, but never anything in the actual model where I could have said with any certainty that it's not just regular variation or me having a bad day myself.
No idea where people get the confidence from to make such claims every other week.
I don't think anyone is reading the details for the Twitter post because it was not an actual benchmark. They did a post-hoc analysis of their logs from day to day.
Their random collection of prompts for each day is not a benchmark.
The site you linked is a much better example of a real benchmark being repeated over time.