upvote
It's still a bit weird though. For obvious reasons, The MO for benchmarks like this has been that the models remain pretty low until suddenly it's done. So even for a 'should i be worried yet' reality check, it's pretty terrible. You can't really keep track of what models are actually able to do. By the time you can replace years of research by specialists with a few api calls then...
reply
> incompetent people in decision-making positions will think researchers can be replaced by AI.

This is a certainty, not a fear =[.

reply