upvote
> Today's frontier models can't even handle 100% of the SWE benchmarks that have been around for longer than they've been training the models. Companies that are benchmaxxing their models against the benchmarks haven't even been able to get them to 100%.

Is average developer 100% the benchmarks ?

reply