upvote
save qwen3.8 27B which is outclassing much larger models and is in spitting distance of the top 10 in https://artificialanalysis.ai/models#intelligence
reply
I wonder why they removed DeepSWE from their incorporates evaluations
reply
They didn't afaict https://artificialanalysis.ai/agents/coding-agents?coding-ag...

It seems it takes some time to run a new model on all the benchies, not sure they run all models on all of them either

reply