upvote
That’s a good visualization, although I am a bit mistrustful of Arena’s scores. It does get around the fact that models are getting trained for the benchmarks, but the methodology of letting random people compare outputs side-by-side is a very shallow judgement method in my opinion.

EDIT: Indeed looking at the overall rankings for text again, the list is rather strange, a lot more about writing style than intelligence.

reply