The suite is across many categories, not only coding, and most of the tasks are low-horizon (or what the opposite of long-horizon is), where the max thinking time is around 10 minutes.
Gemini models are really smart, unfortunately they don't play well with any harness, so hard to use in practice.
But try them out for one-shot tasks, they are really good. Don't use them for coding in a harness, but you can ask them to generate code/planning (still, for coding only other models are indeed recommended).
Happy to hear what would make the website more useful.
Would I want it to grow and make money at some point? Sure, why not, then I can test even more models at higher reasoning levels. Meanwhile it's just me testing models when they come out and publishing the results for anyone who finds them useful.
I don't see why posting some info and a link with my own findings, in a relevant discussion is considered spam. Should it be?