Most of the time I don't need what the bench tests and I'm not really giving them completely ambiguous tasks without any refinement.
I only find marginal differences between models at this point and it almost feels like personality quirks in each model than anything.
alias agy="agy --dangerously-skip-permissions"Anecdotally I'd rate Gemini behind Claude and OpenAI models at fiction and I can't find any benchmarks showing Gemini is the clear winner at this task.
I doubt they even intended it to be, but it seems like I kept going from resorting to 3.5-3.8 (over time) to realizing that Claude and GPT, while great at Python, will make rudimentary mistakes with R; even when they compose giant complicated R code.