upvote
Because AA Coding "Index" consists only of two benchmarks (Terminal-Bench v2.1, SciCode) and generally fails to be meaningfully representative of agentic coding capabilities.
reply
AA coding index has been updated to use DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA.
reply
When? It literally says on the page for Gemini 3.6 Flash "Artificial Analysis Coding Index represents the weighted average of coding benchmarks in the Artificial Analysis Intelligence Index (Terminal-Bench v2.1, SciCode)"
reply
Whats a better option for AA Coding Index?
reply
DeepSWE and FrontierCode are more realistic if you read up on what they actually measure. But the most realistic is to try it yourself. Benchmarks can only vaguely represent typical usage, and how you judge the result. Giving the same real task you have to a few models will make you understand them better than chasing benchmarks.
reply
Could be a good tradeoff for the flash model though. 3.5 -> 3.6 is a tiny bit cheaper and maybe faster?

artificialanalysis.ai has it going from 165 tps -> 304 tps. openrouter.ai needs more data but it has it going from ~100 tps -> ~150 tps, though at peak 3.5 has reached 156tps.

reply