upvote
Load-bearing (whoops) quirks were noticeable months back, but haven't most flagship models become predictable and reliable?
reply
it does seem to be moving in that direction. There were really specific things (large, complex json outputs) that gemini-2.5 flash was basically the only model that seemed capable of reliably for a long period. gpt-5+ has covered the usecase for us now pretty well but still evals slightly below what 2.5 could do
reply
100% agreed in the same boat right now. Feeling really screwed over by Google rn
reply