undefined

points

[-]

In mid-2024, Anthropic made the deliberate decision to stop chasing benchmarks and focus on practical value. There was a lot of skepticism at the time, but it's proven to be a prescient decision.

by girvo2 hours ago|

prev|

[-]

Benchmarks are basically straight up meaningless at this point in my experience. If they mattered and were the whole story, those Chinese open models would be stomping the competition right now. Instead they're merely decent when you use them in anger for real work.

I'll withhold judgement until I've tried to use it.

by avereveard56 minutes ago|

parent|

[-]

What's your opinion of glm5 if you had a chance to use it

by metadat1 hours ago|

prev|

[-]

Ranking Codex 5.2 ahead of plain 5.2 doesn't make sense. Codex is expressly designed for coding tasks. Not systems design, not problem analysis, and definitely not banking, but actually solving specific programming tasks (and it's very, very good at this). GPT 5.2 (non-codex) is better in every other way.

by nl1 hours ago|

parent|

[-]

Codex has been post-trained for coding, including agentic coding tasks.

It's certainly not impossible that the better long-horizon agentic performance in Codex overcomes any deficiencies in outright banking knowledge that Codex 5.2 has vs plain 5.2.

by 306bobby1 hours ago|

parent|

prev|

[-]

It could be problem specific. There are certain non program things that opus seems better than sonnet at as well

by 306bobby1 hours ago|

parent|

prev|

[-]

Swapped sonnet and opus on my last reply, oops

by blueaquilae2 hours ago|

prev|

[-]

Marketing team agree with benchmark score...

by HardCodedBias3 hours ago|

prev|

[-]

LOL come on man.

Let's give it a couple of days since no one believes anything from benchmarks, especially from the Gemini team (or Meta).

If we see on HN that people are willing switching their coding environment, we'll know "hot damn they cooked" otherwise this is another wiff by Google.

by drivebyhooting1 hours ago|

parent|

[-]

You can’t put Gemini and Meta in the same sentence. Llama 4 was DOA, and Meta has given up on frontier models. Internally they’re using Claude.

by not_ai34 minutes ago|

parent|

[-]

After spending all that money and firing a bunch of people? Is the new group doing anything at this point?