upvote
> Terminal Bench: Flash 82.7 vs Terra 78.4

Terra 87.4

https://openai.com/index/gpt-5-6/

reply
https://www.tbench.ai/leaderboard/terminal-bench/2.1

> 78.4

The real score is always the official benchmark.

We need to see later if DS4 flash 0731 is going to maintain the score but we need to look at the official benchmarks.

Already seen a PR for DeepSWE to update the benchmark with 0731, so we can verify claimed vs official.

reply
The comment I replied to was comparing numbers released by each lab, and they made a typo with that specific benchmark.
reply
Open flash model is competing against OpenAI's 'Sonnet' model at the price of GPT 3, I am really excited about this release, hopefully it holds up in real work as well
reply
IIRC GPT 3 was priced at per 1k tokens, had to check, the biggest GPT 3 model from OpenAI was $0.06/1k, so $60 / 1M. gpt-3.5-turbo was the first model after ChatGPT and that was $2 / 1M. And no caching. So not really in the same ballpark
reply
Since they did this with their own harness I m not sure it's apples to apples comparison.
reply
> not sure it's apples to apples comparison.

They're literally comparing the previous version of the same model with the new one. It's based on the same architecture, same pre-trained model, just different post-training. It doesn't get more apples to apples than this.

reply
I think the commenter means the Flash vs Terra benchmarks.
reply
Ah, my bad. Yeah that makes sense. They do say "The official V4-Flash natively supports the Responses API format and is specifically adapted for Codex.", so at some point someone will make a "same harness" comparison.
reply
It will be fair if they release the harness though. I think now the future will be paired model-harness releases, not just weight dumps.

The performance changes are so big with the right harness that is makes sense to engineer the harness and fine-tune the model to one another from the start.

reply
I think it's better than GPT Luna.
reply