I recently tried doing a fairly normal task for this codebase with codex, as I have seen a lot of people talking it up on here. A single task running for ~1-2 hours burned through over half of my usage for the week on the $125/month plan, not on a top model (I don't remember which one specifically I used). It struggled to get the basics done, then got absolutely stuck on a follow up. Handed it over to Claude and it 1-shot it.
For most software eng and design work opus 4.6-4.8 just works fine. For everyday joe asking ai to plan a trip or home diy work even sonnet works fine.
Any cybersecurity or other areas are niches that cannot support trillion $ valuations. What am I missing? Genuinely curious
I just did a direct comparison, big change in a quite complex codebase. Same prompt for Opus, same for Fable. Fable clearly won and delivered very good results, while Opus delivered mediocre, so I did not let it finish. I expected both to fail and was prepared to do lots of manual steering, but not necessary with Fable one shotting it, and all this with 35$ of credits for fable. I am still impressed. If I would have had to hire a human, it would have cost me thousands of dollar for the same task - and a way longer time. So maybe the valuations are overblown, but they clearly provide value.
Yes, it's probably comparable to 4.8 if you are just using it to write code and put up a couple pull requests. That's not where things are now.