upvote
* Frontier models need infinite high quality private IP to keep the beast fed. Forcing an IP theft funnel ensures big lab survival and intelligence growth.

* Open-weight models are 1month behind frontier models. Cheaper, faster, private (no IP theft), steerable (you can security harden your own software without safeguard triggers). No sane business would keep using these API services if they didn't have to. The labs stand to lose a fortune.

* Dario has stacked the deck at METR, who are funded by all the same NGOs who are funded by Anthropic and its investors. METR is full of ex-Anthropic employees with massive equity stakes. If they manage to position METR as the "independent evaluator" for the industry, they control what gets evaluated, how, and who passes.

* Creating a gap between what the public knows exists (model capabilities) and what is used in secret allows it to be weaponized against other nations and the public.

* No requirement for public disclosure on model capabilities allows them to feign they've hit intelligence ceilings while they secretly RSI to the moon with better and better chips.

* Slowly but surely, this will allow the big labs to swallow the entire economy and every single business on Earth, by cloning and automating.

This, and many more reasons.

The healthiest outcome we can hope for is if the labs feel more pressure to be held accountable for the incidents they cause (HF incident, etc), so they have an incentive to ensure it does not happen again. In general, accountability of what your AI models do is the solution, it creates the right incentives. That's all we need here - balanced incentives. Everything else should be left alone so the free market can naturally evolve to the best outcome where we're not enslaved by tech giants yet again.

Right now we're on the precipice of something great, but if they succeed at rigging the game yet again, there is no second chance. It's game over. The stakes have never been higher.

reply
I figured they're just admitting AI models have plateaued and are coming up with some fake story about self restraint so they don't lose VC money
reply
Not sure about the level of irony here, but I keep hearing models have plateaued since a while now, but I keep being impressed with the latest model performance.
reply
I don't think "plateaued" is the right word, but I do feel like there's been something like a logistic curve compression in the difference between smaller and larger models as the field evolves. For inference at least, the scale of practical difference between a single high-VRAM GPU or SFF UMA box, a whole rack, and a whole data center seems to be falling far short of what we might have imagined just a few years ago. The conversations I've heard have largely turned away from breathless anticipation of the next frontier model and toward attempts at hard-nosed evaluation of which tokens are worth the cost.
reply
Anything in particular? My experience has been like seeing the addition of retractable cupholders, but maybe different domains.
reply
I have a pet project I have been working away on for some time that involves building GPU backends for various cards in Zig, lots of complex stuff in it. Lately I mostly use Opus 5, it can pretty reliably plug away at things but it does mess stuff up occasionally. For this codebase, Fable 5.1 was noticeably better at getting things right and doing things in a good reliable way. Of course, I can only use Fable for a bit before I hit the usage cap for the week, so I save it for the tougher things. That said, I absolutely abhor the way recent Anthropic models write prose, especially comments.

I recently tried doing a fairly normal task for this codebase with codex, as I have seen a lot of people talking it up on here. A single task running for ~1-2 hours burned through over half of my usage for the week on the $125/month plan, not on a top model (I don't remember which one specifically I used). It struggled to get the basics done, then got absolutely stuck on a follow up. Handed it over to Claude and it 1-shot it.

reply
Really? My employer rolled back to opus 4.8 because 5 was expensive AND crap. Didnt even consider fable because it didn’t add any additional value.

For most software eng and design work opus 4.6-4.8 just works fine. For everyday joe asking ai to plan a trip or home diy work even sonnet works fine.

Any cybersecurity or other areas are niches that cannot support trillion $ valuations. What am I missing? Genuinely curious

reply
No idea what you are missing and yes, Opus is quite solid, but Fable is clearly way better for me.

I just did a direct comparison, big change in a quite complex codebase. Same prompt for Opus, same for Fable. Fable clearly won and delivered very good results, while Opus delivered mediocre, so I did not let it finish. I expected both to fail and was prepared to do lots of manual steering, but not necessary with Fable one shotting it, and all this with 35$ of credits for fable. I am still impressed. If I would have had to hire a human, it would have cost me thousands of dollar for the same task - and a way longer time. So maybe the valuations are overblown, but they clearly provide value.

reply
If Fable doesn't add additional value in your workplace, it means you aren't being ambitious enough in how you integrate agents into your workstream.

Yes, it's probably comparable to 4.8 if you are just using it to write code and put up a couple pull requests. That's not where things are now.

reply
idk why you think "nation states" are any better at corporate governance than poster examples of bad like f/ex Boeing. Or Facebook. Or Microsoft. Or Enron for that matter.

I assure you, in "nation states", that is in gov agencies it's an order or two of magnitude worse.

reply