upvote
Mature training pipelines, plus ever expanding RL datasets of increased quality, and mega GPU clusters to finish training in a few weeks. Automated safety and reliability testing.
reply
I do wonder if people switch back and forth between primary models (GPTvsClaude) that it may be a better idea to simply keep releasing updates as soon as possible in order to keep users from bouncing back and forth.
reply
This is it.

It's because they need subscription money and interaction data and so keeping a version bump in the wings to stop the bleeding from your competitor's version bump is the logical thing to do. It has nothing to do with RSI.

reply
[dead]
reply
Maybe process maturity too.

Like think about a software org with good CI/CD versus one without. The mature org can do consistent incremental releases because each one is safe and low overhead, the messier org will do fewer big releases because each release requires a big effort on its own.

As model developers mature we might expect to see more frequent point releases rather than the big bang evolutions.

reply
Probably one of the factors. Signed up to openai pro a few days ago, deciding between openai and anthropic, then sonnet 5.5 was released and am wondering whether I made a mistake.

Luckily it's not a mistake as now we have access to . . . dots.

(and sol 6.1, it seems)

reply
jokes on me, I pay for all the subscriptions.
reply
Opus 5.5 is better than they anticipated, it's faster, smarter, cheaper. I'm about to change provider for claude and I'm not the only one
reply
It feels like an updated 4.6. It's fantastic.
reply
> I'm not the only one

See, that's an/the issue. As soon as people start to flee to the improved model, they start to serve degraded models to keep up with the demand.

reply
No, we're pacing ourselves to have the time to evaluate the impact each new model could have, obviously.
reply
Response to DeepSeek’s technical paper and competition.
reply
Which paper are you referring to?
reply
What's that in summary?
reply
Not the person you're replying to, but judging by the emphasis on the cost of cached input tokens in the OP article, I'd guess it has to do with DeepSeek v4.1's KV cache efficiency. It uses <1000 bytes per token, so they're able to get 1M token context in under a GB.

Edit: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/...

reply
Both labs are spying on each other and they get jelly when the other is releasing a new model, so they have to ship something at the same time so they don’t look bad.
reply
Probably just the singularity, no big deal
reply
[dead]
reply
Versions is marketing, snapshots/minor variations are easy and the number must go up. Release timing is another OAI's marketing tactic.

>RSI

Recursive improvement doesn't imply increased rate, another word for it is "iterative" but this probably sounds too boring to some people.

reply
They're pacing the frontier
reply
and seems like they're claiming Sol/Opus are not frontier (and only Astra/Fable are)
reply
It is the only way to reduce prices while making it look like a good thing.
reply
Wanting to have the newer model than the competitor, presumably.
reply
The old "the bigger number is better", GPT announces model 6.1, the obvious thing to do next is to announce Gemini 27, and after that Claudé 3000, then a flute album.
reply
We swear, We Really Wanted To Make An "ASI" Model But This Is Literally The Way The Weights Dragged Us This Time
reply
The initial response to 6 Sol was bad, and Opus 5.5 was definitely winning the public vibes war. Makes sense to rush something out
reply
New models are distill from the actual unrelease frontier models. They are just giving us better checkpoints.
reply
No.
reply
[dead]
reply
Its a news cycle more than anything, and its ONLY going to get much, much worse. Daily releases, or multiple daily, 30-45, by EOY. Welcome to RSI!
reply
They're releasing Sol 6.1 because 1. Astra 6.1 got postponed 2. Sol 6 is shitty 3. They have to release _something_ in response to Opus 5.5
reply
Competition
reply
Anthropic’s IPO?
reply
Productivity is increasing as models get smarter; we are ascending the singularity. I'm serious.
reply
What I don't understand is how much people have to say about every single one. Aren't we at the diminishing returns stage yet? Is there really that much to discuss?
reply
If you look closely at various benchmarks, you'll see that often models will improve in certain areas while regressing in others. It suggests we're already at the point of diminishing returns.
reply
Chinese model pressure. Many of my SWE friends switched to Chinese models. I also use QWEN and GLM for many of the api requiring projects and dropped OpenAI and Anthropic. The only reason was the cost.

EDIT: I love getting downvoted by openai and anthropic employees or their bots.

reply
I can't recommend Chinese models enough. My personal favorite is DeepSeek v4.1 Flash but I have tried Qwen 3.8, Kimi 3 and GLM 5.3 which are equally impressive but DeepSeek is the cheapest and fastest regularly hitting 270 token per second.

And yeah I have worked with Anthropic and OpenAI models, they're good but they cost a fortune while Chinese models are already really good at a fraction of the cost.

reply
DeepSeek v4.1 Flash is fascinating and uneven. It's way too chatty in OpenCode to be a collaboration partner. I tried dsh-tui which feels comparable to the codex/claude tui's and it's usable. but it seems to be "brilliant and yet stupid" in a way I can't quite put my finger on. I've got too much real work to get done to dig into it so until the big boys price me out of the market I'm back to my $100/month deal.
reply
I was working exclusively with DS 4.1 Flash until Opus 5.5 got me back to a sub. I was disillusioned with what was available.
reply
I keep hearing about these Chinese models, but what exactly are you doing with the models and coding? I have a need to fully write code with full tool calling capabilities. Not just methods or functions. I want to be able to prompt a feature and it makes the JIRA ticket, and fully implements it and makes a PR. I don't want to babysit it or even read the code. Once it creates the PR, I want it to monitor it for any comments fro Copilot/security review and then fix it as necessary.

Is that what the Chinese models are capable of? If so, how are you using them? API? Or is there an inference provider that is as fast as the big 2? What about the coding harness?

reply
[dead]
reply