upvote
They never nerfed any model after release.

The lackluster GPT-6 Sol has been superseded by this apparently much better 6.1 Sol within a week.

I am very skeptical of claims that old models weren't much worse. Compare this to February's GPT-5.3.

reply
I am comparing to GPT-5.3 and 5.2, and I perceive that things have not been noticeably better since then. I also know that I can predict new model releases with high accuracy when my coding agent suddenly becomes regard-level at following instructions and completing simple tasks. This is how I knew 6.0 was about to be released - 5.6 suddenly got unusably bad.

I could point out that I said 6.0 seemed good only in comparison to nerfed 5.6 - people would say I’m just a RSI denialist - but now it is in vogue to accept that 6.0 sucked now that 6.1 is out.

reply
How large of codebases are you working on? The models have gotten good enough to 1 shot stupid "trivial" throwaway integration projects with 0 handholding (was having RL'd garbage in late 2025), and I'm actually enjoying designing bounded greenfield personal software from scratch with Astra, in my experience. It's quite slow - 2 weeks of credits and constant talking and back and forth with Astra, but it doesn't feel annoying to talk to and is like an intelligent colleague maybe 70% of the time? Which is great. Just push back when it's dumb.

I'm by no means an AI booster, but given 2022 - 2026 progress I'd say it's "exponential" in the sense of, "holy shit, every year I can do more and more genuinely different things", not "RSI mind reading intelligence can do anything is here".

I don't think Navier-Stokes level intelligence translates over to my projects, unfortunately. Yet? Who knows.

> I haven’t seen actual capability growth since ~January, and I’m pretty sure that was all tooling/harness improvements.

Even if that were the case, I'd say that it's improved in practice. And just from a philosophy perspective, if you're trying to imply some kind of mind dualistic way of viewing things, uh, I disagree with those theories of intelligence strongly (which also incidentally also disagrees with AIT-style theories of intelligence on one axis, though I have many bones to pick with the culture there).

reply
I work on a very large code base (millions of LOC) and I've had lackluster results with autonomous work and 1 shotting. AI is definitely fantastic at working on many programming problems but I am not seeing amazing results at refactoring. In fact, I am seeing very poor results, even with Astra, even with extensive planning docs. All the recent models I've used can definitely get that refactor done, but not autonomously. It needs to be small slices. I've yet to see it 1 shot anything really complicated.

Here's a good example with some assumptions on my part: I work in C++ and it really feels like the models are trained so hard to keep everything compiling all the time. That's a huge negative in my opinion because what happens is that the AI will do things like use wrappers to keep things compiling, even when that basically results in creating or hiding abstraction leaks. Or they get sneaky and include a header they shouldn't. Or they actually do see that there should be a layer boundary and they write some kind of abstraction to cross it but the abstraction itself is garbage or doesn't follow existing API patterns. The AI could invent 10 different, new patterns when there is already 1 existing pattern they should use.

I feel like a lot of this involves a lot of babysitting prompts. Not that there's anything wrong with that of course.

reply
It’s 50/50 on whether it will fuck up implementing an integration test suite when given a list of tests to write and examples of existing tests. It still adds needless abstractions (the reference count codelens in VS Code is good for detecting this sort of thing).

On these metrics it is much better than it was in March of 2025 but no better than it was in March of 2026.

5.6 Sol in the last two weeks became much dumber such that what used to be one correction turned into endless rounds of corrections before just giving up and coding it manually. I’m mostly having it do the “chore” part of coding so it is disappointing that it isn’t better at that.

reply
This is 100% absolutely my experience as well. Especially the needless abstractions and endless rounds of corrections. That was literally my entire last week of work.
reply
> It’s 50/50 on whether it will fuck up implementing an integration test suite when given a list of tests to write and examples of existing tests. It still adds needless abstractions (the reference count codelens in VS Code is good for detecting this sort of thing).

Yes, still running into this, but surprised about this

> On these metrics it is much better than it was in March of 2025 but no better than it was in March of 2026.

I was super hyped at the agentic thing a year ago (Fall 2025), but designing functional software was hell. It would not just "grasp" the right level of "here is the essence of what we need" versus "these are all the small impl details". But idk I feel like Astra's the first model in quite a while that I don't feel genuinely annoyed at handholding a toddler with a PhD.

But I totally believe you on the 50/50 thing. Even recently as a few days ago, Astra did the thing where it ran into an error, and instead of making the sensible bounded decision of "make user retry in this case", it silently built an extremely elaborate recovery state machine w/o looking. These pathologies by no means gone, and I'm still careful in the design phases (which themselves are bounded and incremental) to sus out if Astra's gonna do this kind of RL slop failure mode.

For my use cases personally though, it's been better and better. I can't use AI at work, so you have much harier edge cases than I do, but still.

reply