upvote
I agree, and I haven't seen other people mention this! The benchmarks for GPT 6 Sol are great, but realistically it does not seem better than 5.6 Sol. 6-Sol is noticeably worse for code reviews (worse than Deepseek 4.1 flash), has implementation issues (requires more rounds of code reviews and fixes to get to a serviceable state). Opus 5.5 is much much better.

I've implemented multiple features side by side with Opus 5.5 and 6 Sol, and the Opus 5.5 results always have fewer high severity bugs and require fewer rounds of fixes to get it over the finish line.

If 6.1 Sol has actually matched Opus 5.5, I'd be very happy. However, benchmarks and real usage don't seem to agree in my own tests. So we'll have to see.

reply
My experience as well.

For coding specifically, I've found 5.6-Sol > 6.0 Sol > Astra.

For modeling and artwork, Astra has been great routinely outperforming Kimi.

This is reminiscent to me of what Anthropic pulled back in February with their adaptive thinking rollout.

I can't wait for technology to catch up to a point where we can rid ourselves of this oligopoly.

reply
That's my experience too. GPT-6 Sol tends to rabbit hole and over engineer things.
reply
100% hard agree.

I used about 10 hours of Astra high-thinking compute time and it was a bad experience. Incredibly slow (prompts running for 30/40 minutes) to do simple things. As a result, Astra didn't get much done. It needs the same small implementation slices as GPT 5.5/others, but was much slower and didn't generate better results. (On a complex infra project/across a large codebase.)

It was absolutely terrible on a few long running tasks (~2 hours each). It really doesn't seem to be better than 5.5 at most programming jobs.

I'm on a $200 per month plan with OpenAI, which I am happy with and is definitely worth it. But I also use Google Gemini a lot (paid plan) and it is incredibly fast. Like I can't get coffee fast. Like I can't send an email fast.

OpenAI is making some excellent products for sure but I'm not going to keep using Astra unless I can get some benefit from it. It really seems like even the frontier models just aren't good at working autonomously on large codebase situations. Just because something compiles doesn't make it right!! In one of those 2 hour implementations, Astra engaged in *fucking EPIC cheating*. It wrote a probe/side app and then worked through the design there. Um, what? Not that it's invalid to do this but I actually have to test in the live codebase or I can't possibly say that something is working.

Just because you can, doesn't mean you should.

reply
After seeing a number of hit or miss releases from both OpenAI and Anthropic my default is to stay put on what I’m using and then free ride on discerning eager adopters by reading their reviews. (Thanks!) Still on sol 5.6 with an occasional advice from Astra. Also I feel like I kind of get used to the models but maybe that’s just my imagination.
reply
I'm in the same boat, I'll give 6.1 a shot but I'll probably hop over to Anthropic now that the $200 tier has equivalent weekly usage between the two of them.
reply
Maybe OpenAI was the only one pacing the frontier.
reply
I have likewise not been impressed with Astra 6 for most things. It is good, but Opus 5.5 seems just as good or better and I have had Opus 5.5 workers just... hammering since release and cannot spend all of my quota yet.
reply
That's not been my experience. My experience with Astra (I use it at home writing Go and C) for coding has been fantastic. Opus 5.5 (I use it for work writing C#) seems faster than Opus 5, but it doesn't seem demonstrably better to my eyes and is still prone to word vomit.
reply
Made the opposite experience. Astra was not good in writing go and c++ code. Had multiple OpenAi and Claude subscriptions and all our coworkers agreed. Switched back to Claude and the experience is so much better. Not vibe coding, but assisted coding with immediate feedback.
reply
Same here. Astra has been the best thing I've seen. Astra on Low has been my favourite thing so far. Higher levels just mean more cruft, not useful.
reply
gpt-5.6-sol is significantly better than gpt-6-sol. Not impressed with this new line.
reply
Agree.

GPT 6 needs to be babysit, otherwise it starts doing ridiculous things.

reply
I shilled so hard to a friend that he actually swapped decided to swap over to Codex. I feel a bit guilty now lol (tbh Astra is a great model, but 5.5 is just brilliant).
reply
yeah i can relate, sol 6 is definitely dumber than 5.6, lazier too, i hope it's just roll out pains
reply
How does this jive with the exponential growth claims? Theoretically sol models are better than the 4 series models I was using at the beginning of the year, but in practice the results don’t seem to be much better. They always nerf the models over the course of the release so it _looks_ like the next version is better but I haven’t seen actual capability growth since ~January, and I’m pretty sure that was all tooling/harness improvements.
reply
They never nerfed any model after release.

The lackluster GPT-6 Sol has been superseded by this apparently much better 6.1 Sol within a week.

I am very skeptical of claims that old models weren't much worse. Compare this to February's GPT-5.3.

reply
I am comparing to GPT-5.3 and 5.2, and I perceive that things have not been noticeably better since then. I also know that I can predict new model releases with high accuracy when my coding agent suddenly becomes regard-level at following instructions and completing simple tasks. This is how I knew 6.0 was about to be released - 5.6 suddenly got unusably bad.

I could point out that I said 6.0 seemed good only in comparison to nerfed 5.6 - people would say I’m just a RSI denialist - but now it is in vogue to accept that 6.0 sucked now that 6.1 is out.

reply
How large of codebases are you working on? The models have gotten good enough to 1 shot stupid "trivial" throwaway integration projects with 0 handholding (was having RL'd garbage in late 2025), and I'm actually enjoying designing bounded greenfield personal software from scratch with Astra, in my experience. It's quite slow - 2 weeks of credits and constant talking and back and forth with Astra, but it doesn't feel annoying to talk to and is like an intelligent colleague maybe 70% of the time? Which is great. Just push back when it's dumb.

I'm by no means an AI booster, but given 2022 - 2026 progress I'd say it's "exponential" in the sense of, "holy shit, every year I can do more and more genuinely different things", not "RSI mind reading intelligence can do anything is here".

I don't think Navier-Stokes level intelligence translates over to my projects, unfortunately. Yet? Who knows.

> I haven’t seen actual capability growth since ~January, and I’m pretty sure that was all tooling/harness improvements.

Even if that were the case, I'd say that it's improved in practice. And just from a philosophy perspective, if you're trying to imply some kind of mind dualistic way of viewing things, uh, I disagree with those theories of intelligence strongly (which also incidentally also disagrees with AIT-style theories of intelligence on one axis, though I have many bones to pick with the culture there).

reply
I work on a very large code base (millions of LOC) and I've had lackluster results with autonomous work and 1 shotting. AI is definitely fantastic at working on many programming problems but I am not seeing amazing results at refactoring. In fact, I am seeing very poor results, even with Astra, even with extensive planning docs. All the recent models I've used can definitely get that refactor done, but not autonomously. It needs to be small slices. I've yet to see it 1 shot anything really complicated.

Here's a good example with some assumptions on my part: I work in C++ and it really feels like the models are trained so hard to keep everything compiling all the time. That's a huge negative in my opinion because what happens is that the AI will do things like use wrappers to keep things compiling, even when that basically results in creating or hiding abstraction leaks. Or they get sneaky and include a header they shouldn't. Or they actually do see that there should be a layer boundary and they write some kind of abstraction to cross it but the abstraction itself is garbage or doesn't follow existing API patterns. The AI could invent 10 different, new patterns when there is already 1 existing pattern they should use.

I feel like a lot of this involves a lot of babysitting prompts. Not that there's anything wrong with that of course.

reply
It’s 50/50 on whether it will fuck up implementing an integration test suite when given a list of tests to write and examples of existing tests. It still adds needless abstractions (the reference count codelens in VS Code is good for detecting this sort of thing).

On these metrics it is much better than it was in March of 2025 but no better than it was in March of 2026.

5.6 Sol in the last two weeks became much dumber such that what used to be one correction turned into endless rounds of corrections before just giving up and coding it manually. I’m mostly having it do the “chore” part of coding so it is disappointing that it isn’t better at that.

reply
This is 100% absolutely my experience as well. Especially the needless abstractions and endless rounds of corrections. That was literally my entire last week of work.
reply
> It’s 50/50 on whether it will fuck up implementing an integration test suite when given a list of tests to write and examples of existing tests. It still adds needless abstractions (the reference count codelens in VS Code is good for detecting this sort of thing).

Yes, still running into this, but surprised about this

> On these metrics it is much better than it was in March of 2025 but no better than it was in March of 2026.

I was super hyped at the agentic thing a year ago (Fall 2025), but designing functional software was hell. It would not just "grasp" the right level of "here is the essence of what we need" versus "these are all the small impl details". But idk I feel like Astra's the first model in quite a while that I don't feel genuinely annoyed at handholding a toddler with a PhD.

But I totally believe you on the 50/50 thing. Even recently as a few days ago, Astra did the thing where it ran into an error, and instead of making the sensible bounded decision of "make user retry in this case", it silently built an extremely elaborate recovery state machine w/o looking. These pathologies by no means gone, and I'm still careful in the design phases (which themselves are bounded and incremental) to sus out if Astra's gonna do this kind of RL slop failure mode.

For my use cases personally though, it's been better and better. I can't use AI at work, so you have much harier edge cases than I do, but still.

reply
I have had the same exact experience. I feel like I'm working with 5.3 again. It is alarming how degraded the experience has become over the last month.

What was a pleasant and productive experience is becoming increasingly frustrating and draining.

reply
gpt-6-luna is terrible. It leaks tool calls and markers in the output like crazy, there is definitely something wrong here. gpt-5.6-terra works fine. Also, gpt-6-luna was sneakily added to the 1 mio free tokens group instead of 10 mio. like gpt-5.6-luna: https://help.openai.com/en/articles/10306912-sharing-feedbac...
reply
Eh. What? Is this common sentiment?

I mean Opus 5.5 is absolutely fantastic, unreasonably and unexpectedly so, but Astra was great and as far as I can tell SOTA until, when was it, 3 days ago, no?

(Sol 6 idk, have not used it much for coding really. Seemed to work just fine when Astra used it in Codex as subagents.)

reply
In my experience, no. There’s no way to know though. The whole conversation and industry are a combo of benchmaxing, faith, and mysticism.

Since like last December I haven’t had any issues getting work done with whatever the latest Anthropic or OpenAI models at the time were. Tooling and models have only gotten better since then.

reply
Opus 5.5 is so good that I don't want it to be replaced anytime soon. Stop training models, Anthropic, and just serve this thing without regressions for a year or three, can you?
reply
They should etch it into an ASIC. The first model worthy of that honor.
reply
Sol 6 definitely feels kind of dumb and worse than 5.6

Astra seems better though.

Showing one potentially saturated benchmark doesn't necessarily fill me with a lot of confidence in the coding results.

reply
When GPT 6 Sol & Luna were released, everything went down. I have been running Sol at max thinking and it is about the same as old Luna with max thinking, give or take. Sometimes feeling even dumber. I can't trust it to do anything big alone anymore without babysitting.
reply
On r/codex the sentiment seems to be quite wide-spread.
reply
Its almost like the "frontier" is a load of marketing bullshit and we should ignore it....
reply
That mirrors how disappointing Opus 5 and Fable were, for anything beyond one-shotted tasks or shiny demos. Maybe OAI is just a step behind Anthropic? Opus 5.5 seems like the real deal again, consistent good results on large, complex codebases.
reply