upvote
> design

I force OpenAI models to use image generation for design, then an iteration loop until it matches the image gen.

This is frustratingly manual and takes many more repetitions compared to Claude (and especially Claude Design) which "just work", but it's a big step change over the default.

reply
The UI design gap is something I’ve noticed as well, in things as simple as ASCII diagrams. Claude has a more human touch. All the diagrams GPT 5.6 generated for me were dressed up lists with too many pipe symbols.
reply
have you tried using lower efforts?
reply
Not the person you're responding to, but I have the same feelings they do and to answer your question for me at least, yes.

IMO 5.6 Sol had this weird dead zone between medium and high where medium under engineered and took short cuts and high over engineered and ignored instructions it didn't agree under the guise of trying being helpful. The whole 5.6 line was the first release from OpenAI where it felt like reasoning level really mattered and was incredibly finicky.

I haven't felt similar issues with GPT 6 though and am very happy with Astra low/med/high as my default choices depending on the task.

In general, I felt like with 5.6 the effort level did less than previous to make the models smarter and more just increased the complexity of the response. I have a half joke theory based only on vibes that OpenAI splitting 5.6 into Sol/Terra/Luna is where the intelligence split happened and so the effort levels were just like "think harder about the decision you already made". So like if the model decided the earth was flat on low effort it'd just say something like "the earth is flat because the horizon is flat". If it was on xhigh reasoning it'd give you a massively complex answer about how the sun reflects light because of the ozone layer and why people flying in planes can see a curve. In both cases though, adding more effort wouldn't get it to realize the earth was round. It just made the answer about it being flat more complex.

To be clear, that theory is not meant to be taken too seriously. It's not based on anything other than vibes. It's just my way of explaining to myself something I'm frustrated about to myself.

reply
Astra 6 was a huge improvement over Sol 5.6 for UI work. I haven't tried Sol 6 yet for it (it's only been a few minutes).

The GPT / Codex models have always been "overengineer" personalities. I prefer that to "I left a pile of race conditions lying around and big gaps in testing" though, which is what I was getting from Opus at times.

But yes both Astra and Sol veer on the side of paranoid. And honestly that's better for team work. For solo work where you just want to yeet something, it can be tiring.

You learn to tame the GPT "personality" on this front by combing over once a week and asking it to find and exterminate pointless tests, clean abstractions etc.

reply