You’re guessing that it’s a result of advertising, and I agree that that’s probably a component, but it’s a mistake to assume that they are interchangeable when you have people saying to you directly “I use both and they’re not.”
What matters most in state of the art models isn't simply the final destination, it's the process of how one arrives to that destination.
I would argue the process these days has more to do with the harness than the model, at least when we're talking about the SOTA options. Claude Code's biggest advantage isn't Opus, rather it's the shared knowledge the community has been building and sharing around using it effectively. Almost all of the out-of-the-box tutorials and skills and frameworks are build for Claude first, then Codex maybe.
I'd go further and say that CC and Codex are not even the best harnesses available, they just offer the most subsidized rate plans.
This. Never underestimate the ability of a large number of power users to substantially improve the actual utility of a complex software product.
They always have more time (and sometimes more skill) than a product's developers.
Sometimes the quantity of monkeys matters more than the quality of the typewriters.
To bring ng this back to the discussion at hand (and to be redundant, as it's been mentioned here already), there are many aspects of using an LLM that are not purely about the output from a single or few well formed prompts. Additionally, if the end results are very similar, these othrr aspects will have an outsized influence on people's perspective of the tools, as they're the only differences worth choosing one model over another.
In fact, after seeing all these comments about the amount of effort, you redirected at calling that mere "vibes:
> Edit: i bet 99% of people here, if presented with a test where i gave 5 models but all of the results came from one, would not be able to discern this. Just vibes all the way down
Which, again, is a highly emotional way to view people trying to say that the process matters too. Calling people "vibes based" or "highly susceptible to marketting" and saying they take part in "tupperware parties" rather than evaluating their experience with tools is quite a thing to see, a complete dismissal of professionals' core experience as "vibes" rather than something intrinsic to how they perform labor.
Some examples are blind wine tasting tests. There are instances whereby some journalists invited renowned/established wine tasters and subjected them to blind wine tasting tests. Turns out the judges couldn't tell which was which. Pretty embarrassing.
It speaks volumes as to how people can accurately judge the value of things. There is research by some network scientist that says you can't generally can't tell the 1% from the top, though you can tell the really bad from the generally good. What OP's experiment might tell us is that the LLM competitive advantage is so small no one can tell which is objectively better.