upvote
I fully believe that these models perform better in benchmarks versus their predecessors, but in real world usage inside of real, production codebases? They feel just as flawed as ever. I honestly have not seen any significant improvement in a few months. The last thing where I felt "wow" was `/fast` mode and Deepseek.
reply
Wholeheartedly agree. Astra was some improvement over 5.6-sol in the sense that I'd "argue less" with it, but still frustrating and still sloppy. I'm starting to feel people are not honest about their experiences, they do very simple things or have very low standards. The biggest improvement i've seen from Astra so far is speed.

My experience with agentic coding on projects I care about (because my responsibility in my firm is to care about these things, at least for now) has not changed a lot in the past few months, and I have kept up with every single model update / experimented with harness a great deal.

reply

    > I'm starting to feel people are not honest about their experiences...
I think it falls under:

1. They don't actually look at/care what the agent is producing as long as it works (not planning on maintaining/ops yet).

2. They are using 3rd party benchmarks (which is fair given how widely real-world workloads change from day-to-day, feature-to-feature, making it difficult to really know how well the models would perform).

3. They are doing greenfield work where there is no scaffolding, no existing code, no legacy code, nothing to guide the agents along. I believe in these cases, new models can possible do better from a blank slate. But in existing codebases, I feel like the agents are more likely to simply follow existing patterns and existing guidance to begin with so things are a wash and more reliant on harness and existing code hygiene.

reply