upvote
The author explains this very well tbh:

> Today's models are now good enough for high-quality unattended tasks. Chasing the latest and greatest is silly. It is fun to see the new Fable capabilities, but the tasks we throw at them are usually ridiculous (maybe even insulting) if you believe in LLM sentience. It's like asking a math PhD to organize the files on your desktop.

I'm using DS V4.1 Flash as my main model since their release and it works great for all my coding tasks. My setup is OpenCode Go subscription and obra/superpowers skill.

The only times I try to change models are on general planning tasks (like research this codebase for tech debt mitigation opportunities) or if I need deep research which would benefit from searching the web, in which I still think Gemini is still the best because of the speed and access to google search index. But these are not even 20% of my daily tasks.

reply
I have had good luck rotating between the three tiers of GPT 5.6 with occasional jumps up to Astra or Fable. Most of my work is with GPT-5.6-Sol. Simple tasks like data extraction or trivial refactors (rename this variable etc) I often push down to Terra or Luna, or even to self-hosted Gemma4:31B. Very tricky stuff, like planning a new feature, design review, or code review of a complex change across multiple repositories is where I leverage Astra.

I've had middling success with models like DS V4.1 Flash and free Gemini. They tend to be pretty good at very easy stuff ... but they're more likely to go down rabbit holes, confidently assert falsehoods, or fix bugs with changes to my test harness rather than my code.

I asked about comparison to the well-known SotA models specifically because I use either Astra or Sol for ~85% of my daily tasks. When I try to use smaller/cheaper models, I have had very mixed success. Sometimes it's perfect, while other times it fails in subtle and hard to catch ways.

reply
Of course there's still a huge performance gap

but DS 4.1 Flash is good enough for most tasks

reply