upvote
Same exact experience. I worked with both Fable and Sol foe the last two months, daily for several hours, and got used to the very bright, quick thinking, proactive even.

As of last 2 weeks or so both models are nearly on par with DeepSeek4.1 now, which I also use a lot. They're still better, but that difference is not as pronounced as before and, importantly, the frustration level is now on par.

Whatever they're doing will surely drive people to less advanced but predictable, self hosted open models. I sure would rather use DS4.1 with Qwen/GLM in adversarial mode than deal with this b/s I pay significant amount of money.

Me and my friends have been contemplating on getting an Ultra M5 256 and splitting the cost. PI harness is so good now that this is really a viable alternative.

reply
Same experience with Sonnet on low effort. It used be when I used a "table_name/id" format to reference a db record, it knew exactly how to find it using connected mcp tools. Today it failed 4/4 times (I tried the exact same prompt in 4 separate sessions and each time it replied "I don't have access to [...]"). On medium effort it got it right the first time.
reply
& the nice thing about Hermes (since its open source) is you can be reasonably sure that behavior change is coming from the model and not the harness. (probably)
reply
I'm fairly sure it's just luck of the draw if you get put onto a quant'd model or not. I've seen luna xhigh change intelligence fairly drastically on a day to day basis.
reply
Really this is the base problem. You have zero idea where and how your prompt is being executed.

If for example AWS sells you a 2xLarge server there may be some variability in performance but it's going to be averaged out very well.

When it comes to AI services executing your model there is absolutely no information on what and with what settings your model is being executed. Hell, you have no idea if it even is the model you're paying for. Add that models are not deterministic so variability can be pretty large.

This leads to a common set of dynamics that induce cheating behavior in humans. For example, is there a mix of different hardware. Does lessor hardware use different settings? How do you know xhigh is what your prompt ran under. Anthropic has a proven history of running your prompt silently under different models.

This is a huge mess that needs and will be regulated or sued heavily over. Hell, with as many people out there that hate AI it might be easier than one thinks to have a state sue the providers on this and elicit a huge amount of discovery.

reply
deleted
reply