upvote
I literally only make it halfway through the week until my weekly usage runs out. This is using only Opus, no fable, and I'm on the max x20 plan. It's become ridiculous.
reply
These comparisons are meaningless

I use Opus every day and easily have most of my weekly limit left over at the end of the week

reply
just buy two 20x, no?
reply
or switch to codex
reply
I am on the Claude Max 20x plan, and this still happens when using Fable 5/Opus 5. I would run out of weekly quota in 2 days, whereas Opus 4.8 would last the entire week, and sit at about 80-90% at the end.
reply
Agreed. The area I think will become more prevalent in the future for organizations are cost per intelligence -- effectively efficiency. An unoptimized model that costs 90x more than another that is only 10-15% less intelligent is something I would say is not a good deal.
reply
GLM 5.3-flash fits the bill
reply
What is it equivalent to?

What kind of things are you using it for?

I haven't tested it yet but on all the benchmarks it looks like it's 5-7x slower for agentic tasks.

reply
I made some webapps with it, and have it running my hermes agent (which also does a lot of coding, but not webapps).

Not sure what it's equivalent to, but it's super cheap and I am happy with the results

reply
It's a mix of slightly worse kimi k3 for UI work and slightly smarter than luna for everything else.

But yeah, it's very slow. I've put it to work as an LLM-as-RAG agent.

reply
I'm with you, for what I usually do most models are already more than enough.

What I'm really keen on is better auto-reasoning so I don't have to constantly have the constant inner debate on which reasoning effort to pick for each task.

I seriously hate the none-low-medium-high-xhigh-max-ultra etc that we have now, with companies frequently recommending different ones on each new model release, etc.

It's apparently called Adaptive Test-Time Compute or Dynamic Test-Time Compute and companies are apparently working on it (according to some LLM :shrug:)

reply
Adaptive reasoning is known to be an extremely hard problem to solve, though. It requires you to predict whether a certain LLM, with a certain effort level, with a certain prompt, will give you the right answer.
reply
Have you tried Grok 4.6, if you're focused on token budgets? In a league of it's own for tokens/intelligence.
reply
Not for enterprise. Can't trust the company behind it with my data.
reply
SuperGrok quota is garbage for anything coding. I burn through my quota in a few hours with very mild use.

SuperGrok Plus is slightly better but doesn’t last me more than a few days. Even Claude Max feels leagues more generous in usage…

I haven’t tried SuperGrok Heavy because it’s too expensive

reply
All of the subscription AI platforms are trimming down quotas across the board to push users into higher tiers. Whatever they can do. Local inference needs to meet pricing sooner
reply
Try gpt 5.6 Luna max
reply