The cost is unbelievably low, and the quality of intelligence I get is equivalent to when I was working mostly with Anthropic models (late last year/early this year). I'm fully invested in MiMo and I'm very happy with it.
-- PS: I also check almost daily to see if other models are capable of doing such great work. And they do – DS4F is powerful and DS41 is impressive, GLM 5.3 Flash gets a job done well, etc. – but when I add cost of M-token in the ROI math, Jeez! MiMo is an order of magnitude better.
These days, there are more intelligent models like DS4.1, but Mimo is very obedient, so I plan things with another model and give the implementation to Mimo.
I’ll give Mimo a try.
UltraSpeed was absolutely awesome. I miss it.
DS 4.1 Flash is amazing. Well worth the extra cost.
And I am not a web developer! It's an extraordinary model.
(Mouse and keyboard required)
It’s good enough that I’m considering a second spark, or selling this and buying an M5 Ultra with 256GB for it
I always found that those Mimo models to be really good at tool calling and following instructions
API may be expensive, but I do 900m tokens (95% cached, ~0.4% output) on Z.ai's $18/mo coding plan with GLM 5.3 Flash.
For comparison, I am currently at 6.6B tokens, 95% of monthly quota on a 10$ command code plan, mostly using DeepSeek flash 4.1, or some of the free models for easier tasks.
now we r just noticing the grave getting dug deeper.
Trying to understand why users are using Luna when Sol seems essentially unlimited on the pro plan. Unless you have jobs running 24/7.
I had used 2.5-pro for a hefty chunk of development, and found it to work like a somewhat forgetful senior engineer who was new to my project. Very capable, would almost always choose a reasonable option, if not always the best one for the project, and not great at multi-tasking. Generally, made me comfortable not scrutinizing the code line-by-line, but still needed a bit of steering once projects got to a reasonable size.
The next model is a clear step up in the multi-tasking capability at least, with me very rarely having to steer the implementation of a well-defined issue. In terms of code, I found MiMo-V.2.5-pro to be extremely conservative, implementing minimal solutions. The next model seems a little bit more ambitious, in positive ways, making good guesses about gaps/next steps. It also seems to be a fair bit better at design, at least for the little bit I've done, it was good at translating my concepts to practical elements on screen, and cleaned things up nicely as I made suggestions.
Fable scores 70%, Kimi K3 69%, Astra 74% (all on max effort).
Even flash reached 60.7% by step 12, and it's on step 16 now.
This is so exciting lmao.
Also, Anthropic and OpenAI probably want to keep each other on their toes so they don’t end up on the wrong side of another Opus 4.6 / GPT-5.3-Codex situation, where one lab releases a model only for the other to drop a better one hours later.
I still opus 4.6 though not for code
I’m saying who has a million dollars for me, so I can make my own model?
I'd guess everybody uses at least some benchmarks as stopping criteria, which is kinda sensible, but it also does induce some benchmaxxing, and explains partly why the newest models always tend to eke out in benchmarks.
https://en.wikipedia.org/wiki/Training,_validation,_and_test...
It's not the direct feedback loop of RL but its not far.
For some reason I thought training took much, much longer than what the progress bar suggests.
This is really neat, I'm currently using mimo 2.5 pro, and it's decent (or great given the price). Hopefully their next one is multimodal.
If you are going to develop a near frontier model, and you don’t think you have special sauce up your sleeve, why not making training runs and RL environment scores etc. visible to the world?
I’m genuinely learning quite a bit just from the dashboard
Keeping the garage door open, or at least making the door translucent. It's always cool.
Google had this GPT long go and a wise man within Google noted:
"We don't have any maot neither does anyone else."
The AI bubble burst is guaranteed and is only delayed by IPOs.
Open models have not yet caught up with February's Mythos checkpoint.
Meanwhile OpenAI is solving millennium problems, and their compute is still fully utilized.
Where is the cool shit from the US labs?
The US population is much more pessimistic and doomsday driven these days, whereas the Chinese are more optimistic and future driven.
In the short term, true.
In the long term, unknown but typically when you hold progress that way while other countries don't you at best end up becoming siloed while the rest of the world continues on without you.
Anthropic and OpenAI literally stole from every human in history and youre out here complaining that the Chinese are distilling models and releasing them to the public?
Why do you care?
/s