upvote
Real question: is there anybody that is both maintaining alpha-dev capability by keeping abreast of all these daily changes, while also reserving enough time to actually work?

Seems like we've reached the event horizon of whether AI advances are worth paying attention to.

reply
I enjoy using opencode go to play around with a lot of different models. I wind up using deepseek v4 flash for most everything, stepping up to minimax m3 if that doesn't cut it, finally preferring GLM for complex tasks or important planning I want to go right the first time

I recommend opencode or something akin to it to play with models. Any big model updates or hot new ones will naturally run across your desk that way

reply
I do but that's become my work is routing between all the models and end to end encryption, and making new models from these models https://trustedrouter.com/blog/synth-iris-prometheus-zeus
reply
I think the play now is to just try out whatever the best new model is every time you see a headline that fundamentally reorganizes your conception of what's possible.
reply
Are you saying we've reached peak Bike shedding?
reply
How about: The yaks have started shaving themselves, who can keep track of how good a job they are doing?
reply
I don't think you need to be keeping abreast of them really, you just need to be using the best model you can get enough tokens from, which for many people is Fable 5 @ $200ish, ideally fanning out implementation to cheaper models
reply
Yeah either the benchmark isn't very useful anymore or V4 Flash is a really, really good model.
reply
GPT 5.6 Luna is an extremely cheap and still very capable model.

A chinese model being in the same ballpark of capability at half the price sounds believable to me.

reply
It's significantly worse than Luna and quite a bit slower in some fairly involved tests I run.
reply
In my use, DeepSeek v4 Flash (which replaced the quite excellent MiniMax M3) lags behind GLM 5.2 & Muse Spark 1.2 (let alone Kimi K3). Also, K3 is a much bigger multi-modal model, while Flash is text-only and likely optimised for coding tasks.
reply
Yep, and the v4 flash final is about 2.5x slower than preview making it no longer a fast model, in fact slower than Luna and bigger models in many cases.

Spark is actually the interesting one imo. It's significantly better, also significantly faster. If you are ok with letting Meta soak up your data (which DS does too) it's also the same price.

reply
And now nobody seems interested in it because the price hasn't gone down

it's still $3/$15 for all providers on openrouter

because of some Kimi license

https://openrouter.ai/moonshotai/kimi-k3#providers

reply
Synthetic is offering $7/month subscription for this weekend (which includes K3), insane value for this price !

https://synthetic.new/?referral=kwjqga9QYoUgpZV

reply
Morph has it for a slight discount, apparently.

Uptime looks crap, though.

reply
I believe it's because they are below $20 Million revenue limit (which Kimi K3's license has)

So we won't see any price decrease unless Kimi changes the license of K3

reply
Not for long, Deepseek is saying they will have a significant price jump soon. They really shouldn’t do it because they are on the cusp of capturing the scalable API market.
reply
They need to be able to serve their market. The price increase is partly load shedding. If they improve their ability to serve their load, they can always drop it again, as OpenAI did with Luna recently.
reply
> as OpenAI did with Luna recently

My read is, OpenAI is neither able to claw b2b money (away from Ant) nor are they able to stave off open weights on the other. In short, they're struggling to hold onto their distant #2 position in the coding market, and these pricing changes reflect a (desperate) change in strategy.

reply
and i still won't use it, because they log and spy on your prompts XD.

the private endpoint costs 10x (azure).

private endpoints for deepseek (lots of providers) also cost about 10x more.

but 10x more for deepseek is $0.028 cached input, and 10x more for luna is $0.10.

reply