undefined

points

[-]

They've been around a couple years. This is the first model that has really broken into the anglosphere.

Keep a tab on aihubmix, the Chinese openrouter, if you want to stay on top of the latest models. They keep track of things like the Baichuan, Doubao, baai (beijing academy), Meituan, 01.AI (yi), xiaomi, etc...

Much larger chinese coverage than openrouter

by tarruda12 hours ago|

parent|

[-]

> This is the first model that has really broken into the anglosphere.

Before Step 3.5 Flash, I've been hearing a lot about ACEStep as being the only open weights competitor to Suno.

by Havoc16 hours ago|

parent|

prev|

[-]

>first model that has really broken into the anglosphere.

Do you know of a couple of interesting ones that haven't yet?

by kristopolous16 hours ago|

parent|

[-]

doubao (bytedance) seed models are interesting

Keep your eye on Baidu's Ernie https://ernie.baidu.com/

Artificial analysis is generally on top of everything

https://artificialanalysis.ai/leaderboards/models

Those two are really the new players

Nanbeige which they haven't benchmarked just put out a shockingly good 3b model https://huggingface.co/Nanbeige - specifically https://huggingface.co/Nanbeige/Nanbeige4.1-3B

You have to tweak the hyper parameter like they say but I'm getting quality output, commensurate with maybe a 32b model, in exchange for a huge thinking lag

It's the new LFM 2.5

by admiralrohan13 hours ago|

parent|

[-]

Never heard of Nanbeige, thanks for sharing. "Good" is subjective though, in which tasks can I use it and where to avoid?

by kristopolous13 hours ago|

parent|

[-]

it's a 3b model. Fire it up. If you have ollama just do this:

    ollama create nanbeige-custom -f <(curl https://day50.dev/Nanbeige4.1-params.Modelfile)

That has the hyperparameters already in there. Then you can try it out

It's taking up like 2.5GB of ram.

my test query is always "compare rust and go with code samples". I'm telling you, the thinking token count is ... high...

Here's what I got https://day50.dev/rust_v_go.md

I just tried it on a 4gb raspberry pi and a 2012 era x230 with an i5-3210. Worked.

It'll take about 45 minutes on the pi which you know, isn't OOM...so there's that....

by Havoc12 hours ago|

parent|

prev|

[-]

Thanks!

by 0x199720 hours ago|

prev|

[-]

https://en.wikipedia.org/wiki/StepFun

by SilverElfin20 hours ago|

parent|

[-]

Thanks. Do they sell any of these products today or is it more like research? I am not able to find anything relating to pricing on their website. Just a chatbot.

by 0x199719 hours ago|

parent|

[-]

Princing can be found on their docs website https://platform.stepfun.ai/docs/en/pricing/details

by tarruda12 hours ago|

prev|

[-]

They seem to be the same company that released ACEStep music generation model: https://acestep.io/

Though the only mention I found was in ComfyUI docs: https://docs.comfy.org/tutorials/audio/ace-step/ace-step-v1

by deaux20 hours ago|

prev|

[-]

Might want to give it a search.