upvote
This model was likely trained months before deepseek released their paper.
reply
Models are generally posttrained to a window shorter than you think.
reply
Doesn't mean they didn't apply something similar. They could have also come up independently with their own version, the speculation is not they copied it, rather that they have performance breakthroughs which perhaps is a result of work in same domain
reply
Unless they already have something similar of their own, which is always possible, they'd be stupid not to. I don't suppose we'll ever know, though. It would not be a good look if after the trillions of dollars that have been thrown at US labs, investors found out that they're down to copying Chinese tech.
reply
I don't think it's particularly relevant?

They might be using something like this, or they might be using some other "increased sparsity" techniques, of which there are a great many. They also might be optimizing for something else - like less RAM use for KV cache.

Alternatively, they might be cutting into their margins and dropping the price because of stiffer competition from Astra. I do think that's unlikely though.

reply
Getting a 403 on that link. Mind checking it once?
reply
https://web.archive.org/web/20260922172456/https://miraflow....

tired: AI startup attempting to build their own website

wired: a nonprofit founded in 1996

reply
I'm able to access it on my laptop at home. Maybe a misconfigured bot protection rule, try a different user-agent or IP?
reply
It's working for me (based out of California).
reply