upvote
It is not super good in long horizon tasks and worse than 5.6 in our evals. It really failed the agentic evals where DeepSeek, Kimi and Opus are the winners.

It is great on creating summaries and content.

reply
Yeah, very similar benchmarks at 1/4 the price. I've been pretty happy with Luna 6 though I still think the gap between small and frontier models is larger than many people want to admit.
reply