I don't really care that much about benchmarks, but having tested it on one of my puzzle prompts I can tell you that it solves it well, writes clearly, isn't noticeably slower than Qwen 3.6 27B and is
much more terse in its reasoning (which will help with preserve-reasoning).
It also has a knowledge cutoff inside this year.
The main limitation is the smaller maximum recommended context.