We are working on running our own benchmark of Jev and some of the other models. Our use case is classification that runs in a UI. Currently LLMs have good accuracy, but are too slow (and expensive).
Jev not being available through a cloud provider (Bedrock or similar) makes it more challenging for us to start testing and rolling it out.
So it's a new subprocessor. Which can often be painful to onboard, especially if not compliant according to your needs.