Have I understood correctly that you trained only the logistic classifier, but didn't need to train the embedding model?
If so, I'm curious whether you compared that approach (A) with:
B) Jev only, with a single output.
C) Jev with multiple outputs fed into a logistic classifier.
Obviously C has cons (can't be self-hosted, needs some up-front work on deciding the shape of the output) but it might be somewhat more interpretable. (And I suppose it might have better performance?)