Here's a hint: confidence is not generated by a model.
But confidence value is just a function applied to probabilities. It is not coming from the model, and it carries no additional information.
It is documented btw, and yet you will see plenty of claims that Jev is better than LLM because it returns both.
Do you think though that Luna being a model post-trained for chat produces over-confidence in logprobs?
Typesafe claims that Jev is calibrated, but there are plenty of examples where it completely fails (predicting die roll being the most obvious one).
Unfortunately calibration is hard to benchmark.
If you instead give it a list of probability for each number and ask it whats the probability of each number, the result will be accurate.