upvote
Maybe I'm missing something, but why couldn't it be generated by the model? In older classification tasks with transformers like BERT, you could absolutely obtain a confidence score.
reply
Jev API returns both confidence and probabilities.

But confidence value is just a function applied to probabilities. It is not coming from the model, and it carries no additional information.

It is documented btw, and yet you will see plenty of claims that Jev is better than LLM because it returns both.

reply
Thanks, fixed my understanding!

Do you think though that Luna being a model post-trained for chat produces over-confidence in logprobs?

reply
Yeah, but I wouldn't be surprised OpenAI's decision API is a post-trained Luna with confidence calibration.

Typesafe claims that Jev is calibrated, but there are plenty of examples where it completely fails (predicting die roll being the most obvious one).

Unfortunately calibration is hard to benchmark.

reply
The dice roll prediction is about the way the prompt is setup misunderstanding how Jev works (they treat the confidence score as a probability score, which it isn't).

If you instead give it a list of probability for each number and ask it whats the probability of each number, the result will be accurate.

reply
> give it a list of probability for each number and ask it whats the probability of each number

Did i hear that correctly? In order for Jev to be accurate you have to give it the answer before asking for the answer?

(btw this is exactly how Jev is playing games).

reply