upvote
The idea is that a thread can have many reasonable follow-ups that the user would've accepted, so it is wrong to punish the model for predicting a follow up that is different from the user message, as that prediction could've been accepted by the user if it was given.
reply
The user actual answer vs the user actual answer after seeing the suggested answer are different points of data
reply
You can press Tab+Enter to accept it.
reply
Seeing the suggestion influences the decision
reply
i believe they sometimes show no suggestion at all, fwiw.
reply
RL training, the second phase of LLM training, is based on "I did X, was that good/bad?" and that 1 bit of information is the training data.

So you give the user a suggestion, and the user accepts -> good

You give the user a suggestion, and the user refuses and types something else -> bad (plus some supervisory training data)

The main performance enhancer in LLMs is getting high quality training data. So, first, any extra training data will help. Second this is training data that's directly relevant to their product, and thus higher quality than many other sources.

I'd believe any model provider is mining the shit out of every last customer interaction they can get, not just this.

reply