upvote
I guess the agentic coding benchmarks don't have many rewards for stopping and clarifying what the user wants?
reply
They do not, as they're aiming for full replacement rather than augmentation of human users.

Personally, I think this is a bad idea, but someone's gotta build the Machine God I guess.

reply
> aggressively useful ... in the name of helpfulness

But is that because of training, or can that be (also? mostly?) an effect of the "system prompt"?

reply