Nonetheless, this is very cool work! If I can offer a small suggestion to the team at Cactus, it would be to evaluate your releases on some usability criteria (including false positives). Any serious integrator or adopter of these models would want to have that information available.
OP and the linked page talk about the confidence score and using it as an action threshold, so it looks like an appropriate total response to me.
I am not sure how a micro model will fundamentally solve it. Would love to understand what dannyw and team did there?
> I'm hungover
{ "function_calls": [ { "name": "lock_door", "arguments": { "door": "front door" } } ], "reasoning": "User wants to lock the door. 'hungover' implies a security door. No specific door named, so use 'front door' as default.", "confidence": 0 }