upvote
I don't see it fundamentally any different than knowing when to use a tool. Is this tool like RAG an important enough corner case to train for it? I dunno.

LLMs already shell out and write code to solve certain problems. This is just a special case of that.

reply
It's a special case for an LLM, and you can use an LLM with structure output to get similar results, but you can engineer specifically for that case to get better results per dollar for it. That's why there is little reason to adapt GPT 5.6 Sol or wathever for this task; it can already do it (at a high cost). For OpenAI to compete with Jev they have to maintain another line of models, something like "GPT-5.6-instant-decision", that is small, fast and cheap, in the scale of GPT-5 nano.

Note that I don't think OpenAI is incapable of doing it, but I just don't think they will bother with it.

reply
Keeping people looped into your product is pretty important, but yeah, there's not clean way currently to separate "structured" outputs from the token stream and to start using a different billing structure there. And I also appreciate that they aren't going to be keen on gving free or near free output either, so gotta figure that.
reply
If there's money to be made, I'm sure sama will find a righteous reason to offer it.
reply
system 2 is just an llm with a forced toolcall IMO
reply
In the olden days we call this classifier, usually assignment 2 of Machine Learning 101. BERT (well, GLiNER specifically) and diffusion are calling and want their Large Classifier Models back.

https://github.com/vllm-project/vllm/pull/57250

reply