upvote
It might be sensitive to the prompt and tool definitions the harness provides, I can imagine it overusing the tool if being directly asked to resolve ambiguities. But in a straight Fable vs Sol comparison where I control the context and make sure I don't push the model, Fable consistently uses the ask user tool and Sol ignores it for me. Moreover, GPT has been known to do something like this since 4o if not earlier, it tries to ignore anything it perceives as orphaned context piece (e.g. XML sections with meaningful names but no explicit instructions on what to do with them). I suspect they specifically train it for that.
reply