* speed - much fewer hops back to the LLM
* fewer tokens - intermediate execution steps in the script don't leak into context, only the final result does.
* repeatability - if the LLM needs to repeat work, it can reuse a script it wrote last time.
If you have a harness that has access to a full shell and knows how to use bash or python, you'll often see it writing little scripts. For setups that don't (ie normal model API requests with tool calls), you can give it an lightweight secure execution environment like just-bash, or quickjs.
What if I have an MCP Tool LookupZip(City) and want to chain it with a bash tool that prodcues a list of 100 cities. And then I want to filter again to the largest Zip code.