Codemode is a way for the LLM to orchestrate harness level tools. The reason this happening now, is because the models by the labs are increasingly trained on this. Codex for instance in responses lite requires codemode to even perform parallel tool calling.
It's pretty effective because of the reasons you noted, but there's a composability problem since each MCP has its own sandbox and can't call into the other ones.
IIUC Pi offer a workaround for this, the harness runs the sandbox and populate it with the MCP tools, that way the composability problem is solved and every MCP do not have to implement their own sandbox.
codemode lets you execute scripts in a runtime where your MCP tools are made available as function calls
this matters for cases where the MCP tool is the only way to do something and you do not have an equivalent CLI, API, whatever to script with
* speed - much fewer hops back to the LLM
* fewer tokens - intermediate execution steps in the script don't leak into context, only the final result does.
* repeatability - if the LLM needs to repeat work, it can reuse a script it wrote last time.
If you have a harness that has access to a full shell and knows how to use bash or python, you'll often see it writing little scripts. For setups that don't (ie normal model API requests with tool calls), you can give it an lightweight secure execution environment like just-bash, or quickjs.
What if I have an MCP Tool LookupZip(City) and want to chain it with a bash tool that prodcues a list of 100 cities. And then I want to filter again to the largest Zip code.