upvote
The harness instructs them to behave this way. Also this approach saves tokens. The scripts allow to edit files in bulk, and most of the session cost is in cache reads (e.g. for 300K context each command costs the same as 30K input tokens).
reply
> The harness instructs them to behave this way. Also this approach saves tokens. The scripts allow to edit files in bulk, and most of the session cost is in cache reads (e.g. for 300K context each command costs the same as 30K input tokens).

I understand the reasoning, but at that point wouldn't the LLM be better off creating `sed` commands and executing those? I mean, if it's already executing Python, it can literally do anything to the environment, so using `sed` is at least as safe, with a bonus that it (or a subagent, or a human) can double-check the intention with the sed script and flag incorrect or missing changes.

reply
I've experimented quite a bit with giving agents python vs sed + awk. They make mistakes with both, a lot. The only thing that has stood out is that agents reach for python too quickly if it's available, and that awk causes the least problems, while sed might take several attempts to get results, similar to python.
reply
Also it's the only way that makes sense when you need to work with big files, or large amount of files, or documents that look small when fetched through a RAG tool, but then you read one and get hit with couple megabytes of base64-encoded binary data you didn't expect because RAG tool stripped out embedded images...

Ask me how I know. Or don't. I have a standing rule for all agents warning about that failure mode (and related, doing `ls` in `/tmp` and few other directories that like to accumulate files by the hundreds..)

reply
My harness forbids it, they end up spending time debugging their scripts
reply
Why would you use a constrained edit tool when you are also allowed to use the complete power of python?
reply
Because the complete power of Python also includes the power to fuck things up.
reply
So does using an LLM.
reply
In fact, that's kind of the whole point of using LLMs in the first place. Their value is in their general capabilities.
reply
…which we attempt to constrain by encouraging the use of tools that make it harder to fuck shit up.
reply
Having an agent edit 100 files means the job will definitely get done correctly. When it writes a script to bulk edit things it fucks up and spends ages debugging their script.
reply
Have you ever counted the number of times Claude fucked up quoting/escaping and had to issue a corrected tool call? Or get stuck in some tricky quoting situation for two minutes, throwing a couple piles of shit at the wall to see what sticks. IIRC I’ve even seen it eventually using the edit tool out of frustration once.
reply
Simple is better than complex Complex is better than complicated

Or something, I don't remember...

reply
... simply the best, better than all the rest (Tina Turner)
reply
Ușor, Burebista :-)
reply
That was pre LLM. Now everything is a prompt that you type into AI lol
reply
Why even offer the edit tool in that case? Also, what kind of editing could they possible do what wouldn't be possible with POSIX ed?
reply
You can chain a lot more commands together with this technique than with a single Edit tool call.
reply
The funny thing is that... POSIX ed is composable :-)

You can do a gazillion edits with it in one shot.

Of course, LLM edit tools are probably small bits of their custom code, I just find it funny. I wonder if it's a desire for certain technical characteristics that require custom code or just a lack of info on basic tools. Heck, if it's about platform availability, using an LLM to port ed to Windows (for example) should be trivial[1].

* * *

[1] And there are probably a million existing ports. Also, sed, ex, vi, whatever.

reply
This is an instruction by the harness. It re-injects the prompt every other message, so that's why it "forgets" to use the Edit tool.
reply
this is intentional, afaik agents do better with python and alike than the harness tooling.
reply
Using python or any other stone-age approach for search and replace is stupid when your language provides you with a complete, fully typed AST, like .NET does.
reply
I use AST replacers, much more reliable.
reply
Sounds like you don't have enough experience with coding agents. Deterministic scripts must always be preferred instead of LLM tool calls. In fact, you should instruct your agents to write code to execute instead of letting them call tools.
reply
Plenty of experience ;)
reply
Then why did you comment what you commented, good sir/madam. Claude and Codex are good at remembering to use scripts instead of tools these days, especially if your <32kb .md file mentions it. Not even talking about the skills designed to catch such issues.
reply
Sounds like you completely lack all reading comprehension ability

LLMs sometimes like to execute one-off Python scripts to make edits to files rather than just calling the edit tool directly. Both are tool calls so saying that you should have it write code instead of doing tool calls makes no sense because writing code is a tool call for it...

reply
[flagged]
reply
Are you saying scripts from agents are deterministic? :)
reply
Why don't you try to dispove me. Yes, they are _more_ deterministic than tool calls and consume less tokens.
reply
There is nothing to disprove as you don't understand what does a word mean. Deterministic is not a spectrum, they can either be deterministic or not. In both cases, they are not.
reply
Oh, I'm so sorry I touched your paper feelings.

How dared I to imply that some LLM output is more deterministic than the other, your LLM majesty. Shame on me and my entire family! For generations to come!

So sorry I implied that the code that doesn't work and has to be fixed later is deterministic in its execution and can be reused later instead of being re-generated from scratch!

Will I ever wash it off my name, your grace?

reply
They're talking about writing a file with a harness-native Edit tool. They're saying the agents aren't doing that, but are using ad-hoc methods of writing the files. (My agents seem to prefer see these days.)
reply
Why do you think your agents prefer to create scripts instead of doing tool calls these days?

I wonder, is it easier to modify a script that agent wrote before to satisfy your prompt, or is it easier to write a new one from scratch each time a retry happens?

Are input tokens more expensive than output tokens?

reply
My god. They are not reusing the scripts. They are adhoc, inline Python scripts just used to make a single edit. You seem to fundamentally not understand what everyone else is talking about
reply
They do reuse scripts, though. Maybe it's you who is too lazy to _comprehend_ the output?

Or, maybe your prompts are not good enough. And it's not my problem to fix, as you claim you are very experienced.

reply
deleted
reply