upvote
The thing with watching CI in an agent loop is that it burns tons of tokens. At work I ended up writing a deterministic, traditional CLI tool to poll GitLab CI pipeline+job state changes on a branch and exit with an appropriate status code, and then updated my `/glab-ci-feedback` skill to use that. Saved a ton of token churn, and now I have a runbook a human could just as easily use if they don’t want to (or can’t) use an agent loop.

… but walking away to make a coffee and coming back to the robots auto-fixing bugs only found in CI is definitely some flavor of magic, regardless of the execution order to get there.

reply
FWIW, Claude Channels[1][2] are probably going to be the solution for that, eventually. While I'm not sure how the WebHook receiver example will work with, say, GitHub and a local Claude, the Chat side of things _would_. So you'd have GH send its web hook to Telegram (for example), and then the Telegram Channel MCP would inject that into Claude, and Claude would start working on the problem. Still experimental, but functional enough to play with.

[1]: https://code.claude.com/docs/en/channels [2]: https://code.claude.com/docs/en/channels-reference

reply
I think they are trying now to to bake CI awareness into Claude Desktop, didn't use it yet.

But meanwhile we also have the scripts - one script to watch CI, one script to fetch comments (without dumping raw graphql into the agent), etc etc. Can't wait for this phase to end already

reply
Yeah the codex app can deterministically poll and watch for you too. Consider it like an event based trigger, where the event can be anything you can dream of (like webhooks!)
reply
I think that's what a future dev team is going to look like.

One person doing product management / talking to customers and vibe coding features that solve users' problems, one person keeping the UI/UX in check, one QA person that spends their time clicking through the software, finds the bugs that are obvious to humans but not LLMs and fixes them, and one "harness engineer" who pays off technical debt, observes failure modes and sets the rest of the team up for success.

reply
> Another thing to think about is, what would it take for you to care less about the understanding.

Could you explain why it would be a goal to understand the system less, rather than more?

It seems harder to know if you have good tests while lowering your expertise in the system.

reply
Because humans are currently the bottleneck.

An LLM can produce far more code than a human can understand. And the famous rule that "optimizations are entirely pointless unless you're optimizing at the constraint" is logistics 101.

To accelerate software development, you either need to remove or lessen the need for code understanding, or make it much quicker for humans to gain that understanding. Making the LLM faster won't help you if the LLM isn't the bottleneck.

reply
A human can produce far more code than a human can understand, too, but pre-LLM we always viewed someone overwhelming their colleagues like that as being bad at their job.
reply
There's different layers of understanding the system. I generally care about high level data flow, concurrency and performance (batching, holding transactions too long, back pressure etc.) rather than the mechanics of how the code actually does a thing. I still look to see what the final output looks like and ask my agent questions on how it fits in the larger system and evolve things if necessary, but agents are pretty good at writing code if the rest of the code base looks pretty decent.
reply
deleted
reply
Opus 5.5 on Low seems smarter, cheaper, and faster than sonnet on medium, so what's the point of sonnet?
reply
Being that my first prompt can be something like: for task x/issue y, which model would strike the best balance between cost and capability…

It seems like it would be a better UX to have model and effort selection asked into the system. Of course, I’m not sure in practice if that would be in the best interests of the providers and/or users.

reply
Claude Code has the issue that sub agents inherit the thinking level. This means that to use a smarter or dumber sub agent you need a different model. That's not a particularly good reason, but that's my one use case for Sonnet.
reply
You can also create custom agents with defined effort levels and use those.
reply
Or just use a better harness.
reply
> Run adversarial review.

Be careful about this one if you want to have any level of control over basic stuff like comment style and accuracy. Claude will happily spend 20 review cycles in a row rewriting the same 10 comments for a small bugfix over and over because it can recognize "Claude-ese" in the review cycle but then just immediately and compulsively spew out more of it and drift even further from your style rules in the next "fix".

I'm seriously not joking about the 20 tries, I left it running in the background for what should have been a minor code change and it took 18 out of 20 review cycles to stop writing in more comments that all either broke my ASE-STD100ish style rules or included false statements about the code.

reply
lol yeah, our review bot does a cost based analysis and pauses itself until you re-resume if it goes over a threshold.
reply