upvote
The reason it makes sense is because different models are better at different things. Allowing yourself to be locked in to a particular model by these companies is a bad idea and will lead to worse outcomes for yourself and the market as a whole.
reply
Is there a good place for harness comparison/scoring?
reply
https://www.harness-bench.ai/leaderboard.html

This provides a method, but the data looks stale and perhaps a bit thin compared to say, Cursor, or even AntiGravity data.

reply
never heard of nanobot. how reliable is this benchmark?
reply
I can’t speak to the veracity of the benchmarks but it appears their methodology is sound. Nanobot has 47k stars, fwiw https://github.com/HKUDS/nanobot

It has been more of an OpenClaw or Hermes alternative than a coding agent like OpenCode or Pi, so it’s likely to do well given less context bloat.

reply
IMO it's not. It's benchmarking GPT 5.4 and Opus 4.6. It's also missing Claude Code... one of the most popular harnesses (the most?)
reply