upvote
>suspect that most the real gains are actually taking place in the harness.

Part of the reason harnesses work well is you can run a lot of agents in parallel. That doesn't slow down demand.

reply
That is true, but the eventual realization that more machines doing more coin flips in parallel does not mean "more work gets done" might.

LLMs are amazing tech, but they're terrible without oversight. More agents faster just makes reality collapse on them quicker.

But yeah, you're right, temporarily, this will still push demand. But the topic was about "diminishing returns" as in "tech getting better". Not as in "customer spending".

reply
It's kind of weird because more machines working together does mean more work gets done. Coin flips and weighted coin flips are totally different things. Any biases weights towards reality push you closer to reality when you use them.

New models keep being able to use more and more agents on longer time frames. Your hypothesis doesn't look like what we're measuring.

reply
Who is we?
reply
The people mapping AI capabilities.
reply
Oh cool, so that we includes me! :)
reply
Maybe turn on your light when you use a ruler? Not sure what else to say.
reply
I had actually been thinking more about all the non-LLM functionality that go into the harnesses. I'm not going to name names and I haven't done any rigorous testing, but my general impression is that choice of harness matters more than choice of model. In terms of basic task completion success specifically, not code aesthetics.
reply
A perfect harness will not extract gold from a dumb model. It's a system that builds on each other, though we've not probed that frontier much to have a good intuition on what effects what.
reply
One thing that I really want to know - the better models from today vs a year ago - what has changed. They have already pre-trained on all available public data. Scooping up the last percentage of archaic texts which were never digitized is not going to move the needle.

Is it just that the providers are generating tons of synthetic datasets on coding tasks so that the models get more exposure to the right thing to do? Every time someone points out an LLM stupidity they add some training data to patch over the weakness (trivial to generate "there are two 'l's in llama")?

reply