upvote
If you're actually applying LLMs, all of the things around the LLM that adapt it to coding, for example, that enable it to use existing validation tools for code, and enable it to diagnose and fix tool chain issues that aren't directly coding problems, are what makes the difference between a model that that scores a little higher on a coding benchmark and a model that's useful in a particular code base on a particular platform.

Are there any use cases that have enabled one customer of a frontier LLM to outperform a competitor using a different frontier LLM? Or is this why we are seeing confected points of comparison like solving challenge problems in mathematics?

reply
deleted
reply
I'm afraid it might be the other way around. RSI might pick all of the low hanging fruit soon. There must be a physical limit of how much intelligence you can squeeze out of some amount of parameters and compute.

There are going to still be worthwhile improvements but they are going to be more like not how to make transformers 10x cheaper but how to make next training run cost 9 trillions instead of 10 with a very particular optimization designed at the cost of hundreds of millions for this one specific run.

reply
I think we’re no where near a physical information theoretic limit.
reply
The hardware also is, so there ought to be a whole lot more room for improvement.
reply
It is just occurring to me that “RSI” expands to recursive self improvement. Thought people were talking about repetitive stress injuries; either in regards to programmers writing too much code/not having to write code anymore, or the frontier AI companies and their tendency to applaud themselves.
reply
[flagged]
reply