I think such progress by agents is not a sign of broad generalization but of broad coverage. We have exposure to a subset of deeper scientific subfields and thus can only generate certain attacks to solve a particular problem. Since it is not clear which combination will lead to a solution beforehand it is nontrivial to look at a problem and fill our knowledge gaps. LLMs on the other hand have broad coverage and can generate hypothesis on a wide combination of subfields. With Lean an agentic loop can test these to sift the weak ones. In a way the problems solvable with this setup is also solvable by a human who happens to know the right subfields. These problems are likely to require an esoteric combination so nobody could solve them before. I really am not sure whether all generalization is like this or we can leap and create novelties beyond what an llm can generate. That I guess is the tough question that we need to answer to understand the boundaries of intelligence.
reply