Curious how this ages.
Recursive self improvement, self-play and multi-agent RL could make useful new theories, eventually.
However, at the moment I consider that they stay in the 'convex hull' of their training set + a provided context, and I don't see that much research that made real improvements to the situation.