LLM's are trained on human knowledge and taste. They are actually pretty good at deciding if a conjecture would be found "interesting" by the mathematical community or not.
Note that I am saying LLM, and not chatbot or agent. But even a chatbot can often still reasonably rank a list of mathematical statements by vague properties like "interestingness".
How to RL this is a bit of an open question, but there are interesting conjectures of how to do it.
The fact that it's an open problem is the point I'm making. There's a very high degree of hubris right now, with people just assuming any open problem will be flattened by the AI steamroller soon. And sure, if that's what someone wants to believe that's up to them, but it's not a terribly interesting point of view to me. "What about X" "It'll be solved somehow", "What about Y" "It'll be solved somehow". Not exactly scintillating. If you know of any actual ideas on this I'd be interested to hear them.
Also any specifics on what you said about LLMs rating (preferably novel) mathematical claims for "interestingness" would be interesting.
There is a general idea that beauty in mathematics is about being maximally compressing. Say I have a book with all formally correct logical statements. I could prove everything by truth table, or I can maximally compress my book with all proofs of all statements, and that will make my math beautiful. Because it forces you to reduce everything to a core of very general statements which are powerful compared to the length of the proof.
Math as some kind of condensed crystal from the sea of all possible logic.
There are other ideas of how to do it. The time has come now to just try a bunch and see which ones produce good results.
Your statement that “I no longer have any specific task that I’m confident AI won’t be able to do” is founded on that.
This sentiment has been a recurring theme throughout the history of the field.