upvote
Sure, but the question stands. The ability to evaluate some measure of success seems pretty fundamental to how we train models and iterate with them on tasks like theorem proving. What is that measure for mathematical conjecture generation? How do we evaluate success, either on a particular task for iteration (like we do by eg. setting an agent to produce a lean proof of a specific result) or on a large enough set of training data to learn a set of rules (like we do when eg. using an RL environment to train a model to generate source code that passes automated validity/correctness checks).
reply