upvote
For tasks with no existing ground-truth data, one could try to use an adversarial objective on its own by making AI-generated outputs compete against each other, although this is beyond the scope of our experiments. Adapting our method to this setting is quite straightforward.

For now, we have looked into composite objectives trading off performance on the ground truth against the ability to reject generated samples produced in earlier epochs. For example, when using this system to co-evolve paper writers and reviewers, we reuse generated papers that were accepted by the reviewer of epoch t as adversarial samples in epoch t+1. Then, new reviewers are rewarded for rejecting AI-generated papers. You can have ground-truth performance account for x of the utility of a reviewer, with the adversarial objective then providing 1-x.

reply
deleted
reply