Though, I suppose RL has non-gradient based methods too.
As a mixture, could activation-space search produce useful teaching targets for backprop?
Zeroth-order search would discover candidates, first-order learning would consolidate them. The potentially valuable step is converting a sparse judgment into a reusable training target.
This also changes the relevance of convexity.