Do you have a particular paper in mind that does likelihood-free best-of-K but just calls it GRPO?
Likelihood is not fundamental to the spirit of GRPO, any exploratory mechanism would work.
That sequential LLMs have a step-wise probability is convenient but not critical to this approach (where rejection sampling is widely used in diffusion models).