Mercor, one of the larger vendors for contracting with experts to create bespoke data, says on their webpage they're paying $3M/day to their contractors for data.
So well into the billions of dollars a year for bespoke training data.
That's also ignoring the RLVR data labs can get from software - they can use the vibe coding sessions as training data as well without paying more.
They are just one of many.
A human asks a question, then writes rubrics to judge the LLMs response, so rather than evaluating a specific response, those rubrics can live on as the LLM evolves and gives different answers. There are more complex variants as well, but that's the basic principle.