upvote
Sorry but your argument doesn't seem coherent: How is the cost of RL relevant here?

It would also help if you could substantiate your initial claim (i.e. "internet training data is not where frontier capabilities come from")

reply
RL environment (instruction, stateful container, reward function) is the training data product being bought
reply