the underlying data set needs to be representative
and then you are going to ignore all the research and results that clearly show otherwise? why?
what might we infer about the importance of data from a learning algorithm like decision trees?
Existing datasets, different reward function.