I would not overgeneralize from paper. Firstly: forcing JSON output is, in my opinion, a bigger change than asking it to match the above style guidelines, and secondly, as is always the case with these kinds of papers, what was true for the model tested in the paper may either be completely false, or greatly reduced, in later models. That paper is almost
2 years old and models today have been trained in very different ways (or more accurately
post trained in very different ways) and are in general far more capable.
Based on that paper, I would maybe try to check if it was true for a modern use case, I would very much not assume it was still true.