They're literally comparing the previous version of the same model with the new one. It's based on the same architecture, same pre-trained model, just different post-training. It doesn't get more apples to apples than this.
The performance changes are so big with the right harness that is makes sense to engineer the harness and fine-tune the model to one another from the start.