The disparity was so large that I was certain I must have made a mistake and I spent a few hours debugging, and then a few hours more trying different neural architectures.
Turns out this is just a super common experience for anyone in NNs who would also try the more established learning algorithms.