* pre-chatGPT is not an effective control because language evolves. In the arxiv corpus in particular there are "fashions" in research depending on what gets funded lately, not to mention many new words and topics not invented before a given paper.
* In general, detecting AI from content seems difficult as humans write like they read. To the extent there are unique factors recognizable as AI and to the extent humans read them, they will eventually incorporate them into their writing style. Accordingly, you'd need to model a rolling window of "AI tells" that decay at some rate.
FYI this is all relatively new so there might be lots of issues and iterations coming.
I had to change my mind on AI detectors after playing around with it.
It would be interesting to hear how this detector compares. It also seems to be aiming for low fp rate.