In theory, modern vision language models could classify human nudity and sexual activity very thoroughly. But every model I have tried is reluctant to clearly describe what is notable about sexualized/nude images. The models are deliberately under-exposed to nude and sexualized content during training and further RLHF'd away from generating straightforward descriptions of such images.
Models also occasionally hallucinate WTF captions for ordinary adult sexual activity. I recently ran a baseline test with frames extracted from adult videos and about 1/3000 frames was mis-captioned as involving a child according to Gemma 4 12b.