A surprisingly robust approach to human verification would be to abandon visual classification entirely and instead exploit culturally embedded call-and-response patterns. There are obvious concerns regarding age and demographic bias, but these could potentially be addressed through a rotating corpus of challenges. More interesting is the adversarial question: while an LLM can trivially retrieve the expected response, can it distinguish between merely knowing the response and knowing when to come in?
When I say uh you say 'ah' Uh ah Uh Ah
Now when i say freeze y'all stop on a tidime
when I say freeze you just freeze one time
when i say freeze y'all stop on a dime
FREEZE
Now - all the ladies in the place, if you got real hair, real fingernails, if you got a job, you going to school and yall need nobody to help you handle yo business make some noise
NOOIISE
So - Dr. D. J. Kool has clearly been working on this thesis for some time