upvote
Show HN: JevBench, a reproducible benchmark for typed decision models

(benchmarkheaven.com)

Good project but this one also exists https://huggingface.co/spaces/multimodalart/jev-decision-ind... and the results do not seem to add up and also model sets are different... still needs time to mature likely
reply
https://is-it-ai-slop.app.mintapis.com/ is a fun tool. Is the source or methodology for that in the github repo? I couldn't find it immediately.

We've been experimenting with Jev for classifying email, some thoughts here: https://housecat.com/blog/classifying-email

Flagging AI written email is a much requested feature too.

reply
reply
Interesting; was curious how this didn't fall into trouble with ToS. Apparently the "no benchmarks" clause was intended for "limited preview" audiences and didn't get removed at launch on accident.
reply
deleted
reply