upvote
I can’t see your gist but spam classification is a textbook example of something you shouldn’t measure with accuracy. If 95% of your samples are not spam you can get 95% accuracy by always guessing not spam.

You should use precision (when your model says “spam” how often is it spam?), recall (how many of the spam emails did it catch), or f1 (balanced between those two).

reply
That's a great point. My case is not for spam, the classes are more balanced, but you are correct that precision, recall and f1 would be better measures for some of these tasks
reply
I wonder what numbers you'd get using another system one model - Contrastive Language Model https://contrastive-lm.notion.site/

That model scales very well with quantities of requests.

reply