upvote
For email you can use a classifier

One way: separately embed sender, recipients, subject, body - then use the embedding vectors as input to a logistic classifier

With that setup, I get 95% accuracy on email classification, training on 50-100 base examples. The model trains on CPU in under 1min, and it does inference in under 20ms (most of it is running the embeddings, so you can make it faster if you train your own embeddings model)

Here’s a gist with some sample code: https://gist.github.com/nicobrenner/056a5aaff5d0119c0032ecda...

That code applies the embeddings + classifier setup on the Banking77 dataset. It gets 93-94% accuracy depending on the embeddings you use (SOTA for this is ~95%, with much bigger and slower models)

reply
I can’t see your gist but spam classification is a textbook example of something you shouldn’t measure with accuracy. If 95% of your samples are not spam you can get 95% accuracy by always guessing not spam.

You should use precision (when your model says “spam” how often is it spam?), recall (how many of the spam emails did it catch), or f1 (balanced between those two).

reply
That's a great point. My case is not for spam, the classes are more balanced, but you are correct that precision, recall and f1 would be better measures for some of these tasks
reply
I wonder what numbers you'd get using another system one model - Contrastive Language Model https://contrastive-lm.notion.site/

That model scales very well with quantities of requests.

reply
> hold your horses to paint it as dirt cheap

For a moment I thought this was going to be a metaphor — maybe an ancient Chinese proverb about how paint brushes are made from horsehair and how you can't hold the horse to paint before you've turned the hair into a brush.

reply
Darmok and Jalad, at Tanagra
reply
How are you benefiting from prompt caching for simple classification?
reply
There are two parts in the data you supply to Jev for classification - the prompt describing your classification and the data. The data can be quite small - a simple chat message. And prompt part could be considerable since you need to describe your rubrics well.

With Jev you each time pay for your prompt, you can't cache it.

reply
I mean, it sounds like it's only ideal for cases with significant system prompt overhead. I don't think Jev was built to have a large well described prompt setup. To me its more like a happy go lucky small label classification tool with important decisions left to stronger agentic models or yk humans.
reply
Are you getting better performance from an LLM than a Bayesian classifier?
reply
How many requests per second do you have for spam that you are reliably hitting the Luna cache?
reply
Is that just because the Jev implementation is less mature? Couldn't it also implement prompt caching?
reply
What are the costs compared to an ML model?
reply
[dead]
reply