upvote
They are already solving the problem with search engines, they're just using an LLM as a first pass to create better embeddings to run a similarity match on first. The difference in latency is likely made up for in accuracy.
reply
A 2s LLM call is pretty slow.
reply
Try using Digital Ocean. Minutes spent on inference.
reply
[dead]
reply