upvote
Yeah, the model is small enough that inference is already basically instant for my usecase (only 6 transformer layers for the blog search).
reply