upvote
The main one is mapping images to text and visa versa, e.g. semantic search of images via text, where the images are encoded and the text question is encoded with the same model, then finding nearest neighbors.
reply
I haven’t tried it yet, but, for law practice, I could imagine using it to search a case file for “undamaged roof before Hurricane Katrina” and “damaged roof after Hurricane Katrina” and being able to locate both deposition testimony and pertinent photographs in the body of evidence.
reply