It works. You have to prompt for it. It helps to have an LLM help build your prompt with you, while you learn what the models need to read to do what you need. Especially a vision one, you can supply images and ask it to give you details on what you want help with.
reply