Gemini wrongly called an authentic product a fake, but that was mainly because there really were typographical errors in the authentic product's ingredient list. That's more of an indictment of the manufacturer than it is Gemini.
Finally, a lot of the things that Gemini got wrong seemed to me like they could reasonably be attributed to things like optical artifacts in the photos (glare, shadows, stuff like that). Better/more photos might improve that.
All that being said, this is obviously a tiny study of a single AI tool with a single cosmetic product, so I should probably withhold my "cautious optimism" until we have more data.
Very odd conclusion to an otherwise interesting experiment.