It had previously attempted to create that table as part of the test setup, so it apparently concluded that it was a test table.
During human review, it explained that it had simply chosen a table name inspired by the codebase.
Many other models get things wrong, but Gemini is the only one to go on the defensive.
https://www.fastcompany.com/91383271/googles-chatbot-apologi...
https://www.businessinsider.com/gemini-self-loathing-i-am-a-...
In my opinion still the most egregious example in history of a commercial LLM going off the rails in production. Never any technical postmortem from Google on this.
Point is, if Gemini is flawed then there's a very good chance that it's still deeply flawed today, and getting smarter at the same time - that is a very bad combination.