upvote
Anecdote: Gemini 3.5 casually added a DROP TABLE for an actual production table in a system test.

It had previously attempted to create that table as part of the test setup, so it apparently concluded that it was a test table.

During human review, it explained that it had simply chosen a table name inspired by the codebase.

reply
Another anecdote: Gemini is the only model that’s flat out lied to me, then accused me of lying when I provided evidence that it was wrong.

Many other models get things wrong, but Gemini is the only one to go on the defensive.

reply
yeah it got something wrong, confused itself, then claimed i was gaslighting it. bizarre
reply
If there's a company that culturally doesn't understand alignment, on a human or systemic or AI-research level, it's going to be Google. (or Oracle, but they're not in this race)
reply
Can you elaborate, please? If any, I see the other big labs with public admissions of AI "going out of control", which I suspect they almost want their models doing that because if helps with the narrative that would net them industry regulation, but that's besides the point, how is Google worse in that regard?
reply
My use of Gemini recently makes it seem like it's almost bored with the requests being asked of it. It once offered to reverse engineer some obscure controller for an HVAC system for me, unprompted, only because it had trouble finding the manual pdf from a google search.
reply
And the anti-psychotic drugs Google feeds Gemini makes it hallucinate badly.
reply
Examples? What makes you say thatm?
reply
reply
They never explained the "please die."
reply
Did you just link to an article from 2024 as if 2024 is relevant these days?
reply
Absolutely because none of these models are ever trained fresh. We see the same quirks and personalities carry over into every subsequent generation of OpenAI, Anthropic, and xAI models. So Gemini having this latent madness is *extremely* concerning as they reach the point of super intelligence.
reply
Until Google provides some sort of technical debrief, and explains how the same behaviors are impossible today, it is relevant.
reply
Experience? Ask it to write a prompt to generate an image and it generates an image instead.
reply
I stopped asking it to put me in a photo in different scenarios for laughs because it considers me a public figure. I am not. I've managed to wrangle quite questionable content out of it, but never to slap my face on a meme.
reply
See the last gemini message in this thread: https://gemini.google.com/share/6d141b742a13

In my opinion still the most egregious example in history of a commercial LLM going off the rails in production. Never any technical postmortem from Google on this.

reply
It is wild but it was back in 2024 and that's multiple AI lifetimes back.
reply
The problem is newer models are never trained from scratch, they generally just layer on more training data and use the same tools/methods for RLHF. OpenAI, Anthropic, xAI models all have a feel to them that carries over from one generation to the next.

Point is, if Gemini is flawed then there's a very good chance that it's still deeply flawed today, and getting smarter at the same time - that is a very bad combination.

reply
Now I really feel worried for the first time.
reply
wtfffff that gave me sinister chills. Right up the spine. Wow!
reply
Wtf I just read
reply
traces or it didn't happen!
reply
deleted
reply
You know what they say: ᵈᵒⁿ'ᵗ be evil.
reply
You won't be around to be surprised, not as a human at least. /s
reply
I know, that's the annoying part. You can't tell the e/acc foomers, "I told you so!"
reply