If you just talk to it over API (no web search) the Gemini models are extremely resistant to thinking the user may be living in a universe outside their training data. Try to discuss any news etc and they assume it’s fake or fiction
Out of the three main US AI companies' models, Gemini is obviously the less aligned (read: censored). So I really don't know what you're talking about.
I'm saying AI researchers have a bias towards thinking what needs to happen is prompt -> [crunching tokens] -> response rather than prompt -> [orchestrates 5 tools] -> response
In other words 'just add a calculator tool' is not as sexy research-wise as making the model accurately eyeball arithmetic in its chain of thought. Maybe I'm wrong but that seems to be the case