(Answer number one before that is usually "I don't have internet access, from memory it is either A or B, but I cannot recall what you want to know." ChatGPT or Gemini can often do the search, while google.com AI assistant or perplexity just tell blatant lies. Copilot.com can do the search, but external links are invalid made-up stuff for harder questions, which seems to be the case 9 out of 10 times.)
Which is great, since it could answer with made-up BS, but does not.
AI, except for doing better web searches for a year now, hasn't really improved for my tasks in the last three years, except for coding. Then again, AGI benchmarks seem to go through the roof only above Sonnet 5 and self-hosting, so perhaps the questions I ask not too hard for long now.
And self-hosting, eve 1bit/1.5bit models are a pondering a little too long to comfortable run in summer, but cheap on the RAM and insanely good at coding since a month now all of a sudden.
I have mentioned this here before, but the majority of my organization has reacted viscerally to this verbosity that LLM-text has been forbidden: in comments, in PR/commit messages, in correspondence, in Jira tickets.
A couple non-coders who want to make PRs without writing the description are now rebelling and saying this can be fixed if we spend our time writing skills for Claude so it becomes readable again.
It's important to remember that we are talking about a calculator that doesn't have an understanding of common sense. Unironically, this is common sense.
“The question of whether a computer can think is no more interesting than the question of whether a submarine can swim.”
Whatever these things are doing, it’s not the same as what a person does. Trying to decide if whatever they do fits into the box we label as “intelligence” is completely uninteresting, in my view. What’s interesting is figuring out just what they can do and how best to use them, which sounds like a related question but really isn’t.