I see plenty of anecdotal evidence that models have been trained fantastically well—and getting better—at writing to trigger the right neurons in the human population to produce “This is interesting/informative/correct” responses in bulk.
Could their ability to produce those responses run far ahead of their ability to actually achieve the last in reality? Sure seems plausible, and then where are we?
My anecdotal evidence is the LLM-generated, inchoate technical dross that is routinely upvoted onto the hn front page. Much of it isn’t even coherent enough to be wrong, but the readership here finds it interesting!
That's a big "If".
If a research is good, the author still has to clear all the hurdles in publishing. "Writing your own paper" is just one more hurdle.
> A possible conclusion for this could be: If the majority of CS papers is AI written, let's just accept this reality universally and stop worrying about it altogether.
That's just a different way of saying "if the majority of CS papers are crap, lets just accept this reality".
So, go on, publish away all your AI-induced "research", but the bar is slowly going to be raised anyway to reject that. That's how science always worked - when a bar is not sufficient to exclude the crap, it is raised.
Reviewers were also totally unengaged. Of 20 reviews I read (from my reviewers or from reviewers on the same papers), maybe 2 were mediocre, and the rest were crap (though likely not AI).
The notion that science will somehow benefit from this is about as stupid an idea as you can have. Science relies on skepticism. AIs are not skeptical, and many folks are submitting papers because they stand to gain something, not because they are motivated to do good research or develop new understanding. Fields are being inundated with garbage that is maximally indistinguishable from real work (that's the training objective for LLMs). This in turn maximizes the cost of identifying bad work.
This is the same enshittification process that we see everywhere else. You get spam phone calls because there is no reason for a spammer not to call you. "Researchers" are submitting spam papers because there is no cost to doing so with some possible gain. Absent intervention, this eventually drives the community value of the network to zero (or potentially negative, if friction costs to switching are high).
Other problems include: Signal to Noise Ratio going through the roof.
Yes, I feel like there's room to improve things, I just strongly doubt that "using AI to detect AI" is a particularly useful thing to do here.
That’s a big if. We all know that’s not what’s happening.
That's a big if. ArXiv is not peer reviewed and LLMs basically interpolate and extrapolate text, which makes them essentially fluff generators. Even in the most charitable interpretation, LLMs enable those with nothing to say to say nothing while meeting surface-level style guides.
Why don't you read them and see? The ones I looked at were clear slop.
That's a huge assumption, and one that goes against the whole notion of using LLMs to generate text. AI slop is by far the norm.