upvote
I'm not an academic, and I understand by osmosis that organized science has many problems related to publishing and connected mechanisms, but I have to say there's been quite a few moments in my life and career when I found the inspiration or solution I needed in a good paper.

For condensed knowledge on highly specific topics (i.e. more specific than is economical in book form) I often know nothing better.

Most recently I did sort of a self-taught crash course in ground-penetrating radar and applications for an archeological endeavour I write software for, and a number of papers have been really invaluable.

reply
I think one of the end-game components for AI is to be able to read all the literature and generate a reliable list of which ones are not correct and a convincing reason why.
reply
> generate a reliable list of which ones are not correct and a convincing reason why

That is not realistic, but I suppose where things are heading is that you have some indicator of the strength of evidence -- see Fig. 13 in the following insightful take:

https://news.ycombinator.com/item?id=49407226

Though, "strength" should probably be "reliability" and "validity", and I suppose those indicators are more for picking signals from the noise; i.e., what is even worth clicking and reading. That would be increasingly valuable already today due to the volume (and, yes, slop and other related stuff).

reply
> That is not realistic, but I suppose where things are heading ...

And to expand on this, it's not realistic because science is not armchair philosophy. You have to go out and measure the world.

Sometimes, through force of will, a person can think deeply about a problem and come up with beautiful theories that explain our measurements. Many scientists had careers like this, probably most famously, Einstein. But it's worth noting that Einstein also got a lot wrong! [1]

Even if we somehow give an LLM the ability to go out and measure things, I seriously doubt that the role of humans in science is done. There's a big difference between "an explanation" and "a good explanation." Ask any physicist. There's a surprising amount of aesthetics involved. Good theories are consistent with the evidence, but it's more than that-- there's a great deal of "taste" involved. And there's a good reason for that. For any real problem, there are effectively an infinite number of alternative hypotheses. From a "theory of science" standpoint, this should cause scientists nightmares, but it doesn't. Because by the time you are a practicing scientist, you've developed a feel for what constitutes a satisfying explanation. If you spend time with scientists, especially in the "hallway track" at a conference, "taste" is a frequent topic of conversation!

[1] https://en.wikipedia.org/wiki/Einstein%27s_unsuccessful_inve...

reply
I think you misunderstood. I'm not asking for an oracle that can determine whether a paper is correct, I want an oracle that can find real mistakes in papers (thus invalidating them).

I work full time on "lab in the loop" AI, so I'm pretty familiar with the need for real-world experiments. I am not proposing a fully autonomous scientist that could read an arbitrary paper and emit whether it's universally true without some verification method.

Also, to your statement: " Because by the time you are a practicing scientist, you've developed a feel for what constitutes a satisfying explanation."

I'm a practicing scientist (well, ex-scientist) and it seems like most "satisfying explanations" end up being wrong or incomplete simply because they seem so satisfying.

reply
> I'm not asking for an oracle that can determine whether a paper is correct, I want an oracle that can find real mistakes in papers (thus invalidating them).

What's the difference? How do you find real mistakes without a model? Either you have a trusted mathematical model (in which case you already have a complete explanation) or you have to compare it against the ultimate oracle: the world. Or are you proposing something like "let's use an LLM to convert this hand-wavy English paper into a formal proof and then check it for logical fallacies?" In which case, fine, that would be useful, but that's not exactly the same thing (and also not as important) as saying that a paper advances a bad explanation. Just that the explanation is flawed in some way.

reply
How do you determine that a mistake is serious enough that it actually invalidates the paper? How do you know that it's not possible to correct it and reach a similar conclusion with fundamentally the same approach? Especially when the fix would be complex enough to justify a follow-up paper.
reply
This is realistic. We already do this today: it's called "journal club". A bunch of grad students read the same paper and then criticize it. I've read papers that I thought were amazing only to. have somebody else notice a key issue in a method, or a conclusion that didn't follow, or outright omission of an important detail, in a way that could be verified by both the students and the authors of the paper.
reply
You can be served this inspiration from LLMs
reply
I've tried here and there, but even the top-end models always tend toward shallow generalist takes. I've not been able to get more than basic primers out of LLMs, and nothing close to the ability of a professional author to stay on course with an idea, detail level, specificity etc.

People in the early days used to often whine that LLMs just regurgitate text snippets (unfounded of course), but I think the way we currently train and RLHF them actually seems to largely make them unable to reproduce the knowledge they have been trained on, since they seem to just always want to please the mean with their output. I'm oversimplifying the mechanisms, but you get my drift.

reply
Alan Sokal style hoax is what comes to mind first, of course. Wouldn't be the first computer-generated hoax either (see SCIgen). However see also the P.P.P.P.S. in the article, his co-author denies that in the comments.
reply