upvote
> Data from physics, chemistry, biology experiments

There are PetaBytes of important scientific data locked in archival file formats. The first step is to make this efficiently readable.

https://www.earthmover.io/blog/virtual-zarr

https://news.ycombinator.com/item?id=46659254

reply
> Frontier models can probably not advance much further with the datasets we have currently available

I’ve been reading this sentiment on HN since GPT4o, yet models got better and better

reply
There’s a difference between things getting incrementally better and a step function.

LLMs will obviously keep getting incrementally better, but in order to get the kind on jump that LLMs themselves were, the sentiment is that we need something more.

reply
Fable is a clear step function for me
reply
That’s because a significant portion of the work and spending going on at frontier labs is generating new, more curated data, in select domains like software engineering and now biology[1].

[1] https://www.synbiobeta.com/read/anthropic-is-hiring-biologis...

reply
I think that was his argument as well. It looks like there is no more data to find, but then it becomes important to find more data - lo and behold we can make more.

Sounds like the oil scare from 90' - we thought we were gonna run out of oil. But as oil gets more expensive it pays to dig further down to find the stuff that didn't make sense to dig up before.

reply
Yes. The web is enormous, but just think of home much information within a company doesn't even make it into the internal knowledgebase, let alone anywhere public.
reply