upvote
> a database that predicts the effects of every possible single nucleotide variant in the human genome. We used the AlphaGenome AI model to pre-calculate the regulatory impact of all 9 billion single-letter genetic changes, resulting in a massive, 1-petabyte dataset.

This is for a database, no? While Borzoi is a model?

> Here, we introduce Borzoi, a model that learns to predict cell-type-specific and tissue-specific RNA-seq coverage from DNA sequence.

https://www.nature.com/articles/s41588-024-02053-6

reply
Any model can be expressed as a database.
reply
Kolmogorov looks at this with a ‘duh’ face
reply
> Meanwhile, everyone in the field of genomics knows that AlphaGenome provides essentially zero improvements over the previous SOTA, Borzoi…

Can you elaborate on this? I'm confused why Google would build something that provides zero improvements over SOTA, Borzoi... as you mention. I'm not familiar with this field, just curious.

reply
Speaking as somebody who has worked within Google Research before: the researchers are under tremendous pressure to publish SOTA and sometimes they juice their results a bit to look competitive when they can't match. This is not uncommon in the field- it's remarkably easy to edit a paper to make yourself look good by omitting information.
reply
One of the most egregious cases of this, in my opinion, is only publishing metrics that cover part of the confusion matrix. “The false-negative rate? That could not possibly matter for a variant effect prediction model; why would we include that in the paper?” Example: AlphaMissense.
reply
The use of an exact quote in an ungrammatical fashion is a bit of a language model smell. I can’t help but be reminded of the purely nonsensical AI interview answers. “It’s a pleasure to meet you, Chick Bongo”
reply
Creating an account 12 minutes ago (from the time of this posting) to comment on how another comment seems to you like a "language model smell" is itself, a language model smell, or a scammer.

Please stop accusing or hinting at others being a language model or bot. Not only is it a dumb waste of time, it's wrong in this instance and you are not only going to continue to be wrong but you have no way to prove or demonstrate that any single post comes from a bot nor the ability to do anything about it if you did in fact believe some comment to be attributed to a bot.

reply
They did benchmark the model and beat the SOTA on every metric though
reply
Is it so bad to have another entrant, especially with the resources Google could bring to bear?

Imagine if the Apple EV had actually happened, you think the EV enthusiasts would roll their eyes like you are?

reply
Are you really taking a holier than thou approach on a google article?
reply