undefined

points

[-]

But is arc-agi really that useful though? Nowadays it seems to me that it's just another benchmark that needs to be specifically trained for. Maybe the Chinese models just didn't focus on it as much.

by sdenton42 hours ago|

parent|

[-]

Doing great on public datasets and underperforming on private benchmarks is not a good look.

by Deegy2 hours ago|

parent|

[-]

Is it though? Do we still have the expectation that LLMs will eventually be able to solve problems they haven't seen before? Or do we just want the most accurate auto complete at the cheapest price at this point?