upvote
We already have a 1gb model that is as capable as it will ever be, there's a proven ceiling that cannot be passed. For example: you can't make a mice-sized brain as smart as a human brain no matter how hard you try.
reply
Please source this claim. What 1GB models are capable of has increased generation-on-generation.

> For example: you can't make a mice-sized brain as smart as a human brain no matter how hard you try.

Sure. We don't know where the ceiling is for our digital minds, though.

reply
They have not increased in capabilities, they have increased in specialization.

If you train a small model in another domain it will begin losing capabilities in the former domain. This is effectively the sigmoid problem.

Although I will admit that if we discover a higher information density algorithm that it might change, but not by a substantial amount to where "super intelligence" in 1gb would be possible.

reply
Over the last two years, this weight class has doubled its scores and/or saturated several benchmarks in the Qwen lineup alone without loss of generality: https://claude.ai/public/artifacts/9f249169-3623-417e-86cd-7...

There is undoubtedly a limit somewhere (there is only so much you can pack into a given size) but it's really not particularly clear where that limit is. I don't think it's superintelligence - that much I agree with you - but I think "We already have a 1gb model that is as capable as it will ever be" is strictly false.

reply
The measured entropy of the model remains nearly unchanged though which means we have lost capabilities we have not measured, the model hasn't become "denser" it just became more specialized.

It's like comparing two person A and B of similar intelligence where A is smarter and B is a genius at signing, but signing was not on the test so person A won.

reply
Assumes facts not in evidence. Please show your working.
reply
https://en.wikipedia.org/wiki/Model_collapse - you want to use sigmoid 1.0, but the closer you are to 1.0 the higher the chance your model will collapse so you use 0.99-0.98, but those lead to data loss so after n passes all the original data becomes lost so you have a strict data limit there.

The rest is just the general reality I am sure you are familiar with:

- https://en.wikipedia.org/wiki/Catastrophic_interference

- https://en.wikipedia.org/wiki/Fine-tuning_(deep_learning)

- https://en.wikipedia.org/wiki/Entropy_(information_theory)

reply
You... haven't shown any evidence that we're near collapse. That's what I'm asking you for. Show me some evidence that we are losing capabilities with we have today.
reply
What proof of a ceiling are you talking about? Wouldn’t proving this require a good definition for intelligence, which I don’t think there is consensus on?
reply
- https://en.wikipedia.org/wiki/Model_collapse

- https://en.wikipedia.org/wiki/Catastrophic_interference

- https://en.wikipedia.org/wiki/Fine-tuning_(deep_learning)

- https://en.wikipedia.org/wiki/Entropy_(information_theory)

As for intelligence, the only way we have that is by allowing the model to fill the blanks which have to come from the training data. The models cannot have true intelligence for as long as they are linear models, what we see with reasoning is "boxed" intelligence where the models are effectively "modifying" themselves by feeding it's own reasoning data back into input deriving most plasible output given known information. However, the model is not able to retain what it has learned therefore that intelligence is gone the moment the session is 'full'. You can go pretty far by continiously distilling discovered information, but again all that has to come from the original training data and models own outputs, which it has to take for granted as the 'intelligence' gained is lost creating what we see is the maximum possible benchmark performance and why smaller models are not able to score as high while theoretically having the same capabilities. We can see this with larger models where they can solve tasks much faster than smaller ones as it does not require to generate the solution due to the fact that the solution is already in the training data as 'baked' intelligence and it doesn't have to 'create' it during reasoning.

reply
What model is that?
reply