I assume 100 is the max, meaning it's impossible to be 2x as smart as Muse Spark 1.1
It did far better at some tasks compared to Sol (e.g. the ARC 3 benchmark). And at those tasks, it's not just "a bit smarter": It got 30% vs less than 8% - so you're talking 2.75x more for almost 4x the coverage.
There’s also the frustration of it not quite being enough sometimes. It’s extremely capable, but I still find that it needs more concrete guidance and boundaries than other models.
If you don't believe checking the opt-out box actually opts you out, then this sentence could be said about literally any provider.