why the heck do you need 100+ instances of bert. do you even attempt to research about this before?
the laya paper show that you can do the similar stuff with jev using modern bert only: https://laya.convaiinnovations.com/
and even without the newer wave of applying llm techniques to the older bert models, even flan-t5 was trained for handling 1800+ tasks.
You really need to try that to find out?
And again have you actually tried Jev? It has a ton of world knowledge: it's able to infer user personas based on TV show watch histories using fairly recent titles... where the hell do you think that capability is emerging in 395M params?
The irony is if you really want to die on this hill, there are much better angles by focusing on LLMs that've had diffusion heads attached for fast inference with as much of a constrained decoding intelligence penalty: at least that'd put you in the ballpark.
I was being charitable that you know the field and are clueless about Jev, mea culpa for giving you the space to think I'm the one that's missing something.