Somewhat hilariously, our model is actually surprisingly bad at telling jokes in German. Guess that's not required to solve tasks in training environments.
Disclaimer: I am part of the team that trained Kolibri
It is a good approach to promote it as a German model. But I wonder if it really thinks like the German think. I speculate they just translated texts in other languages, likely English and Chinese, into German.
Hey, I worked on pre-training data Kolibri. We spent considerable time and effort to go beyond just translating. For example by building a pipeline that processes Common Crawl dumps specifically for German. You might be interested in a related blog post: https://aleph-alpha.com/en/blog/sauerkraut-not-burgers-why-g...
Very long words in German are just compounds, made from individual words or morphemes. It is the same as in English if you remove spaces. Subtokens will be equivalent with or without examples.
This is an interesting question, Chinese is way more compact in Character count vs English (about 40%?), let alone german, but yeah.. Chinese makes up for it with… total character count that runs into the thousands…