By the way, convex hull permits extrapolating past the training data. LLM won't invent a new word that could not be defined by a sequence of known words. Just if it's meaningless and fully random/hallucinated, the new knowledge won't work with other known information blocks (breaks convexity).