Here is a project that guides you through it if you want to prove to yourself that it works https://github.com/arcee-ai/DistillKit
GPU kernel optimization is just the kind of well-bounded problem with clear success criteria that AI loves.
It's about evidence this is an active force in competition in LLMs.
[1] https://www.anthropic.com/news/detecting-and-preventing-dist...
It's also how providers build their smaller models out of their larger ones; they publicly talk about the process.
Make sure to stay updated!