Getting it to write well is really hard because there’s no real way to verify whether it’s good prose or not. You and I can tell, but we can’t write a verifier that codifies our judgment.
Maybe they’ll find a way to improve this, but for now it’s certainly one of the harder problems to solve for LLMs.
Part of it is that I think they also have poor theory of mind, which I imagine is also a hard thing to train it to do.
In any event, other LLMs may not automatically have the problem.