But I think it might be more about the LLM's lack of creativity in coming up with the prompt for the image (or "embellishing" a user's prompt), or directly the embedding vector if the LLM is itself the image model's text encoder. Then whatever the image gen outputs is simply an instance of the GIGO principle. It would be interesting to test whether similar cliches and motifs also occur if you ask the model to create an SVG rather than a raster image.