I'm assuming a language barrier on the part of the author. I wonder if the models would do better being prompted in the author's main language.
Also, I feel like the LLMs would have done better if they had started from scratch each time, rather than being burdened by the output from the previous attempt.