This is well past prompt engineering and into process engineering like six sigma. Just like in an early industrial revolution factory, we’re all still figuring out what works in the process of making stuff except this is so early that even simple things like “make this screw standardized” (or “no, keep going” in this case) is really high impact.
The degrees of freedom an LLM has is so large that we're going to be exploring their capabilities for decades, especially if they continue to get better. This is why IMO experts are always going to be better at LLMs in their field because they can force them LLM into processes (think prompt engineering -> CC dynamic workflows) that follow their work processes and get much better results out of them than “keep going.”
What a weird species of halting problem…
in my head, the comparison is the multi-paragraph prompts (borderline essays) i would read in various communities on reddit and similar forums, that people (often self-proclaimed "prompt engineers") said were "required" to get good output. or some of the prompts ive read in various logs that are like a thousand words of setup.
even looking back at the first prompts i was sending when i started to use chatgpt were (in hindsight) crazy long and full of unnecessary guidance/caveats/"ignore xyz"/etc.
Its quite likely they now found the counterexample with a more serious prompt, and then for virality re-tried a few times with meme-prompts like "you should do a breakthrough", knowing that the model is capable of solving this particular one. Worst case the meme-prompts don't work and they share the real one they initially used.
i am not sure why this is "quite likely". it'd be pretty silly to get a mathematical breakthrough and then hide it for an undisclosed amount of time to get a few more likes on a tweet, when the impressive part is the breakthrough.
not saying your theory is impossible, but i think the simple answer is that the model is just smarter than o1 and o3.
and, in any case, the model ended up getting the result with the meme prompt and "keep going", which was what i find fun. just like how the crypto results were from prompts of, more or less, "keep going", and that's pretty damn cool.
I agree with you that obviously no prompt engineering was needed just "solve this problem", but imagine it was you doing this problem with every model, wouldn't you have tested a new model with the best prompt you had from previous iterations, maybe with some partial previous results in it, exactly to maximize your probability for a mathematical breakthrough?