But oddly enough its still great at coding. Just like a lot of people it either interfaces well with people or machines but not both.
That’s the step that causes the most significant gains in agentic performance.
But the RL doesn’t care about anything except maximizing the score, so if you only score based on coding benchmarks, anything can happen to the writing style (as long as it doesn’t hurt the coding performance).
That’s why it often gets worse on models that simply had more RL post training from the same base.
Apparently it helps generalize skills between areas, which makes sense when you compare it to how humans learn but I don't know if it's the same for LLMs.
People that produce slop have to be fired asap, they're just human relays anyway.