upvote
No, I'm referring to the "we're a massive enterprise and we don't make money by actually figuring out issues, so if service XYZ has troubles after 30 hours of uptime, then we'll restart it every 20 hours and problem solved" phenomenon.

Or the phenomenon where I piss blood explaining why having a p99 that is 30x worse than our p95 is maybe possibly a problem that should be looked at.

Or the phenomenon where concurrency exists, and so issues are no longer reliably reproducible, meaning everyone just throws their hands up and tries to ignore and downplay them as much as humanly possible. That is until a dickhead like me comes around, and does something like a scripted 300 restart cycle test overnight until a clustering resiliency defect finally reproduces, and i can capture enough debugging data that would never be possible on a live environment. All the while the product vendor is twiddling their thumbs, waiting for us to provide said data on a silver platter, because for some reason this completely stock issue doesn't reproduce on their end in a pretty much identical environment, or so they say.

reply