Unfortunately this is rarely clean. Its also easy to make mistakes. Sometimes two arms are fundamentally incomparable. The quality and rigor of the comparison is often determined by a lot of extra checks and validations, and different journals demand different levels of rigor. I need to read it carefully to judge if this is good or not.
I haven't worked on these designs, but I remember the methodologist that taught me this in grad school giving us a lecture about this.
EDIT: the BMJ article (laudably) provides access to the analyis code, although I won't have time to review it:
github.com/nilskruger/Tirzepatide-and-the-Risk-of-Atherosclerotic-Cardiovascular-Events
Given the size of the dataset, the effect size, significance and sensitivity testing they did I think it’s very strong evidence for GLP1s causing this and it would be very very surprising to me to see the effect disappear even if they had perfect socioeconomic data.
Otherwise, it's just a waste of our time.
(I am not a medical doctor)
(You can look this up regarding biofilms, I just did today.)