Most observational drug comparisons on this site fail on the same problem: people who get the newer drug differ from people who do not, in ways no adjustment fully captures. This one built three defenses against that, and it is worth reading for the method as much as the result [1]. It is the standard this site applied in the study that checked its own work.
First, the comparator. Tirzepatide initiators (35,353) were compared not against untreated people but against sitagliptin initiators (17,618) — a drug chosen because its cardiovascular effect is neutral, making it a proxy for placebo among people who are being actively treated. Second, propensity score overlap weighting balanced baseline characteristics. Third, the authors tested two negative control outcomes, lumbar radiculopathy and abdominal hernia, which no drug in this class should affect. Neither showed an association.
On the result, major adverse cardiovascular events occurred in 2.9% of the tirzepatide group and 4.4% of the sitagliptin group at one year. That is a risk difference of 1.4 percentage points, a number needed to treat of 70, and a hazard ratio of 0.68. Heart attack drove it, at a hazard ratio of 0.67 and an NNT of 130 — an absolute gain of the kind this site keeps separating from relative ones in sixteen percent lower, six hundredths of a point.
Two findings sit outside the cardiovascular story and deserve reporting without explanation. Infections requiring hospital admission were lower on tirzepatide (HR 0.64, 95% CI 0.55 to 0.75, NNT 48), as was infection-related mortality (HR 0.40, 95% CI 0.26 to 0.61). All-cause mortality came in at 0.55 (95% CI 0.42 to 0.72). Those are large and this site will not speculate about why they appear.
The limits are ordinary and real. Follow-up ran one year and was censored at discontinuation or switching, so this measures the first year of continuous use rather than a durable effect. And claims data records prescriptions and diagnoses, not people. What makes the study unusual is that its authors asked their own data to prove it was not lying — the practice this site wants more of, and rarely gets, as in the result extrapolated to everyone.