Nobody has run that trial. A network meta-analysis pooled 93 randomized trials of every approved type 2 diabetes class [1]. Against placebo, metformin had the lowest odds ratio for major cardiovascular events. It was 0.69, 95% CI 0.53 to 0.89. GLP-1 agonists came in at 0.87, 95% CI 0.83 to 0.92. Tirzepatide and SGLT2 inhibitors were both 0.82. Those four intervals overlap. The authors grade the metformin evidence as lower quality. Every trial pooled here enrolled people with type 2 diabetes. That is also the population in the kidney failure numbers needed to treat.
What was pooled
Ninety-three randomized trials of at least 40 weeks. [1] Each compared a diabetes medication against placebo or an active comparator. Cardiovascular events and deaths were formally adjudicated. Both pairwise and network methods were used. That lets drugs never tested against each other be compared indirectly.
In the network against placebo, metformin came in at OR 0.69, 95% CI 0.53 to 0.89. SGLT2 inhibitors were 0.82, 95% CI 0.75 to 0.90. Tirzepatide was 0.82, 95% CI 0.72 to 0.93. GLP-1 agonists were 0.87, 95% CI 0.83 to 0.92. Insulin, DPP-4 inhibitors and sulfonylureas showed no significant effect on events.
On all-cause mortality, tirzepatide came in at 0.72, 95% CI 0.64 to 0.82. SGLT2 inhibitors and GLP-1 agonists were both 0.85, 95% CI 0.80 to 0.91. In the pairwise analyses, sulfonylureas were associated with increased cardiovascular events. That was the only class moving the wrong way.
Why this is not a ranking
Read the four intervals, not the four point estimates. Metformin runs 0.53 to 0.89. SGLT2 inhibitors run 0.75 to 0.90. Tirzepatide runs 0.72 to 0.93. GLP-1 agonists run 0.83 to 0.92. Almost all of that overlaps.
An analysis can produce an ordering it cannot defend. The ordering is the part that gets quoted. That is the same failure as a pooled average that describes nobody. These numbers support four classes reducing events and three reducing death. They do not support a first place.
An indirect comparison is not a head-to-head
The same limit shows up inside the GLP-1 class. A network of four randomized trials covering 14,348 participants compared tirzepatide with dulaglutide [2]. One trial supplies 13,165 of them, or 91.8% of the evidence. Every published difference is measured against tirzepatide 5 mg, not against dulaglutide directly. No difference appeared in all-cause mortality or major cardiovascular events. The full reading is in one trial that is 92% of a network.
Where the GLP-1 evidence is strongest
It is strongest against placebo, not against metformin. FLOW randomized 3,533 people with type 2 diabetes and chronic kidney disease [3]. The primary composite hazard ratio was 0.76, 95% CI 0.66 to 0.88. Major cardiovascular events were 0.82, 95% CI 0.68 to 0.98. All-cause death was 0.80, 95% CI 0.67 to 0.95. The dose was 1.0 mg weekly, which is the diabetes dose. The condition attached to all of it is set out in the trial that stopped early.
The population, and who is reading
Everyone in these 93 trials had type 2 diabetes. A person buying a GLP-1 for weight loss without diabetes is not in this evidence. The cardiovascular case for that group rests on different trials in different populations.
Metformin is not a weight-loss drug and is not sold as one. A table where it leads on cardiovascular events is a table about diabetes care. It is not a shopping list. The menu a seller does carry is counted in the six obesity drugs on sale.
What a reader can use
One rule. When a comparison puts an old generic near the top, read it as a statement about evidence. Old drugs have old trials. Old trials have wider intervals and looser methods.
It cuts both ways. The generic’s apparent advantage should be discounted. The newer drug’s advantage over it has also not been demonstrated, because nobody ran that trial. And an active comparator answers a narrower question than its sentence suggests. Almost nothing a seller publishes distinguishes the two, which the disclosure scorecard counts.