Most reviews end by telling you which option is best. This one ends by explaining why it will not, and the explanation is more useful than a ranking would have been.
What was measured
Eleven randomized trials in Asian study settings, in adults with type 2 diabetes and metabolic dysfunction-associated steatotic liver disease, all adding a study drug to comparable background glucose-lowering therapy. [1] 692 participants were randomized into eligible comparisons and 642 entered the coprimary analyses. The coprimary outcomes were change in ALT, a liver enzyme, and change in HbA1c.
Adding a GLP-1 reduced ALT by 6.71 U/L, 95% CI −11.53 to −1.89, graded moderate certainty. It reduced HbA1c by 0.48 percentage points, 95% CI −0.83 to −0.14, graded low certainty. Adding an SGLT2 inhibitor reduced ALT by 7.56 U/L, 95% CI −11.24 to −3.88, at moderate certainty, with little or no clinically important further effect on HbA1c.
Two grades for one drug
Notice that the GLP-1 estimates carry different certainty ratings. Moderate for the liver enzyme, low for the blood sugar measure. That is not an inconsistency; it reflects how many trials contributed to each comparison, how consistent they were, and how precise the pooled estimate is.
A summary that reported both numbers without both grades would treat them as equally settled. They are not, and the authors went to the trouble of grading them separately — which is the sort of distinction a pooled figure usually hides.
The three refusals
First, no class ranking. The authors state the available evidence does not support ranking SGLT2 inhibitors, GLP-1 agonists and pioglitazone against one another — several comparisons rested on sparse evidence, and the two classes were not tested head to head.
Second, no claim about durability or clinical significance. Follow-up was short and the endpoints were surrogates, so the paper says both of those remain uncertain rather than projecting forward.
Third, and most striking: comparative safety could not be assessed reliably, because adverse-event reporting across the trials was heterogeneous and incomplete. Eleven randomized trials, and the authors could not compare how safe the drugs were relative to each other. That is a statement about the literature rather than about the drugs, and it is the kind of thing that never survives into a summary.
Why a refusal is worth reading
Because the demand for rankings is what produces bad ones. A reader wants to know which drug to take; a marketer wants to say theirs; and a review that declines to answer is doing the harder and more honest thing — a comparison nobody ran cannot be reconstructed from two separate ones.
The practical version for anybody reading a claim about these drugs and the liver: ask which endpoint moved, how certain the authors were, and whether they compared the drug to the alternative or only to background care. Almost nothing published by a seller answers any of the three, which the disclosure scorecard measures from the other end.