There is a real question underneath this paper, and it is a good one: people with diabetes lose toes and feet to ulcers that begin as small wounds, so anything that changes that rate would matter enormously. The trouble is that the instrument used here can answer a question about harm appearing more often than expected, and it cannot be turned around and asked whether something appears less often than it should.
What was measured
The FDA’s adverse event database was queried for every report naming a GLP-1 drug between 2005 and 2024, and for every report naming one of the other diabetes drug classes. [1] There were 1,819 of the former and 17,206 of the latter. Diabetic foot complications made up 6.38 per thousand of the GLP-1 pile and 11.31 per thousand of the comparison pile, giving a proportional reporting ratio of 0.56 with an interval running from 0.54 to 0.59, which is narrow enough to look like certainty about something.
Semaglutide came out at 0.45, dulaglutide at 0.49 and liraglutide at 0.30, and the pattern held across sensitivity analyses. The consistency is genuine. What it is consistent about is the part that needs care — and the mechanics of the database itself we have covered in a separate piece.
The direction the method runs
Disproportionality analysis was designed to find signals: to notice when a drug’s reports contain a strange term far more often than the background does, so that somebody can go and investigate properly. Running above the expected share is a reason to look. Running below it is not a finding, because the things that suppress a share are so numerous and so ordinary that no ratio can separate them.
This is the same asymmetry that makes an absent result different from a negative one, arriving through a different door.
The genetic half
The authors also ran a Mendelian randomization, using inherited variation in the GLP-1 receptor gene as a stand-in for the drug, and report that it broadly agreed with the database analysis.
That is a genuinely independent instrument, which is worth more than another pass through the same reports. It also answers a different question. Small lifelong differences in a receptor are not eighteen months of a therapeutic dose in a fifty-year-old, and a method that models a lifetime cannot tell you what happens after a year — much as a model trained on one thing keeps getting read as evidence about another.
What would settle it
A cohort with feet in it. Count the people taking each drug, examine them, and record the ulcers — the trial evidence already exists for wound-adjacent outcomes in other conditions, so the design is not exotic.
Until then, the honest summary is the one the authors themselves wrote: preliminary, requiring prospective validation. Weight loss and better glucose control ought to help a diabetic foot, so the hypothesis is reasonable on its face. Reasonable is not the same as demonstrated, and a large database is not the same as a large study.