It depends on the procedure, and the four largest studies do not agree with each other. On knee replacement, a 415-pair comparison of semaglutide against tirzepatide found no significant difference on any outcome [1]. On shoulder surgery, three alarming raw signals — up to a near-fourfold rise in urinary infection — vanished once the comparison was adjusted for who actually gets prescribed a GLP-1 [2]. On spinal fusion, the same drug class looked protective in one database, neutral in another, and nearly five-times worse in a third [3]. On neck fusion, doubling the dose changed nothing measured [4]. No single answer covers all four, and that disagreement is itself the finding.
A study that finds nothing is often the most useful kind. The knee replacement comparison found nothing, and then described what it found in a way worth separating from the data.
What was compared
Adults with type 2 diabetes who had a primary knee replacement between mid-2022 and the end of 2024, and who had an active prescription for semaglutide or tirzepatide within the 90 days before surgery. [1] They were matched one to one on age, sex, race, BMI, glycated hemoglobin, comorbidity burden and concurrent diabetes medication, leaving 415 pairs. Outcomes were tracked to 90 and 180 days.
Nothing reached significance. Medical complications at 90 days came in at OR 1.122, 95% CI 0.736 to 1.710. Surgical complications at 180 days at OR 1.632, 95% CI 0.845 to 3.152. Individual events — heart attack, stroke, pneumonia, sepsis, clots, kidney injury, wound problems, joint infection, death, revision — were all similar, every p above 0.05.
The sentence that goes further than the data
The authors conclude that the two drugs demonstrated similar short-term safety and utilization profiles, and that the findings support continued perioperative use of either agent.
The first half is a reasonable summary of a null result if you read “similar” loosely. The second half is a recommendation, and a failure to detect a difference among 830 patients does not establish that no difference exists. Concluding equivalence takes a study designed to rule out a difference of a stated size, and this was not one. It is the same distinction between what a design was built to answer and what gets claimed from it.
Not the opposite claim either
None of this suggests either drug is risky around surgery. The point estimates cluster near no difference, nothing looked alarming, and the study is perfectly good evidence that no large difference was detected.
The honest position is narrower than either headline: on this evidence nobody can tell the two apart for knee replacement, and the reason is that nobody has looked at enough patients. Whether these drugs should be paused before surgery at all is a separate question with its own page here, and it is the question a patient is more likely to face.
Shoulder surgery: the alarming numbers that did not survive adjustment
Among 5,876 people with type 2 diabetes having a rotator cuff repair, 233 were on a GLP-1 beforehand. Compared directly, they looked worse on three outcomes: respiratory complications at an odds ratio of 1.97, urinary tract infection at 3.85, and postoperative stiffness at 2.20 [2]. Once the two groups were reweighted to resemble each other on measured characteristics, every one of those differences disappeared. Nothing about the patients had changed; only the comparison had. The raw figures were measuring who gets prescribed a GLP-1, not what the drug does — the full breakdown is in the shoulder surgery adjustment.
Spinal fusion: protective in one database, worse in another
Twelve retrospective cohorts on GLP-1 use and cervical spine fusion give three different answers depending on which database and which procedure [3]. For anterior cervical fusion, one large claims network put the odds of failed fusion at 0.52 against non-users; a second, independent network found a neutral 0.86. A semaglutide-specific cohort of posterior cervical fusion went the other way entirely, at 4.79. Pooling all twelve gives an odds ratio of 1.28, with a confidence interval from 0.35 to 4.70. That range spans a two-thirds reduction and a near-fivefold increase. It is evidence the studies should not have been pooled, rather than a finding about the drugs. The procedure-by-procedure detail is in the spinal fusion review.
Neck fusion: doubling the dose changed nothing
A fourth study asked a narrower question: among people already on a GLP-1 before neck fusion surgery, does a higher dose change the outcome [4]? Comparing 921 matched pairs on a high dose against 921 on a standard dose, nothing differed across nine measures — readmission, emergency visits, complications, swallowing problems, opioid use, and failed fusion. Both dose groups also fared better than matched patients taking no GLP-1 at all, though that second comparison carries the weaker evidence, for the same reason the knee and shoulder comparisons above do. The full results are in the neck fusion dose comparison.
Why this pattern matters beyond knees
Because “no significant difference” is the most frequently over-read phrase in medical writing. It gets reported as reassurance, quoted as equivalence, and eventually cited as evidence that two things are the same.
The test takes a moment: look at the interval rather than the p-value, and ask what the upper bound would mean if it were true. If it would matter, the study did not settle anything. Almost nothing a seller publishes includes an interval at all, which the census of what goes unsaid measures from the other end.