The two study designs point in opposite directions, and the reason is the comparison group. The randomized evidence is reassuring and narrow; the observational evidence is alarming and confounded. This site reads that split the same way in what pooling does to psychiatric risk estimates.
What the randomized trials found
A post hoc analysis pooled the STEP 1, 2 and 3 trials — 3,377 participants — plus 304 from STEP 5 [1]. Depression was measured with the PHQ-9 and suicidal ideation with the Columbia-Suicide Severity Rating Scale.
Baseline scores were 2.0 and 1.8, which is no or minimal symptoms. At week 68 they were 2.0 on semaglutide and 2.4 on placebo, an estimated treatment difference of -0.56 with an interval from -0.81 to -0.32.
Participants on semaglutide were also less likely to move into a more severe PHQ-9 category, at odds of 0.63. One percent or fewer of participants reported suicidal ideation or behavior, with no difference between groups.
What the health records found
A multicenter cohort across 13 South Korean hospitals compared 2,357 semaglutide and 6,953 liraglutide initiators against propensity-matched people who started neither [2].
| Outcome | Semaglutide | Liraglutide |
|---|---|---|
| Psychiatric disorders overall | 2.02 (1.30–3.14) | 1.66 (1.36–2.03) |
| Anxiety disorder | 2.39 (1.22–4.71) | 1.68 (1.45–1.96) |
| Depressive disorder | 3.42 (1.51–7.74) | 1.51 (1.14–2.00) |
| Gastrointestinal dysmotility | 3.91 (1.42–10.82) | — |
Read the widths. An interval from 1.42 to 10.82 is compatible with a small effect and an enormous one, and the two largest point estimates are the two least precise. The liraglutide figures, drawn from three times as many initiators, are smaller and tighter.
Why the designs disagree
Someone who starts a weight-management drug has usually just seen a doctor, is being monitored, and is often distressed about their weight. Their matched non-initiator is not necessarily doing any of those things.
Higher recorded rates of anxiety and depression in a group with more clinical contact can reflect more diagnosis rather than more disease. Matching on baseline variables does not remove that, because the contact happens after matching. This site set the general problem out in the trials and the records disagreeing on Alzheimer's, where the same shape appears with the signs reversed.
A randomized trial removes the problem by construction: both arms are in the same study, seen on the same schedule, assessed with the same instrument. That is why the PHQ-9 result carries more weight here than the hazard ratios, despite the hazard ratios being larger and louder.
What is still open
The authors of the cohort add their own caveat: semaglutide-associated vision impairment and liraglutide-associated hepatic impairment lost significance in some sensitivity analyses.
And nobody has run a randomized trial in people with established serious mental illness, which is the group a reader worrying about this is most likely to belong to. Until somebody does, the honest position is narrow. The drug has not been shown to worsen mood in people without a psychiatric history, and it has not been properly tested in people with one. The adjacent question of motivation is in a secondary finding on motivation.