Skip to content
This GLP
← Research
Evidence

Everyone lost the same weight. Not everyone felt better.

Re-cutting 1,145 heart failure participants by frailty separated the weight loss from the symptom benefit. The frailest gained the most, and the fittest gained nothing measurable.

Dana Sullivan6 min read
Symptom score vs placebo at 52 weeksnonfrailn=110more frailn=343most frailn=692no difference

Nearly every claim made for these drugs runs through the weight. Lose the weight, the argument goes, and the knees, the sleep, the blood pressure and the breathlessness all follow. It is a reasonable argument and this analysis is the most interesting reason to doubt that it is the whole story — because it found a trial population where the weight loss was the same for everybody and the benefit was not.

How the trials were re-cut

The STEP-HFpEF program randomized people with obesity-related heart failure with preserved ejection fraction to once-weekly semaglutide 2.4 mg or placebo for 52 weeks, with two co-primary endpoints: change in body weight and change in the Kansas City Cardiomyopathy Questionnaire clinical summary score, a patient-reported measure of heart failure symptoms and physical limitation. Whether the drug worked there is the subject of a separate page. This analysis asked a narrower question of the same 1,145 participants. [1]

Frailty was scored with a cumulative-deficit index built from 34 variables, at baseline and at follow-up, and participants were sorted into three strata. The distribution is worth pausing on: 110 people, 9.6%, were nonfrail. 343, or 30.0%, were more frail. 692, or 60.4%, were most frail. A trial of an obesity drug in heart failure turns out to have enrolled a mostly frail population, which is not how these trials are usually pictured.

The weight moved the same for everyone

Semaglutide-mediated weight loss was similar across all three strata, with a P value for interaction of 0.38. Nothing about being frail made the drug less effective at the thing it is sold for. That is a useful negative finding on its own, and it sets up the contrast.

The symptoms did not

The effect on symptom score varied sharply by frailty, with P for interaction below 0.001. In the most frail, the mean difference against placebo at 52 weeks was 11.0 points, 95% CI 8.1 to 13.8 — a large improvement on a scale where roughly five points is usually treated as clinically meaningful. In the more-frail group it was 3.7, 95% CI −0.2 to 7.6. In the nonfrail group it was −1.5, 95% CI −8.4 to 5.4.

That last number is the one worth sitting with. The point estimate is negative, which means the nonfrail participants on semaglutide reported slightly worse symptom scores than the nonfrail participants on placebo. The interval runs from −8.4 to 5.4 and the group is 110 people, so it demonstrates nothing in either direction. It is printed here because omitting an unflattering estimate while quoting the flattering ones is how a result gets sold rather than reported — the same problem composite endpoints create when nobody looks at the components.

Frailty itself improved

The analysis also measured whether the frailty index moved. It did. The odds of being classified nonfrail at 52 weeks were 3.16 times higher on semaglutide than on placebo, 95% CI 2.44 to 4.09, with P below 0.0001. A cumulative-deficit index counts accumulated problems across many domains, so a shift of that size over a year is substantial, and it is a different claim from losing weight.

What to distrust

The primary symptom endpoint is a questionnaire. Participants score their own breathlessness and physical limitation. In a trial where one arm visibly loses weight and the other does not, blinding holds less well in practice than on paper, and a patient who suspects they are on the active drug may score themselves differently. That is a real limitation of every symptom endpoint in this literature and it does not make an 11-point difference disappear.

There is also a subgroup caution. Three strata from one pooled population, with the smallest holding 110 people, is the kind of cut where an interaction can be real and can also be noise. This one was prespecified and the P value is small, which is the better version of that situation rather than an exemption from it. Body composition is a related and separate question, and what the weight is made of matters more in a frail population than in a robust one.

What a buyer should take from it

Two things, both modest. The benefit these drugs produce is not a simple function of pounds lost, so a seller’s weight-loss figure is not a proxy for how you will feel. And the people who improved most here were the ones in the worst shape at the start — which is the opposite of the marketing, and a reason to be skeptical of any pitch aimed at people who are basically well. Almost none of that appears on a product page, and the disclosure scorecard shows how little does.

Frequently asked

Did frail people lose less weight?
No. Weight loss was similar across all three frailty strata, with a P value for interaction of 0.38. Frailty did not blunt the drug's main effect.
Why did the fittest participants show no symptom benefit?
The analysis does not say, and with 110 people and an interval from −8.4 to 5.4 the estimate shows nothing either way. One plausible reading is that people with few symptoms have little room to improve on a symptom score.
Is an 11-point change large?
Yes. On the KCCQ clinical summary score, roughly five points is usually treated as clinically meaningful, and the interval here ran from 8.1 to 13.8.
Does this apply outside heart failure?
Not directly. These were people with obesity-related HFpEF, most of them frail. The useful transferable point is that symptom benefit and weight loss moved independently, which nothing on a product page acknowledges.

Sources

  1. [1] Pandey A, et al. (2025). Frailty and Effects of Semaglutide in Obesity-Related HFpEF: Findings From the STEP-HFpEF Program JACC: Heart Failure. PMID 40956259

Where to get it

Best GLP-1 injections

Every injectable seller we can verify, with the price each one publishes and an honest read of what the trials measured.

Compare providers →

More in Evidence