Measurably, and by less than the weight loss would lead most readers to expect. Cochrane pooled fifteen randomized trials in 8,651 adults and found physical functioning on the SF-36 improving by 2.12 points against placebo, 95% CI 0.95 to 3.28, graded moderate certainty [1]. In liver disease the mapped EQ-5D utility gain was 0.03 points from a baseline of about 0.78, which is the raw material every cost-per-QALY figure for this drug is built from [2]. The exception is heart failure, where the most frail participants gained 11.0 points on a symptom questionnaire whose clinically meaningful threshold is about five [3].
Health utility is the number that turns a clinical result into a reimbursement decision, and quality of life is the outcome most often claimed and least often measured well. It sits on a scale where 1 is perfect health and 0 is death, and multiplying it by years lived produces the quality-adjusted life years that payers price. What the drug does to the scale is a separate and much better-measured question, covered in the Cochrane review of semaglutide.
What Cochrane graded, and how the grades fall
The 2025 Cochrane review pooled fifteen randomized trials in 8,651 adults and graded every outcome separately, which is the useful part [1]. Percentage body weight is high certainty, at a mean difference of −10.73% against placebo. Quality of life is moderate certainty, and the measured difference on the SF-36 physical functioning score is 2.12 points, 95% CI 0.95 to 3.28. Serious adverse events are graded very low. The grades fall away as the outcome moves from the scale to the person, and that ordering is the finding rather than any single estimate inside it.
The number payers actually use
In the ESSENCE trial, 800 participants with metabolic dysfunction-associated steatohepatitis and F2 or F3 fibrosis were assessed at week 72 [2]. Mapped utility improved by 0.03 points against placebo, 95% CI 0.01 to 0.06, from a baseline of about 0.78. Participants completed the SF-36, and EQ-5D utilities were derived from those answers using a published algorithm rather than measured directly. That makes every figure an estimate of an estimate, and a directly administered EQ-5D would have produced a different number that nobody can size.
The size deserves attention because of what happens to it downstream. Three hundredths of a point is small on a scale whose whole range is one, and cost-effectiveness models multiply it by years of treatment and by population size, so a 0.03 input can produce impressive-looking totals. The authors label the analysis exploratory and the p value nominal, which places it outside the trial’s confirmatory hierarchy, and the population it describes is narrower than it sounds — the staging requirement is set out in the liver approval.
Where the effect is large, and who it is large in
The exception is instructive. A prespecified pooled analysis of the two STEP-HFpEF trials sorted 1,145 participants with obesity-related heart failure by frailty, and weight loss was similar across every stratum at a P for interaction of 0.38 [3]. Symptom score was not. In the most frail the mean difference against placebo was 11.0 points on the Kansas City Cardiomyopathy Questionnaire, 95% CI 8.1 to 13.8, where roughly five points is usually treated as clinically meaningful. In the more-frail group it was 3.7, and in the nonfrail group it was −1.5 with an interval running from −8.4 to 5.4.
That pattern is the answer to the question most readers are really asking. The same weight loss produced very different amounts of felt improvement depending on how much room there was to improve, and someone with little functional limitation to begin with should not expect the trial figures to describe them. The fuller reading of that analysis is in the frailty strata, and the trial it re-cuts is covered in the heart failure evidence.
What none of it measures
No trial in this literature asked whether people felt better in the ways they most often describe: energy, mood, confidence, the ability to get through a day. Those show up in questionnaires only indirectly, and the instruments were chosen to be comparable across medicine rather than sensitive to this drug. What was measured of daytime tiredness sits in the fatigue evidence, and it moved a tenth of the distance a person can feel.
So the honest summary is three-part. Quality of life improves, the improvement is small in general populations and large in people with a great deal of functional limitation, and the instrument doing the measuring was not built for this question. A reader weighing what the drug is worth to them personally is better served by that structure than by any of the three headline numbers on its own.