Accepted answer
Because two papers on STEP 8 are usually reporting two different estimands from the same randomisation. The treatment-policy estimand asks what happened to everyone assigned, including those who stopped; the trial-product estimand asks what happens if you keep taking it. The second is always the larger number, and both are legitimate answers to different questions. Then there is the analysis population — randomised, treated, or completers — and the handling of missing data, where a last-observation-carried-forward and a multiple imputation can differ by a point or more. Neither paper is wrong. Read the statistical methods section and you will find both figures defined in it.
Before comparing two trials, check whether they share an endpoint definition. Frequently they do not, and the numbers then are not comparable in any sense.
Duration decides what can be seen. A 68-week trial can measure weight and glycaemia; it cannot measure anything whose event rate is one per cent per year without enrolling tens of thousands.
Headline results, principal programmes
| Trial | Agent | n | Duration | Primary result |
|---|
| STEP 1 | Semaglutide 2.4 mg | 1,961 | 68 wk | −14.9 % vs −2.4 % weight |
| STEP 2 | Semaglutide 2.4 mg, T2DM | 1,210 | 68 wk | −9.6 % vs −3.4 % weight |
| SURMOUNT-1 | Tirzepatide 5/10/15 mg | 2,539 | 72 wk | −15 / −19 / −21 % weight |
| SURMOUNT-4 | Tirzepatide, withdrawal | 670 | 88 wk | Continued loss vs substantial regain |
| SELECT | Semaglutide 2.4 mg | 17,604 | ~40 mo | MACE HR 0.80 (0.72–0.90) |
| FLOW | Semaglutide 1.0 mg, CKD | 3,533 | ~3.4 yr | Renal composite reduced; stopped early |
| SURMOUNT-OSA | Tirzepatide, OSA | 469 | 52 wk | AHI reduced with and without PAP |
Confidence intervals matter more than point estimates when two trials disagree. Two studies reporting fifteen and twenty per cent whose intervals overlap heavily have not disagreed about anything.
Where a result is quoted from a conference abstract rather than a peer-reviewed publication, the numbers routinely move between the two. It is worth checking which one you are reading.
Be careful about generalising from a trial population to yourself. The exclusion criteria are usually the most informative page in the supplement.
The short version: check the endpoint, check the comparator, check who was excluded, then look at the number.
edited 29 Oct 2024 by Dr_Tomas_Kral — added the placebo-arm figures
4Thank you — this is the answer I was looking for. – wren_calloway 2 months ago add a comment