PeptideStack
5.2kquestions
20kanswers
220users

Why does STEP 1 report -14.9% when other papers quote about -17% for the same 2.4 mg dose?

Asked 19 Jun 2024Modified 2.0 years agoViewed 9.1k times
26

I have been trying to build a spreadsheet of published weight-loss results so I can compare programmes properly, and I keep tripping over the same thing. The primary STEP 1 publication gives a mean change in body weight of -14.9% for semaglutide 2.4 mg at week 68. But secondary analyses, conference decks and at least two review articles I have read quote something closer to -17% for the same trial, same dose, same 68 weeks.

These cannot both be the headline result, so one of them must be a different analysis of the same dataset. I have seen the words treatment-policy estimand and trial-product estimand in the statistical appendix but the appendix assumes I already know what those mean, and every plain-English explanation I find collapses it to "intention-to-treat versus per-protocol", which I do not think is quite right either.

Concretely, what I want to know is: (1) what is the actual difference in what the two numbers are estimating, (2) which one should I use if I am trying to answer "what would I expect to happen to a person who starts this drug", and (3) is the same split present in the SURMOUNT papers, because tirzepatide results also seem to come in two flavours depending on where I read them.

I am not asking which number is bigger. I am asking which question each number is the answer to.

semaglutide
semaglutide

A GLP-1 receptor agonist with a fatty-acid-acylated backbone and a roughly one-week half-life, marketed for type 2 diabetes and for weight…

360 questions
clinical-trials
clinical-trials

Reading the primary literature properly: estimands, intention-to-treat versus per-protocol, confidence intervals, absolute versus relative…

913 questions
dosing-math
dosing-math

The arithmetic itself: milligrams to millilitres to insulin units, concentration after reconstitution, dose per draw, and vial-days per vial. Show…

811 questions
shareeditfollowflag
MO
askedmarta_okonkwo87k25819 Jun 2024
8The ICH E9(R1) addendum is the source document for the estimand language if you want the primary reference. – kwn_analytical 4 months ago
7Worth noting the two numbers are not sponsor spin - both are prespecified and both are in the same paper. – tare_weight 2 months ago
add a comment

3 Answers

Accepted answer first, then by votes
71

Accepted answer

They answer two different questions, and neither is "intention-to-treat versus per-protocol" in the classical sense. Both analyses use every randomised participant. What differs is how each one handles the events that happen after randomisation: stopping the drug, and starting rescue therapy.

The two estimands

  • Treatment policy. The question is: what happens to weight if you assign this drug as a policy, and then let real life happen? Data collected after a participant stops the drug or adds another weight-loss intervention still counts, at the value it actually was. Discontinuation is treated as part of the effect of the policy, not as a nuisance to be removed. This is the primary estimand in STEP 1 and it produced -14.9% versus -2.4% for placebo, a difference of about -12.4 percentage points [1].
  • Trial product. The question is: what is the pharmacological effect of the molecule if it is taken as the protocol intends, without rescue medication? Observations after discontinuation or rescue are handled as if that deviation had not occurred, using a model that borrows information from participants who remained on drug. In STEP 1 this shifts the semaglutide arm to roughly -17% while leaving the placebo arm essentially where it was, because placebo participants who stopped were not losing much anyway.

Note what that asymmetry tells you. The gap between the two numbers in the active arm is almost entirely the arithmetic of dilution: roughly one in six participants was off drug by week 68, and their weight had partly come back. In the placebo arm there was nothing to come back from, so the estimand choice barely moves it.

Which one you want

If you are asking "what does prescribing this drug to a population achieve", treatment policy is the honest answer, because discontinuation is a real and large part of what happens. If you are asking "what does this molecule do to adipose tissue in someone who keeps taking it", the trial-product number is closer. Regulators generally want the first. People comparing molecules head-to-head usually want the second, because it is less contaminated by trial-conduct differences.

The failure mode to avoid is mixing them across trials. Quoting the trial-product figure for one drug and the treatment-policy figure for another manufactures a difference out of nothing but analysis convention.

Yes, SURMOUNT does the same thing

SURMOUNT-1 reports a treatment-regimen estimand of -15.0%, -19.5% and -20.9% at 5, 10 and 15 mg versus -3.1% for placebo, and an efficacy estimand of about -16.1%, -21.4% and -22.5% [2]. Same structure, different labels. "Treatment regimen" maps to treatment policy; "efficacy" maps to trial product. The vocabulary is not standardised between sponsors, which is a large part of why this confuses everyone.

Practical rule for your spreadsheet: add a column recording which estimand each row came from, and refuse to compare rows that disagree. If a source quotes a number without saying which, treat the number as unusable rather than guessing.

edited 9 Aug 2024 by Dr_Nadia_Farsi — tightened the wording; no substantive change

shareimprove this answerflag
DF
answered · acceptedDr_Nadia_Farsi90k25817 Jul 2024
2The asymmetry point is the one most reviews miss - the estimand choice moves the active arm far more than placebo. – fib4_reader 30 days ago
Adding an estimand column to my own table immediately killed three comparisons I had been making. – Dr_Malik_Osei 9 months ago
add a comment
Sponsored

Janoshik Analytical - Independent Third-Party Testing

HPLC purity, identity confirmation and quantified content on the vial you actually hold. Reports arrive with the chromatogram attached, not just a number.

Submit a sample
Sponsored — paired listing

GL Biochem (Shanghai) Ltd. - Direct Synthesis

Founded 1998. ISO 9001 and cGMP certified, 1,500+ staff and 200+ patents. The synthesis house behind a great many of the vials that get sent out for testing - batch-specific documentation with every order.

Visit GL Biochem
33

Adding the part the accepted answer leaves implicit: "per-protocol" as normally understood is a genuinely different and worse thing, and it is worth being able to spot it because older literature is full of it.

A classical per-protocol analysis drops participants who deviated. That is a post-randomisation selection, so it destroys the thing randomisation bought you. If the people who stop the drug are systematically different - and in weight-loss trials they are, because non-responders and people with intolerable nausea are enriched among the stoppers - then deleting them biases the result in an unpredictable direction. It usually flatters the drug, but not always.

The trial-product estimand is not that. It keeps everyone and uses a model to answer a hypothetical question about what would have happened without the deviation. That is still an assumption-laden analysis, and the assumption is specifically that the missing on-drug trajectory of a stopper resembles that of a continuer with similar observed history. That assumption is doing real work and it can be wrong. But it is a stated assumption applied to the full randomised set, not a silent exclusion.

Practical checks I apply when reading one of these papers:

  • Find the retention numbers. If treatment discontinuation is under about 10%, the estimand choice barely matters and you can stop worrying. STEP and SURMOUNT are both above that.
  • Find the rescue-therapy rules. A trial that permits rescue in the placebo arm will show a smaller separation under treatment policy, and that is not a weakness of the drug.
  • Look for a tipping-point or multiple-imputation sensitivity analysis. Its presence tells you the statisticians were worried about exactly this, and its magnitude tells you how much to trust the primary number.

Also worth internalising: the completer-only figures that circulate informally are neither of the two estimands. They are the classical per-protocol analysis, and they are the largest numbers in circulation for precisely the reason that makes them least trustworthy.

shareimprove this answerflag
TI
answeredteodora_ilic15k286 Jul 2024
18

One more source of the "same trial, different number" problem that is not about estimands at all: responder proportions and mean change get quoted interchangeably in secondary coverage.

STEP 1 reports that about 86% of participants on semaglutide reached at least 5% loss, 69% reached at least 10%, 50% reached at least 15% and roughly a third reached at least 20% [1]. A slide that says "half of patients lost 15%" and a slide that says "patients lost 14.9%" are describing the same dataset from two directions, and if you record only the number you will end up with a table that cannot be reconciled.

The distributional view is also more useful than the mean for most purposes. A mean of -14.9% is compatible with a very wide spread, and it is. The responder curves in these programmes are broad, with a meaningful tail of people who lose almost nothing. If you are trying to reason about what an individual should expect, the quartiles matter more than the mean, and they are usually in the supplementary appendix rather than the abstract.

Nothing here is medical advice and none of these figures transfer cleanly to a research-use-only compound of unverified content; a trial number is a statement about a specific molecule at a specific dose under supervision.

shareimprove this answerflag
DV
answereddead_volume49k388 Aug 2024

Your answer

Ask PeptideStack is a static archive. Posting is closed, but the norms are worth stating: answer the question that was asked, show your working, cite the trial or the certificate, and say plainly where the evidence runs out.

Not medical advice. Research-use-only compounds are not approved for human use.