Accepted answer
Read the 2.4 mg row, not the pooled one. A programme that randomised more than one dose level reports each arm separately, and the figure that circulates afterwards is usually either the top-dose arm or an average across arms nobody was randomised to. If STEP 1 ran a 2.4 mg arm, that row carries its own sample size and its own confidence interval, and both are narrower than the trial-level ones by roughly the square root of however many arms there were. Take the primary publication rather than the press release: one reports by arm, the other reports whichever number is largest. A dose level inside a trial is a protocol decision made under supervision, not a recommendation, and nothing here is medical advice.
The trial answers a narrower question than the headline suggests, and the narrowing is where the useful information is.
Placebo arms in this class are not nothing. Lifestyle-intervention placebo arms in the major obesity trials commonly lose two to three per cent of body weight, so an active-arm figure quoted without its comparator overstates the drug effect by roughly that much.
Non-inferiority and superiority designs are not interchangeable. A non-inferiority result says the new agent is not meaningfully worse against a pre-specified margin — it does not say it is as good, and it certainly does not say it is better.
Where a result is quoted from a conference abstract rather than a peer-reviewed publication, the numbers routinely move between the two. It is worth checking which one you are reading.
I am not a clinician and this is not medical advice; it is a reading of a published protocol.
If a claim cannot be traced to a named trial with a named endpoint, treat it as a claim rather than as evidence.