PeptideStack
5.2kquestions
20kanswers
220users

Why do the A1c reductions in SURPASS look so much bigger than in the obesity trials?

Asked 11 Mar 2026Modified 5 days agoViewed 7.2k times
20

I have been reading across the programmes and the A1c numbers do not line up in any way I can make sense of. SURPASS-2 reports tirzepatide lowering A1c by more than 2 points. STEP 2 reports semaglutide 2.4 mg lowering it by about 1.6. PIONEER trials report oral semaglutide at around 1.0. Those are three different drugs at three different doses, so some spread is expected, but the ordering seems to track baseline A1c as much as it tracks the drug.

What I suspect is happening is that a trial enrolling people at a baseline A1c of 8.3% has more room to fall than one enrolling at 8.0% or 7.4%, and that the reduction is therefore partly a property of who was recruited. If that is right, comparing headline A1c reductions across programmes is close to meaningless, and I would like to know what to compare instead.

Also: what counts as a clinically meaningful A1c difference in a trial? I have seen 0.3 and 0.4 percentage points quoted as non-inferiority margins, which seems small relative to the reference change value for a single person.

a1c
a1c

Glycated haemoglobin as a ninety-day glycaemic average: what a change of half a point means, why it lags, and the conditions under which it…

108 questions
clinical-trials
clinical-trials

Reading the primary literature properly: estimands, intention-to-treat versus per-protocol, confidence intervals, absolute versus relative…

913 questions
tirzepatide
tirzepatide

A dual GIP and GLP-1 receptor agonist. Questions here cover the SURPASS and SURMOUNT programmes, the practical differences from a pure GLP-1…

162 questions
t2dm
t2dm

Type 2 diabetes: glycaemic endpoints, the SURPASS and SUSTAIN programmes, dose ranges licensed for diabetes versus obesity, and interaction with…

20 questions
shareeditfollowflag
SI
askedsample_id17k2711 Mar 2026
5The baseline-dependence effect is real, well characterised, and almost never mentioned in secondary coverage. – a_lindgren 8 days ago
add a comment

3 Answers

Accepted answer first, then by votes
59

Accepted answer

Your suspicion is correct and the effect is large: as a rough rule from meta-regressions across glucose-lowering trials, each additional 1.0 percentage point of baseline A1c buys roughly 0.4 to 0.5 additional percentage points of reduction, whatever the agent. Comparing headline reductions across programmes with different baselines is therefore measuring recruitment as much as pharmacology.

The numbers, with baselines attached

TrialAgent and doseBaseline A1cA1c changeComparator change
SURPASS-1Tirzepatide 5 / 10 / 15 mg, monotherapy~7.9%−1.87 / −1.89 / −2.07Placebo −0.04
SURPASS-2Tirzepatide 5 / 10 / 15 mg~8.3%−2.01 / −2.24 / −2.30Semaglutide 1 mg −1.86
STEP 2Semaglutide 2.4 mg, T2DM + obesity~8.1%−1.6Placebo −0.4
PIONEER 1Oral semaglutide 3 / 7 / 14 mg~8.0%−0.6 / −0.9 / −1.1Placebo −0.3
SUSTAIN-6Semaglutide 0.5 / 1.0 mg~8.7%−1.1 / −1.4Placebo −0.4 / −0.4

Sources: [1] [2] [3] [4] [5].

Notice that the ordering is not purely baseline-driven either. PIONEER 1 had a baseline of 8.0% and produced −1.1 at its top dose; STEP 2 had a similar baseline and produced −1.6. So baseline explains part of the spread and exposure explains the rest. Both matter, and neither alone lets you rank the drugs.

Why baseline dependence happens

Three mechanisms, all real:

  • A floor. Glucose-lowering agents cannot push A1c much below roughly 5.5% in a person with functioning counter-regulation. Someone starting at 7.2% has at most 1.7 points of headroom before the floor; someone at 9.5% has nearly four. Group mean reductions inherit that ceiling on effect size.
  • Curvature in the underlying physiology. At higher A1c, more of the excess comes from fasting hyperglycaemia driven by hepatic glucose output, which responds strongly to these agents. At A1c near 7%, more of the residual excess is postprandial and harder to shift.
  • Regression to the mean at the group level. Trials enrol on the basis of an A1c above some entry threshold, which selects people whose measured value was on a high draw. Some of the fall in both arms is that selection unwinding — which is exactly why the placebo arm in these trials never sits at zero, and why the SURPASS-1 placebo change of −0.04 is unusual enough to be worth noticing.

That third point is the reason you must never quote a single-arm change. The interpretable quantity is always the between-arm difference, because the placebo arm absorbs the selection effect, the trial-participation effect and any secular drift in background therapy.

What to compare instead

  • The placebo-subtracted difference, not the within-arm change. STEP 2's −1.6 becomes a treatment effect of −1.2 against placebo. PIONEER 1's −1.1 becomes −0.8.
  • The proportion reaching a target, such as A1c under 7.0% or at or below 6.5%. This is more clinically legible, but it is even more baseline-sensitive than the mean change, so it only permits comparison when the baseline distributions are similar.
  • Head-to-head arms only. SURPASS-2 is the one entry in that table that licenses a between-drug conclusion, because it randomised participants between tirzepatide and semaglutide 1 mg within a single trial. Everything else is an indirect comparison resting on the exchangeability of placebo arms recruited in different years under different background therapy.
  • Note the comparator dose. SURPASS-2's semaglutide arm used 1.0 mg, which was the highest approved glycaemic dose at the time and is not the highest dose now available. That is a fair trial design and an unfair citation when the sentence "tirzepatide beat semaglutide" is written without the dose.

On the non-inferiority margin

The conventional 0.3 to 0.4 percentage point margin looks small next to a single person's reference change value of roughly 7% relative — about 0.5 points at an A1c of 7.4%. The apparent contradiction dissolves once you notice they are measurements of different things.

The RCV governs whether one person's two draws differ. The trial margin governs whether two group means differ, and the standard error of a group mean falls as one over the square root of the sample size. With several hundred participants per arm, the standard error on a mean A1c change is on the order of 0.05 points, so a 0.3-point difference is enormous in that currency — roughly six standard errors.

The margin is not chosen for statistical reasons anyway. It is chosen as the largest loss of efficacy that would be clinically tolerable in exchange for whatever the new agent offers, and regulators have historically settled on 0.3 to 0.4 by convention rather than derivation. It is worth being sceptical of that convention: nothing establishes that a 0.35-point A1c difference is clinically unimportant, and a chain of successive non-inferiority trials each conceding 0.3 points can drift a long way from the original comparator.

edited 25 Jul 2026 by mz_4113 — removed a claim I could not source

shareimprove this answerflag
M4
answered · acceptedmz_411399k2582 Jul 2026
3The point about successive non-inferiority trials drifting is the classic biocreep argument and it applies here. – ekaterina_volk 26 days ago
2SURPASS-2 using semaglutide 1 mg is the most commonly omitted detail in every comparison I have read. – j_wierzbicki 9 months ago
add a comment
Sponsored

Sigma-Aldrich - Certified Reference Materials

Analytical standards and reagents with traceable certificates. Every quantitative result you read inherits the accuracy of the standard behind it.

Shop standards
22

Worth adding what the A1c endpoint does not capture, because in this drug class the gap between the glycaemic endpoint and the reason anyone prescribes the drug is unusually wide.

A mean A1c reduction hides three things that matter:

  • Hypoglycaemia. Two agents achieving the same A1c with different hypoglycaemia rates are not equivalent. The glucose-dependence of insulin secretion in this class is the reason its A1c reductions come with low hypoglycaemia rates as monotherapy, and the reason that rate rises sharply when combined with sulfonylureas or insulin. The A1c number is silent on all of it.
  • Variability. A1c cannot distinguish a flat 8.5 mmol/L from an average of 8.5 mmol/L composed of 4.0 and 14.0 mmol/L excursions. The coefficient of variation on a sensor record can.
  • Everything non-glycaemic. Weight, blood pressure, lipids, and — critically — the cardiovascular and renal outcomes, which in this class are not proportional to the A1c effect. FLOW used the 1.0 mg glycaemic dose and produced a renal composite benefit; the historical trials that lowered A1c hard by other means did not produce comparable macrovascular benefit. A1c is a marker of glycaemic state, not a proxy for the value of the therapy.

There is also a mundane reporting issue that changes the numbers you read. Trials report either a treatment-policy estimand, which counts everyone as randomised including those who stopped the drug or added rescue therapy, or an on-treatment estimand, which conditions on adherence. The on-treatment figure is always the larger one, sometimes by 0.2 to 0.3 points on A1c and much more on weight. When two sources quote different numbers for the same trial arm, this is usually why, and neither is wrong — they answer "what happens if this is prescribed" versus "what happens if this is taken".

Check which estimand a figure came from before you put it in a comparison table. The mismatch is silent, it is common, and it is always in the direction of flattering whichever arm had better adherence.

shareimprove this answerflag
TA
answeredtri_gly_ala48k3821 Jun 2026
8

One narrow arithmetic point, since the question mentioned wanting to do this properly.

If you want to convert a trial's A1c reduction into average glucose to make it concrete, use the linear eAG relationship and remember that because it is linear, the offset cancels when you take a difference. So:

Δ eAG (mg/dL) = 28.7 × Δ A1c
Δ eAG (mmol/L) = 1.5944 × Δ A1c

Worked for the entries in the table above:

  • SURPASS-2, tirzepatide 15 mg, −2.30 points: 1.5944 × 2.30 = 3.67 mmol/L of average glucose, or 28.7 × 2.30 = 66 mg/dL.
  • STEP 2, semaglutide 2.4 mg, placebo-subtracted −1.2 points: 1.5944 × 1.2 = 1.91 mmol/L, or 34 mg/dL.
  • PIONEER 1, oral 14 mg, placebo-subtracted −0.8 points: 1.5944 × 0.8 = 1.28 mmol/L, or 23 mg/dL.

The −46.7 term never appears, because it cancels: (28.7·A − 46.7) − (28.7·B − 46.7) = 28.7·(A − B). Worth stating explicitly because people occasionally subtract it twice and produce a nonsense figure.

Expressed this way the differences are easier to weigh. Three and a half millimoles per litre off a mean glucose is a large physiological change. One and a quarter is real but modest. That framing is more informative than the percentage-point figures, which compress large differences into small-looking decimals — and it is a useful corrective when a 0.3-point non-inferiority margin is described as negligible, since 0.3 points is about 0.5 mmol/L of average glucose, which nobody would describe as nothing if it were reported that way.

shareimprove this answerflag
KL
answeredkirsi_lahtinen45k3826 Mar 2026

Your answer

Ask PeptideStack is a static archive. Posting is closed, but the norms are worth stating: answer the question that was asked, show your working, cite the trial or the certificate, and say plainly where the evidence runs out.

Not medical advice. Research-use-only compounds are not approved for human use.