STEP 2 extension data: what happens after the trial stops
A design note rather than a result: what the comparator was, and what that permits you to conclude.
TheCompound Journal
Reporting on incretins, compounding & the peptide supply chain
Discontinuation
The regain trajectories, arm by arm, with the estimands named.
An earlier version gave the STEP 4 lead-in as twelve weeks. It was twenty weeks, during which participants escalated to semaglutide 2.4 mg weekly before randomisation.
The question of what happens when treatment stops has been answered more rigorously than most questions in this field, and that is unusual enough to state plainly. Three randomised designs bear on it directly: an off-treatment extension of the pivotal semaglutide obesity trial, a withdrawal design in which participants who had already reached a maintenance dose were randomised to continue or to switch to placebo, and a comparable design in the tirzepatide programme. All three point the same way, and the direction is not ambiguous.
A randomised withdrawal design begins with an open-label lead-in during which all participants receive the active drug and escalate to a target dose. Those who tolerate it and complete the lead-in are then randomised, usually two to one or one to one, to continue the drug or to receive matching placebo, and both arms are followed for a defined period with weight as the primary endpoint.
The design has two properties worth naming. Because randomisation occurs after the response, it isolates the effect of continuing from the effect of having lost weight, which a conventional parallel-group trial cannot do. And because the population has been selected for tolerating the drug, the withdrawal arm is not a general population — it is an enriched one, which makes the arm comparison internally valid and limits how far the absolute figures generalise.
Regulators favour the design for chronic-use products precisely because it answers the duration question. Its cost is ethical rather than statistical: participants who have achieved a substantial benefit are randomised to lose it, which is defensible only where the question is genuinely open and the follow-up is bounded. The Journal notes that all three withdrawal designs in this class published their regain data in full, which is more than can be said for several older obesity programmes.
The pivotal semaglutide obesity trial ran for sixty-eight weeks with a mean weight reduction of approximately 14.9 per cent on 2.4 mg weekly against 2.4 per cent on placebo.1 An extension followed a subset of participants for a further fifty-two weeks after both the drug and the lifestyle intervention were withdrawn, which makes it an off-treatment observation rather than a randomised withdrawal.
By week 120 — a year after stopping — participants who had received semaglutide had regained approximately two-thirds of the weight they had lost, finishing on average around 5.6 per cent below their original baseline against approximately 0.1 per cent for the former placebo group.2 Improvements in glycaemic parameters, blood pressure and lipids reverted broadly in step with the weight.
Two details are consistently dropped from summaries. The residual benefit was real: a mean 5.6 per cent reduction sustained a year after stopping is not nothing, and it is more than most non-pharmacological interventions achieve while they are still being delivered. And the lifestyle support was withdrawn at the same time as the drug, so the extension describes the removal of an entire intervention package rather than of a molecule.
Regain as a share of loss, as a share of body weight, and as a final position relative to baseline are three numbers. They are quoted as one.
On denominatorsSURMOUNT-4 applied the same architecture to tirzepatide with a longer lead-in. Participants escalated over thirty-six weeks of open-label treatment to their maximum tolerated dose of 10 or 15 mg weekly, achieving a mean reduction of approximately 20.9 per cent, and were then randomised one to one to continue or to switch to placebo for fifty-two weeks.3
Continuation produced a further mean reduction of about 5.5 per cent, for a total near 25.3 per cent at week 88. Withdrawal produced a mean regain of about 14 per cent of body weight, leaving that arm approximately 9.9 per cent below original baseline. The between-arm difference of roughly fifteen percentage points is similar in magnitude to STEP 4 despite the much larger initial loss.
The steeper regain in absolute terms is the expected consequence of a larger loss rather than evidence of anything peculiar to the agent. It is nonetheless the figure most often quoted without its denominator, and a fourteen-point regain from a twenty-one-point loss is a materially different statement from a fourteen-point regain from a ten-point loss. Both arms in this trial ended below where they began, and the arm that stopped ended roughly where the continued arm of the semaglutide programme did.
Tapering is borrowed language from drug classes with withdrawal syndromes, and there is no equivalent phenomenon here. With a long half-life every stop is gradual in pharmacokinetic terms whether or not the dose is stepped down.
| Study | Design | Lead-in | Randomised follow-up | Lifestyle support after |
|---|---|---|---|---|
| STEP 1 extension | Off-treatment observation | 68 weeks on drug | 52 weeks off | Withdrawn |
| STEP 4 | Randomised switch to placebo | 20 weeks to 2.4 mg | 48 weeks | Continued |
| SURMOUNT-4 | Randomised switch to placebo | 36 weeks to max tolerated | 52 weeks | Continued |
| S-LiTE | Post-diet maintenance, 4 arms | 8-week low-energy diet | 52 weeks | Continued |
| STEP 5 | Continuous treatment, no withdrawal | — | 104 weeks on drug | Continued |
| The first three are the withdrawal evidence base. STEP 5 is included because it is the only two-year continuous-treatment comparator and is frequently cited alongside the withdrawal data as though it were part of it. | ||||
The parent programmes establish the losses from which the withdrawal arms fall: approximately 14.9 per cent at sixty-eight weeks for semaglutide 2.4 mg in adults without diabetes, approximately 20.9 per cent at seventy-two weeks for tirzepatide 15 mg, and — the only continuous two-year comparator anybody has — approximately 15.2 per cent sustained at week 104 with treatment maintained throughout.45
Read together, the withdrawal evidence supports four statements and does not support a fifth. Regain begins promptly after cessation, within weeks rather than months. It proceeds at a decelerating rate, with the steepest portion in the first three to six months. It does not, within twelve months of follow-up, return participants fully to their original baseline; residual reductions of roughly five to ten per cent persist at one year in all three datasets. And continued treatment maintains and usually extends the loss, with the extension diminishing as the plateau is approached.
The statement not supported is that the drugs cause weight regain, or that stopping leaves a person worse off than never having started. Nothing in these datasets shows overshoot above the original baseline at a group level. Every arm that stopped remained below where it began at the end of follow-up.
The Journal makes this point repeatedly because the contrary claim circulates widely and is often accompanied by a mechanistic story about metabolic damage. The withdrawal trials are the direct test of that claim and they do not support it. What they do support is the unremarkable proposition that a treatment for a chronic condition works while it is being taken.
The withdrawal question changes shape when the drug was prescribed for something other than weight. In the cardiovascular outcome trial of semaglutide in overweight and obesity without diabetes, the reduction in major adverse cardiovascular events emerged over years of continued treatment, and the trial provides no information about what happens to that benefit on cessation.6 The same applies to the renal outcome data in chronic kidney disease with type 2 diabetes, where the effect on kidney disease progression was measured over a median of several years of treatment.7
There is no reason to expect an outcome benefit that accrues over years to persist after the exposure ends, and no trial has tested it. For a person taking the drug for glycaemic control, stopping has an immediate and measurable consequence in HbA1c over the following three months. For a person taking it for cardiovascular or renal risk, stopping has no measurable short-term consequence at all, which makes the decision harder rather than easier.
This is the situation in which the Journal thinks the withdrawal-trial coverage has done the most damage. Framing discontinuation as a weight question invites a person taking the drug for kidney disease to reason about it in the wrong currency entirely.
Extending the interval between doses is a dose reduction expressed in time. Framed that way it connects to the maintenance literature instead of sitting apart from it as an unstudied practice.
The correspondence this department receives on stopping divides almost evenly between people frightened by regain figures they have seen quoted without denominators and people who stopped without difficulty and cannot understand the alarm. Both groups are reading the same trials. The difference is almost entirely a matter of which number was quoted to them and whether anybody explained what it was a proportion of.
Selected from correspondence received on this article. Writers are identified by initial, surname and city, verified before printing. Replies are from the desk that filed the piece or from the standards editor. Write to letters@compoundjournal.com.
A withdrawal design also measures the effect of the surrounding programme falling away. Participants lose the visits, the monitoring and the contact at the same moment they lose the drug, and no published design separates those.
— P. Ahluwalia, Chandigarh
Blinding a withdrawal is harder than blinding an initiation, for reasons that are obvious once stated: the effects being withdrawn are perceptible. Any withdrawal trial reporting successful blinding should be asked how it was assessed.
— P. Kekana, Rustenburg
The dose that produced the change and the dose that holds it need not be the same, and the pharmacology gives no strong reason to assume they are. That is a hypothesis worth testing and at present it is being tested informally by a great many people.
— C. Farquharson, Aberdeen
What is not studied at all is what happens when somebody maintains for years and then has to stop for reasons outside their control. Every account of stopping in the literature is a planned stop, and planned stops are the minority case here.
— G. Rasmussen, Odense
Whatever is decided, the most useful thing anybody can do is write down the date of the last dose and what happened over the following weeks. The observational record in this area is poor mainly because nobody keeps one.
— P. Hollingsworth, Norwich
A design note rather than a result: what the comparator was, and what that permits you to conclude.
Every withdrawal trial compared full dose against nothing. The clinically interesting comparison — full dose against a reduced one — has not been randomised.
Real-world persistence figures, with their definitions stated, because the definitions are doing most of the work.
A design note rather than a result: what the comparator was, and what that permits you to conclude.
The evidence base is thin and the document says so, which is to its credit.
A unit is a volume. A dose is a mass. The bridge between them is concentration, and concentration is a number somebody has to calculate.