What the placebo arms of the withdrawal trials actually tell us
Three randomised withdrawal designs have tested what happens when treatment stops. Their results are consistent and they are consistently misreported.
TheCompound Journal
Reporting on incretins, compounding & the peptide supply chain
Side effects
Trial adverse-event tables count episodes reported to a study nurse. They are the best data we have and they systematically under-record the mundane.
The Journal has taken the view that the tolerability data in this class is better than its reputation and worse than its usage. It is drawn from large randomised populations with placebo comparison, weekly contact and systematic collection, which puts it well ahead of almost anything in the grey market. It also collapses duration, timing and severity into a single percentage, which is why two people can read the same table and come away with entirely different expectations. This file is an attempt to unpack it.
The pivotal semaglutide obesity trial randomised 1,961 adults to 2.4 mg weekly or placebo for sixty-eight weeks. Gastrointestinal disorders were reported by around seventy-four per cent of the active arm and about forty-eight per cent of placebo. Within that, nausea was reported by roughly forty-four per cent against seventeen per cent, diarrhoea by about thirty-two per cent against sixteen, vomiting by about twenty-five per cent against seven, and constipation by roughly twenty-three per cent against ten.1
Three features of that table are routinely lost. The placebo rates are high, which is what happens when a large population is asked systematically about gut symptoms every few weeks. The events were predominantly graded mild or moderate. And discontinuation attributable to gastrointestinal events ran to about four and a half per cent of the active arm, against under one per cent on placebo.
The gap between three-quarters of participants reporting a gastrointestinal event and four and a half per cent stopping because of one is the most informative thing in the table. Most of this effect profile is endured rather than disabling, and any account that quotes the first figure without the second is describing something other than what happened.
A number in an adverse-event table counts participants who reported at least one episode of a coded term at any point during the treatment period. It says nothing about how many episodes, how long they lasted, or how bad they were beyond a three-level severity grade defined by interference with usual activity.
This construction has predictable consequences. A cumulative figure over sixty-eight weeks is the union of many short episodes and cannot be read as a prevalence. Two populations with identical percentages can have entirely different lived experiences. And severity grading captures function rather than distress, so an episode of severe nausea that did not stop somebody working is graded moderate.
None of this is a criticism of the trials, which followed standard practice and reported it transparently. It is a caution about a specific and common misreading: that a forty-four per cent nausea figure describes a state rather than an event count. The published tolerability analyses that break events down by timing and duration are considerably more informative than the summary tables, and are cited far less often.2
Reduced fluid intake is where the real harm in this effect profile lives, and it is prevented by the least interesting measure available.
On the dehydration pathwayGastrointestinal events in this class are concentrated in the escalation phase. Reported incidence rises in the days following a dose increase, declines over the subsequent weeks at an unchanged dose, and rises again at the next increment. Analyses that plot event onset against week show a series of peaks aligned to the escalation schedule rather than a flat burden across the trial.2
Two things follow. The first is that the escalation phase is where discontinuation risk lives, which means the tolerability problem in this class is largely a titration problem. The second is that a symptom appearing eight months into stable dosing should not be attributed to the drug by default, because that is not where the drug-attributable events cluster.
There is a corollary that patients find useful and are rarely told. The worst week of a given dose is usually the first one. A person who has been unwell for four days after an increase is, on the published pattern, at the point where things typically begin to improve rather than at the beginning of a permanent state. That is a statement about a population and not a promise about an individual, and we put it that way deliberately.
Periprocedural guidance in this area has moved quickly, and versions still circulating have been superseded. Any publication describing it should date what it describes, because a reader acting on a withdrawn version is acting on nothing.
| Event | Semaglutide | Placebo | Excess |
|---|---|---|---|
| Any gastrointestinal disorder | ≈74% | ≈48% | ≈26 pts |
| Nausea | ≈44% | ≈17% | ≈27 pts |
| Diarrhoea | ≈32% | ≈16% | ≈16 pts |
| Vomiting | ≈25% | ≈7% | ≈18 pts |
| Constipation | ≈23% | ≈10% | ≈13 pts |
| Discontinuation for GI event | ≈4.5% | <1% | ≈4 pts |
| Cumulative participant incidence from the primary publication, rounded. Excess is arithmetic difference in percentage points and is not a risk ratio. Most events were graded mild or moderate. | |||
Discontinuation for adverse events ran to roughly four and a half per cent on top-dose semaglutide and between four and seven per cent across the tirzepatide dose range, against one to three per cent on placebo. The great majority of those discontinuations were gastrointestinal and the great majority occurred during escalation.13
Those figures should be read as a floor. Trial participants receive weekly contact, free product, a nurse who can be telephoned, and an investigator with a strong interest in retention, and they are pre-selected by their willingness to enter a trial. Real-world persistence data for this class is markedly worse, with a substantial proportion of people no longer filling prescriptions at twelve months, for reasons that combine tolerability with cost and supply.
The Journal draws one inference. If most intolerance-driven discontinuation happens during escalation, and escalation practice is the least evidence-based part of the treatment course, then the largest available improvement in outcomes in this class is probably not a new molecule. It is a better answer to the titration question, which nobody has run a trial to obtain.4
Nausea: the sensation preceding or in place of vomiting; a symptom. Vomiting: forceful expulsion of gastric contents; a sign. Retching: the effort without the expulsion. Early satiety: fullness disproportionate to volume consumed. Dyspepsia: upper abdominal discomfort, often used loosely to include all of the above.
Gastroparesis: a clinical diagnosis of delayed gastric emptying with characteristic symptoms and no mechanical obstruction. It is not a synonym for drug-induced emptying delay, and the two are conflated constantly. Ileus: failure of propulsion without mechanical obstruction. Obstruction: mechanical blockage.
Incidence: proportion of a population experiencing at least one event in a period. Prevalence: proportion affected at a point in time. Adverse-event tables report the first and are read as the second. Adjudicated: reviewed against predefined criteria by a committee blinded to treatment, which is a materially stronger standard than a reported term.
Nothing in this department is medical advice. Severe or persistent abdominal pain is not a tolerability question and is not filed as one here, and the clinicians quoted in these pages say the same in their own words.
A closing note on the market this publication covers. Everything above assumes the symptom is the molecule. For material sold for research use only it may be the content, the counter-ion, the reconstitution, or the endotoxin, and none of those is visible on a purity certificate. A person reasoning carefully about tolerability while holding a vial of unmeasured contents is reasoning carefully about the wrong variable.
Selected from correspondence received on this article. Writers are identified by initial, surname and city, verified before printing. Replies are from the desk that filed the piece or from the standards editor. Write to letters@compoundjournal.com.
Duration of follow-up drives incidence mechanically: a longer study finds more of everything. Comparing rates between studies of different lengths without annualising is a category error and it is done constantly.
— L. Silveira, Belo Horizonte
Trial populations were selected, monitored and supported in ways that the general population is not, and event reporting in a trial is systematic in a way that reporting outside one never is. Both facts push the comparison in opposite directions and neither is usually mentioned.
— C. Wilcoxson, Des Moines, IA
The vagal afferent pathway is the part of the story with the strongest experimental support and the weakest presence in general discussion, presumably because it is harder to draw. Mechanistic coverage tends to follow what can be diagrammed.
— M. Halim, Kuala Lumpur
Three randomised withdrawal designs have tested what happens when treatment stops. Their results are consistent and they are consistently misreported.
What the published pharmacokinetics permit, what the labels state, and where the two diverge.
Almost nothing in the standard management repertoire has been tested in a randomised trial in this specific population. We say what is extrapolated and from where.
The receptor populations that produce satiety and the ones that produce nausea overlap substantially. That is why the ceiling of this drug class is where it is, and it is…
Where the curve flattens, what flattens with it, and what does not.
A drug that delays gastric emptying complicates the assumption behind every fasting instruction in perioperative medicine. The professional bodies have moved twice on this…