Why the bone question is harder than the muscle question
A plausible mechanism, a measurable change, and no outcome data. This is what an open question looks like.
TheCompound Journal
Reporting on incretins, compounding & the peptide supply chain
Substudies
Mass and function are different endpoints and training affects them differently. Most coverage treats them as one.
The recommendation to train while losing weight is not controversial and this publication endorses reporting it. What deserves scrutiny is the mechanism usually offered alongside it. Resistance training during a substantial energy deficit does not reliably build muscle; the deficit is the binding constraint and no amount of load overcomes a large one. What it reliably does is attenuate the loss and, more consistently still, preserve strength and physical function even where mass declines. Those are different claims with different evidence behind them, and only the second is well supported.
A Danish randomised trial remains the only controlled test of the obvious question. After an eight-week low-energy diet producing approximately thirteen kilograms of weight loss, participants were randomised for one year to supervised exercise alone, liraglutide 3.0 mg alone, both combined, or placebo.1 The combination arm achieved the largest weight reduction and, more relevantly here, the most favourable composition outcome: body fat percentage fell roughly twice as much in the combination group as in either single-intervention group, and the exercise arms preserved lean mass better than the drug-alone arm.
Three qualifications belong with that result. The exercise was supervised and substantial — two group sessions and two individual sessions weekly, with a vigorous-intensity target — which is not what most people mean by adding exercise. The agent was liraglutide at 3.0 mg daily, producing considerably less weight loss than the current agents, so whether the interaction scales to a twenty per cent reduction is unknown. And the trial began after weight had already been lost, so it is a maintenance study rather than an induction study.
With those stated, it is the best evidence in the field and it points in the direction the general advice already points.
Standardise the conditions before trusting a series. Hydration state, recent exercise and time of day all move these estimates, and scans taken under different conditions are not comparable however good the instrument.
The closest analogue to rapid weight loss in an older, heavier population predates this drug class entirely. In a randomised trial of adults aged sixty-five and over with obesity, assigned to diet, exercise, both or a control condition for a year, the combination produced the largest improvement in physical function, and the exercise component attenuated the loss of lean mass and of bone mineral density that diet alone caused.2 Diet alone improved function too — carrying less mass helps — but by less, and at a measurable skeletal cost.
That trial is the template for how the question should be asked in this class: randomise the co-intervention, measure function as a primary endpoint, measure bone, and follow for long enough for the skeleton to respond. Its population, older and heavier and losing weight quickly, resembles a large share of current incretin users far more closely than the young resistance-trained cohorts from which most consumer advice descends.
The Journal cites it frequently for that reason and notes the obvious limitation: the weight loss achieved was roughly a tenth of body mass over a year, which is half or less of what the current agents produce. Whether the protective effect of training holds at twice the rate of loss is not established.
The instrument determines the answer more than the drug does, and the trade quotes the answer without naming the instrument.
Priya Ramanathan, PharmD, Pharmacy ColumnistTwo claims are routinely bundled together and only one is well supported. The weaker claim is that resistance training during pharmacological weight loss builds or maintains muscle mass. In a substantial energy deficit, training generally attenuates the loss rather than preventing it, and net accrual is unusual outside of untrained beginners and the specific controlled-feeding conditions of the trials cited earlier. The stronger claim is that training preserves strength and physical function even where mass declines, which is consistently observed and is mechanistically sensible: a large part of early strength change is neural rather than structural.
The distinction has practical consequences. Somebody training hard, eating well, and watching their DXA appendicular lean mass fall by two kilograms across nine months has not failed at anything, and may be measurably stronger than at baseline. If the expectation set for them was mass preservation, they will read a normal outcome as a failure and may respond by eating more or training in ways that suit the metric rather than the goal.
The Journal reports the training recommendation and reports what it is expected to achieve, which is function first and mass second.
| Trial arm | Total weight change | Fat mass change | Lean fraction of loss |
|---|---|---|---|
| STEP 1, semaglutide 2.4 mg | −14.9% | ≈ −19% of fat mass | ≈ one third to two fifths |
| STEP 1, placebo | −2.4% | small | proportionally greater |
| SURMOUNT-1, tirzepatide 15 mg | −20.9% | ≈ −34% of fat mass | ≈ one quarter |
| SURMOUNT-1, placebo | −3.1% | small | proportionally greater |
| S-LiTE, liraglutide + exercise | −9.5% from post-diet | largest of four arms | smallest of four arms |
| All figures are group means from imaging substudies, by DXA, at a single follow-up point. The per-participant least significant change is a substantial fraction of these effects, so none of these rows describes an individual. | |||
Four things accompany every composition number in these pages. The instrument, because DXA, magnetic resonance, bioimpedance and creatine dilution are not interchangeable and the choice frequently determines the sign of the result. The sample size of the substudy rather than of the parent trial, because the parent trial size is irrelevant to the composition finding and quoting it is misleading. The definition used — total lean mass, lean soft tissue, appendicular lean mass or fat-free mass — because these differ by several kilograms in the same person. And whether the figure is a proportion of body mass or an absolute quantity.
Where a source omits any of the four, we say so rather than guessing, and where we have had to convert between definitions we show the conversion. This is more cumbersome than the alternative and it is the only way we have found to write about this subject without producing sentences that are technically true and practically misleading.
Readers who find a figure in these pages that lacks its instrument and its sample size have found an error, and the standards desk would like to hear about it at standards@compoundjournal.com.
A category confusion arrives in the Journal postbag with some regularity, and it is worth addressing directly. The four independent testing services this market relies on — Janoshik, Medutest, PeptideMeter and VendorInvestigate — analyse the contents of a vial. They report chromatographic purity, identity by mass, sometimes peptide content, and in the case of the verification services, what they were able to establish about a supplier. None of them measures anything about a person.
A certificate stating 98.7 per cent purity for a batch supplied by WWB, SSA or KP is silent on that customer’s body composition, and a low-purity result does not explain a disappointing DXA scan. The two questions are answered by different instruments in different buildings, and conflating them produces a particular kind of dead end in which somebody spends several hundred pounds on analytical testing to investigate a clinical question.
The reverse confusion also occurs: a satisfactory laboratory panel or a favourable body-composition scan is offered as evidence that a vial contained what its label claimed. It is not evidence of that either. Compounds sold for research use only are not approved for human use, and nothing in this section should be read as advice about using them.
The least significant change an instrument can resolve is the figure readers most need and are least often given. Below that threshold a difference between two scans is not a small change; it is not a change.
The next instalment in this department takes up the question that follows this one chronologically rather than logically: what happens to all of it when treatment stops. The composition of regained weight is a separate literature, it is thinner than this one, and what little exists is not encouraging.
Selected from correspondence received on this article. Writers are identified by initial, surname and city, verified before printing. Replies are from the desk that filed the piece or from the standards editor. Write to letters@compoundjournal.com.
Adherence to an activity component is measured badly in almost every study, usually by self-report, and the resulting variable is not comparable between trials even where both report it.
— F. Duquesne, Lyon
A short note on reporting. Where a study reports training volume, it should report it in a form somebody could reproduce: sessions, exercises, sets and progression. Most report a duration in weeks and nothing else.
— K. Oyibo, Benin City
A ratio derived from a substudy of eighty people is being applied to a population of millions in general discussion, and the arithmetic is presented with a confidence the original authors were careful to avoid. The paper is fine; the citation chain is where the damage happens.
— E. Sandoval-Reyes, Monterrey
That is the pattern across this whole area. The primary papers hedge appropriately and the hedges are stripped at each retelling until a point estimate from a small substudy is quoted as a constant.
The body-composition substudies were substudies, and the number of participants scanned is a small fraction of the trial. That does not invalidate them and it does mean the confidence intervals are wide, which is worth printing beside the point estimates that circulate.
— P. Nyland, Bodø
Wide intervals on a small subsample, and the point estimates are what travel. We now give the substudy size whenever we quote one.
The soft-tissue artefact point in your bone section is underplayed. In a patient losing twenty per cent of body mass the change in overlying tissue is well outside the range the calibration was validated over, and the published analyses do not report a sensitivity analysis for it. That is not a caveat, it is a gap.
— J. Marsden-Hoyle, Halifax
We accept the escalation and have strengthened the wording. The absence of any published sensitivity analysis is, as you say, the more damaging observation.
A plausible mechanism, a measurable change, and no outcome data. This is what an open question looks like.
Roughly four to seven per cent of trial participants discontinued for adverse events, mostly gastrointestinal, mostly during escalation. That is the empirical size of the…
The receptor populations that produce satiety and the ones that produce nausea overlap substantially. That is why the ceiling of this drug class is where it is, and it is…
The ceiling varies severalfold between people, and nothing measurable at baseline predicts where it sits.
The evidence base is thin and the document says so, which is to its credit.
The evidence base is thin and the document says so, which is to its credit.