Mass spectrometry answers a different question, properly
Where two methods disagree, the conservative convention is to report the lower figure. It is not universal, and whether a laboratory follows it belongs on the report.
TheCompound Journal
Reporting on incretins, compounding & the peptide supply chain
Analytics
Duplicate submissions under different names test within-laboratory repeatability, which is a different quantity from between-laboratory reproducibility.
All four services covered here were sent the institutional comparison table before publication and invited to correct it. Three responded and two corrections were made. One participant asked that the blind duplicate results not be attributed by name; the Journal agreed, on the basis that the methods are the finding and the methods are printed in full, and records the request here.
The headline finding is that within-laboratory repeatability was good in every case and between-laboratory reproducibility was not, and that the entire gap is explained by disclosed method choices rather than by anything hidden. Each service answered the question it was asked using the method it publishes. Each answer is defensible. The spread between them is nonetheless larger than the differences vendors compete on in their marketing, which means the number a buyer is comparing is not a property of the material.
Twelve vials, one lot, purchased at retail without disclosure of purpose. Two vials were sent to each of the three assay services under two different submitter names and addresses, so that each laboratory received two nominally unrelated submissions of the same material some three weeks apart. Six further vials were retained. Each service was asked for its standard purity determination at its standard price and turnaround, with no special instructions.
The design tests two distinct quantities that the trade conflates. Repeatability is the agreement between duplicate determinations within one laboratory; reproducibility is the agreement between laboratories. Interlaboratory studies in analytical chemistry consistently find the second to be substantially worse than the first, and the variance decomposition that separates them is standard methodology.1 The distinction between repeatability and intermediate precision is formalised in the validation guidance,2 and multi-site studies in adjacent fields have repeatedly found between-laboratory agreement on identical samples to be the harder problem.3
All three services were informed after the fact, before publication, and each was given the opportunity to comment on its own method as printed and on the comparison as a whole. All three responded. Two supplied additional method detail that has been incorporated. One disputed the framing of the comparison, and its objection is printed in the correspondence below. None of the three asked for its result to be withheld, which the Journal records because it did not have to be that way.
A genuine proficiency scheme needs a homogeneous material and an assigned value, prepared by a body that is not a participant. Nothing of that kind exists in this market, which is why comparisons here measure agreement rather than accuracy.
Within-laboratory repeatability was good. The two determinations from each service agreed to within 0.3 percentage points in every case, and to within 0.1 in one, which is about what a well-controlled chromatographic method should deliver on duplicate material and is a genuinely reassuring result.
Between-laboratory reproducibility was another matter. The three services returned figures spanning 2.1 percentage points on material from one lot. Every point of that spread is accounted for by disclosed method differences: gradient duration, integration threshold, the retention-time cut-off defining the solvent front, and whether an orthogonal second gradient was run and the lower figure reported. Rerun the raw data from the shallowest method with the fastest method’s integration threshold and the two figures converge to within 0.4 points, which is the strongest available demonstration that the disagreement is methodological rather than analytical.
Identity results agreed completely: all three found a single dominant species at the expected mass, and none reported evidence of an unrelated compound, which is the ordinary outcome of intact-mass confirmation on submitted material.4 The two services reporting peptide content returned 93% and 91% of label, a difference within the stated uncertainty of nitrogen determination. The material, in short, was what it claimed to be, and the disagreement was confined to the second significant figure of the number the market competes on.5
A laboratory cannot fail a vendor. It can only issue a report to whoever paid for it.
On the limits of the word verificationFour limitations, stated because the alternative is letting readers over-read a small study. First, one lot of one compound from one supplier is not a sample from which the performance of these services in general can be inferred; it is an existence proof about method-driven spread. Second, three services is too few for any statistical treatment beyond the descriptive; published round-robin studies of peptide purity use seven or more participants for exactly that reason.6
Third, and most important, the exercise tested reproducibility, not accuracy. All three could be equally wrong: without a certified reference standard of known purity, there is no true value against which to score them, and the compendial approach to validating a purity procedure requires exactly such a reference to establish accuracy rather than mere agreement.7 What we measured is dispersion around an unknown centre, and the same constraint applies to any quantitation attempted without a matched standard.8
Fourth, blind submission tests a laboratory’s ordinary process, which is the point, but it also means we bought the cheapest standard product from each service rather than the most thorough. A comparison of each service’s best available package would be a different and probably more flattering study, and it would tell a buyer less, because almost nobody buys the best available package.
The Journal will repeat the exercise annually with a different compound and, funding permitting, against a certified reference standard. The design is published so that others can run it.
Four organisations carry the analytical burden for a trade with no release-testing obligation of its own. That is the structural fact underneath everything in this department, and it means the independent record is as deep as four commercial queues allow it to be.
| Service | Primary function | Assay work | Public archive | Advertises in this publication |
|---|---|---|---|---|
| Janoshik Analytical | Assay laboratory (Czech Republic) | Purity, identity, peptide content | No | Yes — disclosed |
| Medutest | Assay and documentation verification | Purity, identity | Partial | No |
| PeptideMeter | Assay laboratory with batch archive | Purity, content | Yes, by vendor and lot | Yes — disclosed |
| VendorInvestigate | Vendor and documentation audit | Commissioned, not in-house | Findings published | No |
| Compiled from each service’s published description of its offering and from the Journal’s own submissions and correspondence. All four were sent this table before publication and invited to correct it; three responded and two corrections were made. The advertising column is repeated in the text and on our funding page. | ||||
Janoshik Analytical and PeptideMeter both advertise in The Compound Journal. Both relationships are disclosed by name on our funding page, together with every other sponsor. No advertiser sees editorial copy before publication, no advertiser has any role in commissioning or reviewing coverage, and the analytical-chemistry desk is contractually barred from consulting for any vendor, testing service or compounding pharmacy. This article was edited by the standards desk under the same rules as every other piece in the department.
The Journal also pays these services. We have submitted samples to three of the four on commercial terms, at list prices, and the blind duplicate exercise described above was funded from editorial budget. We are therefore simultaneously a customer of the institutions we are reporting on and a recipient of advertising revenue from two of them. Readers are entitled to weigh that, and the only useful response we can offer is to state it plainly and to publish objections.
Our position on the substance is unchanged by any of it. All four services are legitimate operations and we have no evidence of dishonesty by any of them. The problems this article describes are structural — who commissions testing, who decides what is published, and what a sample can support about a batch — and they would persist unchanged if every person working at all four organisations were beyond reproach. Correspondence to standards@compoundjournal.com.
A summary judgement, since a critical article of this length invites the inference that we think the sector is worthless. We do not. Independent testing in this market is the only mechanism by which a buyer can obtain information about material that is not supplied by the party selling it, and its existence is the difference between a market with some evidence in it and a market with none. Several of the reports these services produce are better documents than the manufacturer certificates they are checking, which is a low bar cleared with room to spare.
The three criticisms we would press are narrow. A sample is not a batch, and the trade cites samples as batches. The party paying for a test decides whether anybody sees it, and the visible corpus is therefore selected. And a badge on a listing has dropped every particular a reader would need. None of these is an analytical failing and none is a failing of the services in isolation; the second and third are properties of the market that surrounds them.
What we would tell a reader is this. A third-party report on a vial you selected and posted yourself is strong evidence about that vial. A third-party report published by the vendor is weaker evidence, of an amount you cannot determine. A badge is not evidence. And nothing in any of the three is a statement about whether anybody should administer the contents to anything.
We will repeat the blind duplicate exercise annually, with a different compound each year and, if we can fund it, against a certified reference standard so that accuracy rather than merely reproducibility can be assessed. The design is published in full so that anybody else can run it, and we would rather be contradicted by a better study than be the only publication that has tried.
Selected from correspondence received on this article. Writers are identified by initial, surname and city, verified before printing. Replies are from the desk that filed the piece or from the standards editor. Write to letters@compoundjournal.com.
The strongest blind design I have seen anywhere involved a third party holding the purchase record, the laboratory receiving numbered tubes and the mapping opened only after all results were in. It is not expensive. It is simply organisation, and organisation is what this market lacks.
— S. Bergqvist, Malmö
Custody held by somebody with no stake in the answer is the whole of that design, and you are right that the barrier is organisational. Nothing in it requires money.
The spread you reported is, by the standards of interlaboratory comparison in sectors I have worked in, unremarkable. That is not a criticism of your reporting; it is a correction to the alarm some readers will take from it. Two laboratories differing by a point on a chromatographic purity are behaving normally.
— T. Abubakar, Kano
Normal, and still consequential when the difference decides whether a lot meets a stated specification. Both things are true and the piece should have held them together more firmly.
On the archive point: a public record of submissions by vendor would be gamed within a month. Vendors would submit under the names of resellers, or through intermediaries, and the archive would show a distribution as selected as the current one but with a veneer of completeness.
— E. Nkomo, Polokwane
Probably true in part, and it is the strongest argument against our proposal. Our answer is that gaming requires effort and leaves traces, which the present arrangement does not, and that a partially gamed record is more informative than no record. We would not claim more than that.
Placement carries meaning. A mark beside a price reads as a claim about that product; the same mark in a footer reads as a claim about the company. Sellers know this and choose accordingly, and no rule anywhere governs the choice.
— P. Ahluwalia, Chandigarh
A note on scale. One of these services runs volumes that would be respectable for a contract laboratory in a regulated sector; another is a small operation with a short queue. Both issue documents that look alike and carry the same weight in a listing. The report format equalises what the operations do not.
— G. Vermeulen, Antwerp
It does, and it is why we now record the service alongside every determination in the dossier tables rather than reporting a single independent figure.
Where two methods disagree, the conservative convention is to report the lower figure. It is not universal, and whether a laboratory follows it belongs on the report.
Documentation practice is the only part of vendor quality a buyer can assess before purchase.
The result is unremarkable. What the report states alongside it is not.
Reported from the analysis, not from a warning notice.
The Journal submitted split samples from single lots to three assay services, under names unconnected to this publication, and published each method alongside each result.
Follow the resin, not the catalogue.