Littermate
A long laboratory bench seen end-on, glassware and racks arranged along its length under even overhead light.
A teaching laboratory bench. The rat tendon studies discussed here were run in rooms like this one, and almost every methodological question about them concerns what happened in the room rather than what the animals were given. Photograph: Milda 444 / Wikimedia Commons (CC BY-SA 4.0)
Vol. 6

One lab, many tendons

Most of what has been published about BPC-157 in rat tendon and ligament models traces back to a small number of connected groups. That is not evidence of misconduct. It is a structural fact about the literature, and it changes what any single reported effect is worth.

Fact-checked by Danny Whitcombe

There is a moment in a rat Achilles tendon study that decides most of the result, and it is not the moment the compound is given. It is the moment, days or weeks later, when someone sits at a microscope with a tray of stained sections and assigns each one a number. Fibre alignment, cellularity, vascularity, collagen maturity: a composite score, out of ten or twelve, produced by a person looking at a picture. If that person knows which sections came from treated animals, the score moves. Not because anyone is dishonest. Because that is what happens to human judgement under a known condition.

This is the ordinary machinery of preclinical research and it applies to every compound ever studied this way. It applies with particular force to BPC-157, a synthetic fifteen-residue peptide whose rodent literature is unusually large, unusually consistent in direction, and unusually concentrated in its authorship.1

The shape of the problem

Search the published literature for this peptide and the return is substantial: dozens of experimental papers across three decades, spanning ethanol-induced gastric lesion models, chemically induced colitis, Achilles tendon and medial collateral ligament transection, quadriceps crush injury, and vascular anastomosis work. Nearly all of it is in rats. The direction of effect is remarkably stable across the whole corpus. Treated animals heal faster, score better on histology, and in the biomechanical studies tolerate more load before the repaired tissue fails.

A reader encountering that volume naturally reads it as weight of evidence, because dozens of papers pointing the same way is the shape a well-supported finding takes.

It is also the shape a single productive research programme takes. Telling those two things apart is the entire job.

Most of this corpus traces back to one network of investigators centred on a group in Zagreb, publishing continuously from the 1990s onward, together with collaborators and former trainees.2 That is not an accusation and it is not unusual. Productive groups own their subjects; that is how research fields get built at all. But a body of work from one network is one body of work. Its papers share a protocol lineage, an animal supplier, a surgical technique, a scoring rubric, and a set of assumptions about what counts as a result. Errors anywhere in that chain propagate across every paper in it rather than cancelling out between them.

What the rodent work actually shows

Take the tendon literature specifically, since it is the part most often cited outside the field.

The standard design is a transection model. The Achilles tendon of an anaesthetised Sprague-Dawley or Wistar rat is cut, sometimes sutured, and the animal is allowed to heal while receiving either the peptide or a control. At set intervals the tendons are harvested. Some go to histology for the scoring described above. Some go to a materials testing machine, which pulls until the tissue fails and records the load at which it did.

The biomechanical endpoint is the better of the two, because a load-to-failure number comes off an instrument rather than out of a judgement. Reported differences on that measure have been real and in some studies sizeable.3

It carries its own problems, though. Load to failure depends on the cross-sectional area of the specimen, on how the tissue was gripped in the machine, on the rate at which the machine pulled, and on how long the specimen sat between harvest and testing. Those parameters differ between laboratories and are frequently reported incompletely or not at all. Two groups can run what looks like the same experiment and produce numbers that are not comparable to each other.

And the histological score, the weaker of the two measures, is the one that appears in more of the papers.

An engraved technical plate showing a nineteenth-century freezing microtome from two angles.
A freezing microtome, drawn in 1873. The instrument for cutting tissue thin enough to score has been in laboratories for a hundred and fifty years. The habit of scoring those sections unblinded has been leaving the field rather more slowly. Photograph: Wellcome Collection / Wikimedia Commons (CC BY 4.0)

Why single-lab dominance matters

The reason independent replication carries so much weight is not that outside laboratories are more honest. It is that they are differently wrong.

Every laboratory has habits that are invisible from inside it. Which rat strain, from which supplier, at what age. How long the animals acclimatise before surgery. Whether the surgeon has done four hundred of these or forty. What the room temperature is, and the light cycle. Whether the control animals were handled as often as the treated ones, since handling alone changes rodent physiology in measurable ways. None of that appears in a methods section, and any of it can generate an apparent effect that has nothing to do with the compound.

When a second, unconnected group runs the same experiment, its idiosyncrasies are different ones. If the effect survives both sets of habits, it is more likely to be about the molecule than about the room. If it does not survive, that is information too, and it is information the original group cannot generate no matter how many further studies it runs. This is the whole mechanism by which a reported finding becomes a known fact, and it is not optional.

The BPC-157 literature has not been through that process at anything close to the scale its citation count implies. There is a scattering of work from groups with no connection to the originating network, and it is thin relative to the volume of the primary corpus.4 For a compound discussed as widely as this one is, that is the single most important fact about it, and it is almost never the fact that gets repeated.

What independent replication would require

It is worth being concrete about what would settle this, because more research is needed is a sentence that can be written about anything and commits nobody to anything.

A useful replication would be preregistered and multi-laboratory. Preregistered, so that the primary endpoint is fixed before the first animal is operated on and cannot migrate afterwards to whichever measure happened to move. Multi-laboratory, so that the result does not belong to one room and one pair of hands.

It would need randomised allocation with the sequence concealed from the surgeon, and histological scoring done by assessors who do not know the group assignment and ideally do not know the hypothesis. It would need a sample size chosen from a power calculation rather than from convention. A great deal of this literature runs on eight animals per arm, which is adequate for detecting a very large effect and close to useless for detecting a moderate one.5

It would need a prespecified analysis plan. And it would need to publish what it found, including nothing, which is the requirement that fails most often.

None of this is exotic. It is the standard the preclinical field has been moving toward for the better part of two decades, and reporting guidelines for animal research exist precisely to make it routine.6 It simply has not been applied to this compound.

Bound ledgers shelved tightly together, spines worn and unlabelled.
Bound volumes. Counting papers is not the same as counting confirmations, and a citation trail that never leaves one network can be very long without ever becoming independent. Photograph: Luke McKernan / Wikimedia Commons (CC BY-SA 2.0)

Where this leaves a reader

Not at the conclusion that the compound does nothing. That conclusion is not available from this evidence either. Under-replicated is not the same as refuted, and it would be its own kind of overreach to write as though it were.

Where it leaves a reader is here. There is a coherent, internally consistent body of rodent work reporting that a synthetic peptide improves healing metrics in injured rats. It was produced largely by one research network over three decades. It carries the methodological weaknesses common to preclinical work of that period: small groups, blinding often unstated, endpoints that vary between studies. And it has not had the independent, multi-site confirmation that would turn it from a finding about a research programme into a finding about biology.

There are no completed, published randomized controlled trials in humans. There is no established human safety dataset. Every question a person might actually have about this compound sits downstream of experiments nobody has run.

That is an unsatisfying place to stop, and it is where the evidence stops. The alternative, which is to round a large rodent literature up into a claim about people on the grounds that it is large, is how a great deal of confident nonsense enters circulation and stays there. A mouse that heals is a fact about a mouse. A rat tendon that fails at a higher load is a fact about a rat tendon. The distance between those facts and anything that matters to a person is not a gap in the writing. It is a gap in the research, and it closes the same way it opened: one experiment at a time, run by people with nothing riding on the answer.

Apparatus

References and further reading

We describe sources rather than manufacture citations. Where we have not read a specific paper, we say what body of work we are pointing at instead of inventing a reference for it.

  1. The preclinical corpus for this peptide, as described in review articles covering rat gastrointestinal and musculoskeletal injury models published between the late 1990s and the 2020s. Return
  2. Authorship concentration, established by reading the author lists on the primary experimental papers rather than on the reviews. The recurring senior authorship is associated with a Zagreb-based group, its collaborators and its former trainees. Return
  3. Rat Achilles tendon transection studies reporting biomechanical load-to-failure endpoints, which appear in the experimental surgery and orthopaedic research literature from the 2000s onward. Return
  4. Work from unaffiliated groups. A small number of studies in rodent musculoskeletal and gastrointestinal models exist outside the primary network. We have not located an independent multi-site replication of the tendon findings. Return
  5. The relationship between group size and detectable effect is a general statistical result, not a claim about this compound. Any standard treatment of statistical power for two-group comparisons covers it. Return
  6. Reporting standards for animal research. The ARRIVE guidelines set out what a preclinical paper should report about randomisation, blinding, sample size and outcome measures. Return

Theo Vance

Research Editor

Left a molecular biology doctorate after a third failed attempt to reproduce a result he had been told was solid. He now writes almost entirely about why replications fail: the parts of a protocol that never make it into the methods section, the strains that are not the strains, the lab that is the only lab.

Corrections to this piece: corrections@littermate.review

Elsewhere in this issue