
The levels of scientific evidence rank studies by how well they can establish causation: anecdote and the cell model at the base, the randomized controlled trial and the meta-analysis at the top. A spectacular finding at a low level carries less weight than a modest result at a high one.
The pyramid, from the bottom up
| Level | Type of study | What it can support |
|---|---|---|
| 1 (base) | Opinion, anecdote, testimonial | Nothing generalizable |
| 2 | In vitro study (cells) | That something happens in an isolated system |
| 3 | Animal study | That something happens in that species, in that model |
| 4 | Case series | That it happened in a few individuals, with no comparison |
| 5 | Observational study | Association, never causation |
| 6 | Randomized controlled trial | Causation, in the population studied |
| 7 (top) | Systematic review and meta-analysis | A synthesis of every available trial |
The rule that sums up the whole pyramid: the higher you go, the less room there is for the result to be explained by something other than what was studied.
A common mistake is to treat the levels of scientific evidence as a matter of prestige. They aren't. They are a matter of which competing explanations have been ruled out. An observational study is not worse science than a trial; it simply leaves more alternative explanations open.
Cell models, animal models, and why almost nothing translates
Most of the literature on research peptides sits at levels 2 and 3. It pays to understand what that means.
In vitro means cells in a dish. The concentration of compound applied is usually far higher than anything living tissue would reach, there is no liver metabolism, no intestinal barrier, no immune system, and none of the other forty things that happen inside an organism. An in vitro effect is a hypothesis, not a finding.
In animals there is an organism, but three gaps still separate the result from a person: the species is different, the disease model is induced artificially and does not necessarily resemble human disease, and doses are usually scaled by body weight, which is a debatable simplification.
The rate at which promising preclinical findings become approved drugs is low across every therapeutic area. This is not peculiar to peptides; it is how drug development works. Most molecules that look exciting in an animal model go nowhere, and that does not mean the research was done badly.
The four phases of clinical research
| Phase | Question it answers | Typical participants |
|---|---|---|
| Phase 1 | Is it tolerable? What does the body do to it? | Dozens, often healthy |
| Phase 2 | Does it do anything? At what dose? | Hundreds, with the condition |
| Phase 3 | Does it work against a comparator, in a large population? | Hundreds to thousands |
| Phase 4 | What shows up with widespread, long-term use? | After approval |
Two clarifications that change how headlines read:
Phase 2 confirms nothing. It is built to explore dosing and detect a signal, not to demonstrate efficacy. A spectacular phase 2 result often shrinks in phase 3, and this happens often enough that the methodology literature has a name for it.
Phase 4 is where rare effects turn up. An adverse effect that occurs in one in ten thousand people will not be detected in a trial of two thousand. It is detected when millions use it.
Topline results vs. peer-reviewed publication
A topline is a statement from the trial sponsor with the headline results. It is legitimate information and often accurate, but it has two structural limits: it has not been through independent review and the sponsor decides what to highlight.
A peer-reviewed publication includes full methods, secondary analyses, detailed adverse events, and the limitations reviewers required the authors to acknowledge. It is where you see what a press release leaves out.
The difference is not academic. When this site cites preliminary data, it says so explicitly, precisely because a number can be qualified in the publication that follows.
The seven red flags
1. Small sample. A trial with twelve participants can suggest; it cannot prove. The smaller the sample, the more likely the result is noise.
2. No control group. Without a comparator there is no way to know what would have happened without the intervention. Many conditions improve on their own, and the placebo effect is real and measurable.
3. No blinding. If participants or assessors know who is getting what, expectation contaminates the measurement, especially for subjective outcomes like pain.
4. Surrogate endpoint. A marker is measured instead of the outcome that matters. It gets its own section below.
5. Undeclared conflict of interest. It does not invalidate a study, but it is information the reader needs in order to weigh it.
6. No replication. A single result is a hypothesis until an independent group reproduces it.
7. The result does not match the prior registration. Trials are registered before they begin, declaring what they will measure. When the published paper reports an outcome different from the registered one, ask what happened to the original. It is called outcome switching, and it is more common than it looks.
Surrogate endpoint vs. clinical outcome
This distinction is probably the most useful thing in the whole article.
A clinical outcome is something a person cares about directly: living longer, having fewer heart attacks, regaining function.
A surrogate endpoint is a marker assumed to predict that outcome: cholesterol, blood pressure, bone density, a lab value.
Surrogates are appealing because they are fast and cheap to measure. The problem is that improving the marker does not guarantee improving the outcome, and the history of medicine is full of cases where the marker improved and the clinical result got worse.
A concrete, current example sits in the phase 3 readouts for retatrutide: TRIUMPH-3 markedly improved several cardiovascular risk markers (triglycerides, blood pressure, hsCRP), and yet the major adverse cardiovascular event composites did not reach statistical significance. Better markers, unproven events. They are two different things, and confusing them is the costliest mistake a reader can make with a headline.
How we read evidence here
Our editorial policy defines the hierarchy of levels of scientific evidence we apply, and its three operating rules are:
The level is stated in the text, not in a footnote. When a finding comes from an animal model, the sentence says so. Nothing appears as simply "promising."
No number without trial, population, and time point. A percentage that doesn't say which trial, in which population, and at which week is not data; it is an impression with decimals.
What isn't settled gets its own section. Every article includes what remains open, and that section is mandatory.
None of this turns the content into scientific literature. It is educational content with its sources on display, which is a more modest category and, we think, a more honest one.
Frequently asked questions
Is a mouse study useless?
Far from it: it is how the hypotheses later tested in people get generated. What it cannot do is establish that something will work in humans. It is a starting point, not a conclusion.
What sample size is enough?
It depends on the effect being looked for: a large effect can be detected with few participants; a small one needs many. That is why trials calculate their size in advance. A study that doesn't report that calculation is leaving out relevant information.
Is a meta-analysis always the best evidence?
It sits at the top of the pyramid, but it inherits the quality of what it pools. A meta-analysis of bad studies produces a bad conclusion with more decimal places. The methodology literature sums it up bluntly: garbage in, garbage out.
Why does a phase 2 result shrink in phase 3?
Several reasons combine: phase 2 trials tend to use more selected populations, early effects tend to regress to the mean on replication, and a phase 3 measures under conditions closer to real practice.
References
- Guyatt GH, et al. GRADE: an emerging consensus on rating quality of evidence and strength of recommendations. BMJ, 2008;336:924–926. DOI: 10.1136/bmj.39489.470347.AD
- Ioannidis JPA. Why Most Published Research Findings Are False. PLoS Medicine, 2005;2(8):e124. DOI: 10.1371/journal.pmed.0020124
- Perel P, et al. Comparison of treatment effects between animal experiments and clinical trials: systematic review. BMJ, 2007;334:197. DOI: 10.1136/bmj.39048.407928.BE
- Prasad V, et al. The strength of association between surrogate end points and survival in oncology. JAMA Internal Medicine, 2015;175(8):1389–1398. DOI: 10.1001/jamainternmed.2015.2829
- Goldacre B, et al. COMPare: a prospective cohort study correcting and monitoring 58 misreported trials in real time. Trials, 2019;20:118. DOI: 10.1186/s13063-019-3173-2
Written by the Bionic Editorial Team. Last reviewed: August 2026.
How we work: our editorial policy and source hierarchy.
This content is strictly educational and does not constitute medical advice, diagnosis or a therapeutic recommendation. The compounds mentioned are research products (Research Use Only) and are not approved by INVIMA, FDA, EMA or ANSM for therapeutic use in humans. Any health-related decision should be made with a licensed medical professional.