Method and context

How we grade evidence

The four grades used across this publication, what moves an ingredient between them, what we refuse to grade and why we never publish a number we cannot defend.

Reviewed 2026-07-31· Published by Northbank Media
The short answer

We grade the strength of the evidence base, not the popularity of the ingredient. Every claim is placed in one of four grades: Strong, Moderate, Limited or Insufficient. A grade describes how confident anyone should be that the claim holds in people, based on the depth, independence and consistency of the human research behind it. It is never a score, a rating or a rank, and no company can influence one.

1. Why the grades are words and not numbers

A great deal of health publishing gives ingredients a score out of ten or five stars. Those numbers look precise and are almost always arbitrary. They combine incommensurable things, they imply a resolution the underlying evidence does not have, and they are extremely convenient for anyone who wants to sell an ingredient with a good one.

We use four named grades instead, and we accept the loss of apparent precision as a gain in truthfulness. An evidence base is not a quantity. It has a shape: how many controlled human trials broadly exist, how large and how long they typically are, who funded them, who has replicated them, and what is conspicuously missing. A grade is a summary of that shape.

2. The four grades

Strong. Multiple independent controlled trials in people, consistent in the direction of effect, with a mechanism that holds up and, critically, replication by parties with no commercial interest in the result. Where a claim has additionally survived regulatory scrutiny as part of a licensed medicine, that counts. We apply Strong sparingly. At the time of writing, one review on this site carries it.

Moderate. Several controlled human trials exist, and they are typically small, short or narrow in who was studied. The direction of effect is reasonably consistent. The size and durability of it are not settled, and independent replication is patchy. Most useful ingredients live here.

Limited. Human evidence exists but is thin: small, short, single site, dominated by parties with a commercial interest, or leaning heavily on laboratory and animal work to fill the gaps. A Limited grade is not a statement that something does not work. It is a statement that nobody currently knows.

Insufficient. Little or no controlled human evidence for the claim as it is being made. The idea may be biologically reasonable and still be untested at the level being sold. This grade appears most often on individual claims within reviews of ingredients that carry a higher overall grade.

3. We grade claims, not ingredients

This is the most important structural decision on the site, and it is why every review opens with a claim-versus-evidence table rather than with a single verdict.

An ingredient does not have one evidence base. It has one per claim. Zinc is essential for wound repair, which is settled physiology, and the claim that extra zinc speeds healing in people who are already replete is unsupported. Both statements are about zinc. Collapsing them into one score for zinc destroys the only information the reader needed.

So each review carries an overall grade for orientation and a table that grades the individual claims separately. When those diverge sharply, the table is the part worth reading.

4. What moves a grade up

  • Independent replication. The single most powerful factor. A finding reproduced by a group with no stake in it is worth more than several findings from the party selling the ingredient.
  • Adequate size and duration. Trials long enough to speak to the timescale on which the claimed benefit would matter. For skin ageing that means months at an absolute minimum, and years to be convincing.
  • Outcomes a person would notice. A change in an instrument reading is not the same as a change a participant can perceive. We weight perceptible outcomes more heavily than surrogate markers.
  • Adoption into clinical guidance. Guidance bodies such as NICE review trial bases rather than marketing. Appearance in mainstream guidance is meaningful corroboration.
  • Consistency across populations. An effect that holds in different groups is more likely to be real than one that appears only in a narrowly selected sample.

5. What holds a grade down

  • Funding concentration. When a large share of research in a field comes from the parties selling the ingredient, the field has not been stress tested by anyone with a reason to find nothing.
  • Laboratory evidence used as a substitute. Cell and animal work describes mechanism. It does not demonstrate benefit in people, and a large volume of it never lifts a grade on its own.
  • Surrogate outcomes only. Instrument readings and biochemical markers are proxies. Proxies fail often enough that a body of evidence built entirely on them stays capped.
  • Formulation confounding. A trial testing a finished commercial product cannot attribute any result to one ingredient in it.
  • Publication bias. Positive results are more likely to be published, and more so when the funder has an interest. A field where every published study agrees is a field to be suspicious of.

6. What we will not do

We will not invent a study, an author, a journal, a sample size, a p-value or an effect size. If you find a specific statistic on this site attributed to a specific paper, we have made a mistake and want to hear about it. Everything here describes the shape of an evidence base qualitatively, because that is what can be stated without misrepresenting the literature.

We will not review or grade named commercial products or brands. Grades attach to ingredients and claims. This removes the mechanism by which a review site becomes a shop.

We will not accept payment, samples, sponsorship or affiliate commission in exchange for coverage or for a grade. No company previews a grade before publication.

We will not soften a grade because an ingredient is popular. Several ingredients on this site with substantial markets behind them carry Limited grades, and they will continue to do so until the evidence changes.

7. Revisiting grades

A grade is a description of a moment. Evidence bases change, and ours are dated on the page so you can see when they were last reviewed. We revisit a grade when adequately sized independent trials appear, when a systematic review substantially changes the picture, when a regulator or guidance body acts, or when a reader shows us we were wrong.

We would like to raise grades. A publication that only ever downgrades is as unreliable as one that only ever promotes. What we will not do is raise a grade in anticipation of evidence arriving.

8. Corrections

If a grade on this site is wrong, tell us. Contact the editors with the page and the reason. Substantive corrections are made on the page and noted. Our editorial policy sets out independence, funding and the single disclosed exception that applies to four archive articles.

No commercial links on this page

This article contains no affiliate links, no sponsored placements and no links to any commercial product, brand, retailer or clinic. Nobody paid for it, nobody previewed it and nobody can have a grade changed. Our editorial policy sets out the single disclosed exception that applies to four archive articles, none of which is this one.

Nothing here is medical advice. Speak to a pharmacist, a GP or a dermatologist about your own circumstances.

Sources

Institution level references. We link to bodies that publish their methods, not to retailers. External links open on those bodies' own sites.

  • Cochrane LibraryThe methods underpinning systematic review and evidence synthesis, including risk of bias. www.cochranelibrary.com
  • NICE: guidanceHow UK clinical guidance is developed from underlying trial evidence. www.nice.org.uk
  • EFSA: health claimsHow claims about foods and supplements are assessed before they may be made. www.efsa.europa.eu
  • PubMedThe primary biomedical literature, and the place to check anything stated here. pubmed.ncbi.nlm.nih.gov

Frequently asked questions

Why do you not score ingredients out of ten?

Because a number implies a precision the evidence does not have and combines things that cannot be combined. Four named grades describe the shape of an evidence base without pretending to measure it.

Can a company pay to change a grade?

No. No company pays for coverage, previews a grade, supplies samples we accept, or has any route to influence one. We take no affiliate commission and run no sponsored content.

Why does one ingredient have several different grades?

Because we grade claims rather than ingredients. An ingredient can be essential in deficiency, which is Strong, and useless as a general supplement, which is Insufficient. Both are true and collapsing them loses the point.

Does a Limited grade mean the ingredient does not work?

No. It means nobody currently knows, because the human research is too thin, too short or too commercially concentrated to support a conclusion. Absence of evidence is not evidence of absence, and it is also not a reason to buy something.

How often do you revisit grades?

Whenever the evidence moves: new independent trials of adequate size, a systematic review that changes the picture, or regulatory action. Every page carries the date it was last reviewed.

Read next

Get the method updates

We email when the grading method changes or a new method page is published.

We email when a grade changes or a new review is published. Nothing else, and no product recommendations, because we do not make any.