How we weigh the evidence
Every claim we assess carries the same two marks: how strong the evidence is, and which way it points. This is the method behind them, and the reason we sometimes say "we don't know."
Ask two people whether a habit is "good for you" and you will often get two confident, opposite answers. The problem is rarely the facts. It is that the word good hides two different questions: how sure are we? and how much does it matter? Almost every health argument collapses those two into one. We keep them apart on purpose.
Two marks, not one
Every review and verdict we publish carries two signals.
The first is certainty, our confidence that the real effect is close to what the studies report. We grade it on four levels: high, moderate, low, very low. High certainty means many well-conducted studies point the same way and further research is unlikely to overturn the conclusion. Very low means the evidence is thin, inconsistent, or indirect, and an honest reader should hold the claim loosely.
The second is effect, the direction and rough size of what was measured: a benefit, a harm, no meaningful difference, or mixed, when it depends on who, how much, or for how long.
Separating them matters because the two do not move together. There is strong evidence for small effects and weak evidence for dramatic ones. A headline usually sells you the drama and hides the certainty. We show both, side by side, so you can tell them apart in a glance.
Where certainty comes from
We start from the shape of the evidence, not a single paper. A claim supported by several consistent randomised trials, or by a careful systematic review that pools them, starts high. A claim resting on one small study, on animal work, or on a survey that can only show correlation, starts lower.
Then we look for reasons to trust it less: studies that disagree with each other, results measured on a stand-in rather than the outcome you care about, samples too small to be sure, and signs that unflattering results never got published. Each of those pulls the certainty down. None of it is about whether the finding is exciting. It is about whether it is likely to survive.
Why we sometimes say "we don't know"
The most useful thing an evidence desk can do is refuse to round up. When the honest answer is very low certainty, we say so plainly, even when a confident answer would read better. An inflated verdict feels good until the day someone acts on it. A calibrated one is duller and far more useful.
What this is, and isn't
This is an evidence summary, written for the general reader. It is not personal medical advice, and it cannot account for your particular situation, that is a conversation for you and a clinician who knows your history. What we can do is tell you, without spin, how much the science actually supports a claim, and how far.
That is the whole promise of Dr Schwartz: the two marks, honestly assigned, every time.
This is an evidence summary for the general reader — not a diagnosis or personalised medical advice. For any health decision, talk to a qualified professional.
More stories
The evidence pyramid: why not all studies are equal
"A study found" tells you almost nothing until you know what kind of study it was. A quick tour from anecdote to meta-analysis.
Relative vs absolute risk: the trick behind scary headlines
"Doubles your risk" and "raises it from 1 in 10,000 to 2 in 10,000" can describe the exact same finding. One sells papers; the other tells you what to do.
Does vitamin C stop you catching colds?
Decades of Cochrane trials find no meaningful drop in cold risk for ordinary adults, only a modest shortening of colds once they start.