Likelihood Ratios: How Much Should This Finding Change Your Mind?
A practical guide to one of the key ideas in Bayesian updating—and one of the most useful tools in diagnostic reasoning.
A patient has some finding—a symptom, an exam finding, a lab result, an imaging result—and you want to know how concerned or reassured you should be.
Every Bayesian update needs two ingredients
First, you need a starting point: the prior probability, or in clinical language, the pretest probability. This is your best estimate before the new result arrives. It may come from prevalence, the history, the exam, your experience, or some combination of all four.
The second is the data you obtain to either support or refute the original hypothesis. The strength of that data is best quantified as a likelihood ratio. The prior tells you where you are starting. The LR tells you how far the new evidence should move you. Put them together and you get the post-test probability.
The same test result can land very differently in two patients who started with different pretest probabilities. The evidence has not changed, but the starting point has—and that changes where they end up.
So what is a likelihood ratio?
I think of it this way: how much more at home is this result in one group than the other?
An LR above 1 favors the disease. An LR below 1 argues against it. An LR near 1 barely changes the odds because the result occurs at roughly the same rate in both groups. So, the farther an LR moves from 1—in either direction—the more weight that finding carries.
Where the LR comes from
Let’s start with the two familiar test characteristics.
Sensitivity is the true-positive rate. Among people who have the disease, how often is the result positive?
Specificity is the true-negative rate. Among people who do not have the disease, how often is the result negative?
That gives us the two rates on the other side of the table as well. The false-positive rate is 1 minus specificity. The false-negative rate is 1 minus sensitivity.
Interactive LR lab
Build the ratio from the rates
Choose a test or finding, then follow the same result through two equal reference groups and into the likelihood-ratio equation.
1 · Start with published performance
- Result counted as positive
- Cough present
- Published performance
- 92% sensitivity · 22% specificity
- Studied population
- Symptomatic respiratory infection or suspected influenza across 72 inpatient and outpatient studies and multiple age groups.
- Source and limitation
- Ebell et al., 2025 systematic review and meta-analysis. The pooled population and the definition and measurement of cough varied among studies.
2 · Separate two equal groups
These are two separate reference groups, not one population with 50% disease prevalence. Equal groups make the conditional rates easy to see.
100 people with the disease
Sensitivity determines how this group separates.
- true positives
- about 92
- false negatives
- about 8
100 people without the disease
Specificity determines how this group separates.
- false positives
- about 78
- true negatives
- about 22
3 · Turn the rates into likelihood ratios
LR+
92% ÷ 78% = 1.2
A positive result is about 1.2 times as likely in someone with influenza as in someone without it.
LR−
8% ÷ 22% = 0.36
A negative result occurs about 0.36 times as often in someone with influenza as in someone without it.
4 · Apply this LR to a starting probability
Cough when evaluating influenza
Keep the published performance above fixed. Change the starting probability, choose the result you observed, and predict how far it will move you.
Use the slider or enter a value from 5% to 95%.
Optional. Enter a value from 0% to 100%. This is a prediction, not a scored answer.
The starting probability is a teaching assumption, not a prevalence estimate for every patient. This exercise does not recommend testing, treatment, or disposition for an individual patient.
Now the formulas make more sense:
LR+ = sensitivity ÷ (1 − specificity)
LR− = (1 − sensitivity) ÷ specificity
One important boundary on the influenza option: the Sofia example is an antigen RIDT, not a rapid molecular assay. Antigen RIDTs are less sensitive than molecular influenza assays, and a negative antigen result does not independently exclude influenza when clinical suspicion remains. Test performance also depends on the assay, specimen, collection timing and quality, patient population, and reference standard. See the current CDC guidance on rapid influenza diagnostic tests for clinical interpretation.
LR+ compares positive results with positive results: the true-positive rate divided by the false-positive rate.
LR− compares negative results with negative results: the false-negative rate divided by the true-negative rate.
This is why LR+ and LR− from the same test can be very different. A test can be much better at confirming a diagnosis when positive than excluding it when negative—or the reverse.
The easy mistake: counts instead of rates
Here is a common near-miss: “LR+ is true positives divided by false positives.” The missing word is rates. Raw patient counts answer a different question, and they change when the mix of patients changes.
Same test · different patient mix
The rate stays put. The raw-count ratio does not.
1 · With disease
100 people who have the disease
- 90 TP
- 10 FN
True-positive rate90 ÷ 100 = 90%
2 · Without disease
100 people who do not have the disease
- 20 FP
- 80 TN
False-positive rate20 ÷ 100 = 20%
3 · Random sample
100 people: 10 with disease, 90 without
- 9 TP
- 1 FN
- 18 FP
- 72 TN
TPR 9 ÷ 10 = 90%
FPR 18 ÷ 90 = 20%
LR+ = 90% ÷ 20% = 4.5
The random sample contains 9 true positives and 18 false positives, but 9 ÷ 18 is a ratio of raw counts—not the LR. The rates stayed the same; only the patient mix changed.
That distinction matters in the ED because we are not testing random samples of the population. By the time a test is ordered, selection has already happened through the chief complaint, history, exam, epidemiology, and clinical setting. Sometimes that leaves us with a fairly high starting probability. Sometimes we deliberately test at a low starting probability because the test is part of a rule-out strategy. Either way, the patients being tested are a selected group in whom the disease is plausible enough to investigate.
As that starting probability changes, so does the number of true positives relative to false positives in the tested cohort. That is the territory of predictive value. The LR is doing a different job: it describes the weight of the observed result by comparing its rate in people with the disease with its rate in people without it.
This does not make an LR universal. Test performance can still shift with the assay, threshold, spectrum of illness, setting, and study methods. The point is narrower: do not let the disease mix in one cohort masquerade as the evidentiary weight of the test itself.
A Bayes factor in clinical clothing
If you have encountered Bayes factors elsewhere, this is the same basic move wearing clinical clothes.
A Bayes factor asks how likely the observed evidence is under one hypothesis compared with another. In a binary diagnostic question, those hypotheses are disease and no disease. The likelihood ratio for the result you observed tells you how the evidence should change the odds between them.
The LR is not the final probability. It is the multiplier that helps you get there.
Bayes factors are used far beyond diagnostic testing, so the terms are not interchangeable in every setting. But in this specific setting, the mathematical role is the same.
Same evidence, different starting point
Let’s go back to our LR+ of 4.5 and apply it at three different starting probabilities.
Start at 5%, and a positive result raises the probability to about 19%.
Start at 20%, and the same result raises it to about 53%.
Start at 50%, and it raises it to about 82%.
Same test. Same result. Same LR. Different destination.
This is why asking whether a test is “good” is usually an incomplete question. Good enough to do what, and starting from where?
What the LR cannot decide for you
A finding can move probability a lot without settling the clinical decision. Another finding can move it only a little and still matter because you started close to a decision threshold.
The LR does not choose that threshold. It does not tell you how serious a missed diagnosis would be, what harms come with more testing, or what the patient values. It answers one narrower question: how should this result change the odds?
That leaves three questions worth keeping separate:
- Where was I before the result?
- How far should this result move me?
- What should I do with the updated probability?
The math is most helpful when it clarifies the second question without pretending to answer the other two.
Making the decision explicit—and putting a number on your prediction—also gives you something you can revisit. Over time, that feedback can help you make better decisions. The SOLVED Perspective on making your predictions explicit is a useful next step.
A caution about stacking evidence
Once you get comfortable with LRs, it is tempting to start multiplying them together: one for the history, one for the exam, one for the lab, one for the imaging. This can be useful, but only if the findings are contributing reasonably independent information.
Often they are not. Fever and chills may be telling you much the same thing. Several components of a decision rule may arise from the same physiologic process. A test result may already be partly reflected in the way you formed the prior.
Before multiplying LRs
Ask what is actually new.
Three findings, one underlying signal
- Fever
- Chills
- Tachycardia
A shared inflammatory response
Multiplying all three may count the same evidence more than once.
The better question
Does this result add information that was not already in the prior?
If the answer is uncertain, the final number should be treated with less precision—not more.
Multiply correlated findings and you count some of the evidence twice. The final number may look beautifully precise while being conceptually shaky.
Before stacking LRs, ask whether each finding is actually adding new information. If you are not sure, be modest about the precision of the answer.
The point
Likelihood ratios are not a replacement for judgment. They are a way to make one part of judgment less mysterious.
Start with the prior. Look at the result that actually occurred. Ask how often that result appears with and without the disease. Then let the LR move you an appropriate distance—no more and no less.
Given what I thought before, how much should this finding change my mind?
Continue Learning
Explore the frameworks behind better emergency medicine decisions.