2×2 Diagnostic Table
| Condition Present (Reference standard +) |
Condition Absent (Reference standard −) |
|
|---|---|---|
| Test Positive | True Positive (TP) Has the condition, test said positive. | False Positive (FP) No condition, test said positive. |
| Test Negative | False Negative (FN) Has the condition, test said negative. | True Negative (TN) No condition, test said negative. |
Target prevalence lets you see what PPV and NPV would be in a different population using the same sensitivity and specificity. Run the main 2×2 table first — if no values are entered there, these calculations cannot proceed.
Enter values in the 2×2 Table tab first, then return here to see every calculation written out step by step.
What Are Sensitivity and Specificity?
Sensitivity is the proportion of people who truly have the condition who receive a positive test result. Specificity is the proportion of people who truly do not have the condition who receive a negative test result. Both are calculated from a 2×2 diagnostic contingency table, and both condition on reference-standard status — not on the test result.
That last point is the one that trips most people up. Sensitivity tells you how good the test is at catching disease in people who actually have it. It says nothing about what a positive result means to the person who just received one. That question — given a positive result, what is the probability of having the condition? — belongs to PPV, and PPV depends on prevalence in a way that sensitivity does not.
How to Use This Calculator
Fill in TP, FP, FN, and TN from your 2×2 table. These must be non-negative whole numbers. The cells are color-coded: green for correct classifications (TP, TN), red for errors (FP, FN).
These appear immediately in the top summary band. The Step-by-Step tab shows the numerator, denominator, and arithmetic in full so you can verify every number.
PPV and NPV are shown with the sample prevalence from your table. If your study prevalence differs from the population you care about, use the Advanced tab to enter a target prevalence and get Bayes-adjusted estimates.
The Wilson method is more reliable than the simple Wald interval near 0% or 100%, which is exactly where many real diagnostic tests sit. The calculator uses Wilson intervals for all four proportions.
The bottom of the results panel generates a reporting sentence in the style common to methods sections: “Sensitivity was X% (95% CI: A–B%), specificity was Y% (95% CI: C–D%)...” Hit the Copy button to paste it directly into a manuscript or report.
Worked Example (n = 1,000)
This is the example loaded by the “Example (n = 1,000)” button. Verify the arithmetic by hand, then cross-check the calculator output.
Table
TP = 90 FP = 20
FN = 10 TN = 880
Total N = 1,000
Sensitivity
90 / (90 + 10)
= 90 / 100
= 90.00%
Specificity
880 / (880 + 20)
= 880 / 900
= 97.78%
PPV
90 / (90 + 20)
= 90 / 110
= 81.82%
NPV
880 / (880 + 10)
= 880 / 890
= 98.88%
Accuracy
(90 + 880) / 1,000
= 970 / 1,000
= 97.00%
LR+
0.9000 / (1 − 0.9778)
= 0.9000 / 0.0222
= 40.50
Youden’s J
0.9000 + 0.9778 − 1
= 0.8778
Reading these numbers: Of every 100 people with the condition, 90 are caught (sensitivity = 90%). Of every 900 people without the condition, 880 are correctly cleared (specificity = 97.78%). But — only 81.82% of positive results are true positives (PPV), because there are 20 false positives among the 110 test-positives. That gap between sensitivity and PPV grows even wider at lower prevalence.
The Denominator Is Everything
The most common error in interpreting diagnostic test statistics is using the wrong denominator. Each metric describes a percentage within a specific group, and which group that is defines what the number actually means.
Sensitivity
Denominator: TP + FN All people the reference standard confirmed as condition-positive. “Among those who truly have it, how many did the test catch?”Specificity
Denominator: TN + FP All people the reference standard confirmed as condition-negative. “Among those who truly don’t have it, how many did the test correctly clear?”PPV
Denominator: TP + FP All people who tested positive, regardless of true status. “Among those who got a positive result, how many actually have the condition?”NPV
Denominator: TN + FN All people who tested negative, regardless of true status. “Among those who got a negative result, how many are truly disease-free?”This is also why sensitivity cannot be confused with PPV, even though both are expressed as percentages and both involve TP in the numerator. They divide by completely different totals and answer completely different questions. Sensitivity uses the reference-standard column; PPV uses the test-result row.
Why the Same Test Has Different PPV in Different Populations
Sensitivity and specificity are conditional on reference-standard status, so they describe the test regardless of how common the condition is. PPV and NPV are conditional on the test result, so they change whenever the underlying prevalence changes — even if the test itself is unchanged.
Here is a concrete way to see it. Take a test with 95% sensitivity and 95% specificity. In a population where 50% have the condition, the PPV is about 95%. In a population where 1% have the condition, the PPV drops to roughly 16%. The test is identical. Only the population changed.
Sensitivity vs. Accuracy: They Measure Different Things
| Metric | Numerator | Denominator | Answers |
|---|---|---|---|
| Sensitivity | TP | TP + FN (all condition-positive) | How many with the condition did the test find? |
| Specificity | TN | TN + FP (all condition-negative) | How many without the condition did the test correctly clear? |
| Accuracy | TP + TN | N (everyone) | What fraction of all observations were correctly classified? |
| PPV | TP | TP + FP (all test-positive) | If the test is positive, what is the probability of the condition? |
| NPV | TN | TN + FN (all test-negative) | If the test is negative, what is the probability of no condition? |
Accuracy looks attractive because it is a single number from 0% to 100%, but it can be badly misleading when the condition is rare. A test that always returns “negative” achieves 99% accuracy in a population where only 1% of people have the condition — while having 0% sensitivity and catching nobody. Sensitivity and specificity are the right tools for evaluating test performance because they are not distorted by prevalence in the same way.
Likelihood Ratios: Updating Probability Without Knowing Prevalence
Unlike PPV and NPV, likelihood ratios are independent of prevalence and can be applied to any pre-test probability to estimate a post-test probability.
Positive LR (LR+)
LR+ = Sensitivity
/ (1 − Specificity)
High LR+ = strong evidence
for the condition when
the test is positive.
Negative LR (LR−)
LR− = (1 − Sensitivity)
/ Specificity
Low LR− (close to 0) = strong
evidence against the condition
when the test is negative.
Post-Test Odds
Pre-test odds
= p / (1 − p)
Post-test odds (positive)
= Pre-test odds × LR+
Post-test probability
= odds / (1 + odds)
Youden’s J
J = Sensitivity
+ Specificity − 1
Range: −1 to +1
+1 = perfect test
0 = no discrimination
In the worked example above, LR+ = 40.5. That means a positive result is 40 times more likely in condition-positive people than in condition-negative people. Whether that translates into a high PPV depends on what fraction of your tested population has the condition — which is exactly what the Advanced tab models by letting you enter a target prevalence.
Complete Formula Reference
| Metric | Symbol | Formula | Range |
|---|---|---|---|
| Sensitivity | Se | TP / (TP + FN) | 0 to 1 |
| Specificity | Sp | TN / (TN + FP) | 0 to 1 |
| PPV | PPV | TP / (TP + FP) | 0 to 1, depends on prevalence |
| NPV | NPV | TN / (TN + FN) | 0 to 1, depends on prevalence |
| Accuracy | — | (TP + TN) / N | 0 to 1 |
| Prevalence (sample) | p | (TP + FN) / N | 0 to 1 |
| False Negative Rate | FNR | FN / (TP + FN) = 1 − Se | 0 to 1 |
| False Positive Rate | FPR | FP / (TN + FP) = 1 − Sp | 0 to 1 |
| Positive Likelihood Ratio | LR+ | Se / (1 − Sp) | 0 to ∞ |
| Negative Likelihood Ratio | LR− | (1 − Se) / Sp | 0 to 1 |
| Youden’s J | J | Se + Sp − 1 | −1 to 1 |
| Diagnostic Odds Ratio | DOR | (TP × TN) / (FP × FN) | 0 to ∞ |
| Bayes PPV | — | (Se × p) / [(Se × p) + ((1−Sp) × (1−p))] | 0 to 1, for any prevalence p |
Common Mistakes When Working With Diagnostic Statistics
A 95% sensitive test does not mean that 95% of positive results are correct. Sensitivity measures detection among people who have the condition; PPV measures accuracy among people who tested positive. These are two different denominators and two different questions.
PPV from a case-enriched validation study cannot be transplanted to a screening context where the condition is much rarer. Always state the prevalence alongside any predictive value.
In a dataset where 99% of people are condition-negative, a test that always returns “negative” achieves 99% accuracy with 0% sensitivity. Accuracy is only informative when the condition-positive and condition-negative groups are roughly balanced.
Double-check that your TP cell represents people who both tested positive and have the condition, and your TN cell represents people who both tested negative and do not have the condition. The labels in this calculator include one-line descriptions to help you match your data correctly.
Sensitivity of 80% from 5 out of 6 condition-positive cases is very different from 80% out of 500. The 95% Wilson confidence interval will be much wider for the smaller sample, which reflects genuine uncertainty that should not be hidden by quoting only the point estimate.
A 2×2 table reflects performance at one specific threshold. The area under the ROC curve describes performance across all thresholds. You cannot calculate AUC from a single set of TP, FP, FN, TN counts.
Why This Calculator Uses Wilson Confidence Intervals
The standard Wald interval — p ± 1.96√(p(1−p)/n) — breaks down when the estimated proportion is near 0% or 100%. For a test with 100% observed sensitivity from 10 condition-positive cases, the Wald formula gives an upper bound above 100% and a lower bound of 100%, which is clearly wrong. The Wilson score interval handles boundary cases correctly by building the confidence interval around the test statistic rather than the estimate itself. For proportions near 50% with large samples, the two methods give nearly identical results, so there is no downside to using Wilson throughout.
Related Topics on Statistics Fundamentals
Diagnostic test performance connects to probability theory, hypothesis testing, study design, and health statistics. These pages build out the full picture.
Frequently Asked Questions
Sensitivity is the proportion of people who truly have the condition who receive a positive test result: TP / (TP + FN). A test with 90% sensitivity correctly identifies 90 out of every 100 condition-positive cases. The remaining 10 receive false negatives — they have the condition but the test missed them. In classification terminology, sensitivity is also called the true positive rate or recall.
Specificity is the proportion of people who do not have the condition who receive a negative test result: TN / (TN + FP). A test with 97% specificity correctly clears 97 out of every 100 condition-negative people. The remaining 3% receive false positives. Specificity is also called the true negative rate. One minus specificity is the false positive rate, which is the x-axis of an ROC curve.
They look at two different groups of people. Sensitivity is calculated only among people confirmed to have the condition (denominator = TP + FN). Specificity is calculated only among people confirmed not to have the condition (denominator = TN + FP). A test can be high in one and low in the other. Raising a test threshold typically increases specificity while decreasing sensitivity, and vice versa.
PPV (Positive Predictive Value) is the proportion of positive test results that are correct: TP / (TP + FP). NPV (Negative Predictive Value) is the proportion of negative results that are correct: TN / (TN + FN). Unlike sensitivity and specificity, PPV and NPV depend on how common the condition is in the population being tested. The same test will have a lower PPV in a population where the condition is rare, even if sensitivity and specificity are unchanged.
Sensitivity is conditioned on reference-standard status — it is calculated entirely within the group that has the condition. Changes in how many people have the condition do not change the proportion of that group that the test finds. PPV is conditioned on the test result — it is calculated across everyone who tested positive, which mixes true positives and false positives. When prevalence falls, there are fewer true positives but the number of false positives stays proportionally the same, so PPV falls. This is the core of the Bayes theorem adjustment in the Advanced tab.
LR+ = Sensitivity / (1 − Specificity). It describes how many times more likely a positive result is among people with the condition compared to people without it. An LR+ of 10 means a positive result is 10 times more likely in condition-positive people. LR+ is useful because it does not change with prevalence, so it can be applied to any pre-test probability to calculate a post-test probability using Bayes’ theorem in odds form.
Youden’s J = Sensitivity + Specificity − 1. It ranges from −1 to 1. A value of 1 means both sensitivity and specificity are 100% — a perfect test. A value of 0 means the test offers no discrimination: it performs no better than a coin flip under the symmetric model. Youden’s J is widely used to choose an optimal cut-off threshold on a continuous test score, because the threshold that maximises J simultaneously maximises the sum of sensitivity and specificity.
Zero counts make some metrics undefined rather than zero. If TP + FN = 0, there are no condition-positive reference cases in the data, so sensitivity cannot be estimated. If TN + FP = 0, specificity cannot be estimated. If TP + FP = 0, PPV is undefined. The calculator identifies these situations and labels the metric as “undefined” rather than displaying 0%, which would be mathematically incorrect. The Wilson confidence interval is still computed for defined proportions even when the point estimate is 0% or 100%.
The diagnostic odds ratio (DOR) = (TP × TN) / (FP × FN) = LR+ / LR−. It summarizes the overall discriminatory performance of the test in a single number. A DOR of 1 means the test cannot distinguish between condition-positive and condition-negative cases. Higher values indicate better discrimination. Like likelihood ratios, DOR is independent of prevalence. Its limitation is that the same DOR value can arise from very different sensitivity/specificity combinations, so it should not replace the individual metrics in a report.
The Wald interval (p ± z√(p(1−p)/n)) is unreliable when the proportion is close to 0% or 100%, or when the sample size is small. Because many diagnostic tests have very high or very low sensitivity and specificity, the Wald interval frequently produces bounds above 100% or below 0%, or it collapses to a single point at the boundary. The Wilson score interval is built differently — it inverts the hypothesis test rather than using the standard error of the estimate — and it performs much better at boundaries and with small samples. The practical difference is small for proportions near 50% with n ≥ 100.