Two independent accuracy metrics, each answering a completely different question — and a diagnostic device can excel at one while failing the other.
Every diagnostic device — a blood glucose meter, a pulse oximeter alarm threshold, an ELISA assay, an AI-based imaging classifier — eventually gets reduced to a yes/no call against ground truth: does the patient actually have the condition, and did the device say so? From that single 2×2 comparison, two structurally different metrics fall out. Sensitivity answers: of everyone who truly has the condition, what fraction did the device correctly flag? Specificity answers a different question entirely: of everyone who truly does not have the condition, what fraction did the device correctly clear?Both are measured on the same 0–100% scale and both sound like generic "accuracy," which is exactly why they get conflated — but a device can be excellent at one and mediocre at the other, and which one matters more depends entirely on the clinical cost of being wrong in each direction.