← STEM Learning Studio
Concept Explainer · STEM

Correlation vs. Causation

Two variables can move together perfectly and still have nothing to do with each other — a third, hidden variable is often driving both.

Correlation is a purely statistical statement: two variables tend to rise and fall together (positive correlation) or move in opposite directions (negative correlation), measured by a correlation coefficient r between −1 and +1. Causationis a much stronger claim: changing one variable directly produces a change in the other. Correlation is necessary for a causal relationship to show up in data, but it is nowhere near sufficient — data alone, without an experimental design or a mechanism, cannot distinguish "A causes B" from "B causes A," from "C causes both A and B," from pure coincidence.

Same correlation, three different underlying structures

r = 0.9, all three
ABcauses1. Direct causationA really does cause BABcauses2. Reverse causationB actually causes A, backwards from assumedCABno direct link — just correlated3. Confounding variableC (e.g. summer heat) drives both A and BClassic example: ice cream sales (A) and drowning deaths (B) — the confounder is hot summer weather (C)
What the correlation shows
r = 0.9 (strong)
Identical number, all three underlying structures — the statistic alone cannot tell them apart.
What settles it
A controlled experiment
Manipulate A directly, hold everything else fixed, and see whether B actually moves.

Observational data vs. a controlled experiment

How you actually prove causation
OBSERVATIONAL STUDYJust record A and B as they occur naturallyConfounders (C) are free to vary along with A→ can show correlation, never proves causeCONTROLLED EXPERIMENTRandomly assign A, hold everything else fixedRandomization spreads confounders evenly across groups→ isolates whether A really drives B
Observational data
Correlation only
Cheap and fast to collect, but confounders ride along uncontrolled.
Randomized experiment
Can support causation
Random assignment breaks the link between the treatment and any confounder.
Why this works

Randomization is what breaks the confounder's grip.

In an observational dataset, anything that influences both the variable you care about and the outcome will show up as a spurious correlation, and you usually can't even name every confounder that might be lurking. A controlled experiment sidesteps the problem structurally rather than statistically: by randomly assigning which subjects get the treatment (the "A" variable) and which don't, any confounding factor — known or unknown — gets distributed roughly evenly between the treatment and control groups by pure chance. If the outcome still differs meaningfully between the groups after that, the only remaining explanation left standing is the treatment itself, because randomization has already ruled out every alternative explanation, including ones nobody thought to measure. That is the entire logical engine behind why designed experiments can support causal claims that raw observational correlation cannot.

Common misconception
"A really high correlation coefficient basically proves causation."

No — the size of r says nothing about the causal question. r measures only the strength and direction of a linear association; it is entirely blind to which variable, if either, is driving the other, and blind to whether a third variable is driving both. There are well-documented "spurious correlation" datasets — for example, per-capita cheese consumption tracking almost perfectly with the number of people who died tangled in their bedsheets over the same years — with correlation coefficients above 0.9 and zero plausible causal link at all; it is coincidence riding on a shared time trend. A very high r can result from direct causation, reverse causation, a shared confounder, a shared trend over time, or pure chance in a small or cherry-picked dataset. Strength of correlation is evidence causation is possible, never evidence that it is true — that requires either a controlled experiment or a well-justified causal mechanism plus techniques (like controlling for known confounders) that rule out the alternatives.

Related Concept Explainers
Independent vs. Mutually Exclusive Events
Why "can never happen together" means maximal dependence
Sample vs. Population Statistics
Why the sample formula divides by (n−1)

Correlation vs. Causation — Concept Explainer

Explains why a strong statistical correlation between two variables never by itself proves that one causes the other, using the ice-cream-sales-and-drowning-deaths confounder example, and why randomized controlled experiments — not raw observational data — are what actually support causal claims.

Why This Is Commonly Confused

Correlation is easy to compute from any dataset you happen to have, while establishing causation generally requires a deliberately designed experiment or a carefully justified causal argument that rules out alternative explanations. Because correlation is so much more accessible, it gets reported constantly, and readers naturally round "these move together" up to "one causes the other." The statistic itself carries no information about mechanism or direction — it is purely a measure of how tightly two variables track each other across the data you observed.

The Formal Distinction

The Pearson correlation coefficient r quantifies the strength and direction of a linear relationship between two variables, ranging from −1 (perfect negative relationship) through 0 (no linear relationship) to +1 (perfect positive relationship). Causation requires an additional structural claim: an intervention that changes variable A produces a corresponding change in variable B, holding everything else fixed. A nonzero correlation is consistent with (and is in fact necessary evidence for) at least four distinct underlying structures: A causes B, B causes A, a third variable C causes both A and B (confounding), or the association arose by chance, especially in small samples or after testing many variable pairs.

Where This Matters in Engineering and Data Analysis

Engineers doing root-cause analysis on process or reliability data run into this constantly: a sensor reading correlated with failure rate might be the actual cause, a downstream symptom of the actual cause, or simply tracking a shared driver like ambient temperature or production shift. Design of Experiments (DOE) methodology exists specifically to establish causal relationships in engineering — by deliberately varying one factor at a time (or using factorial designs) while randomizing or holding other factors constant, engineers isolate which input variables actually drive an output, rather than relying on correlations mined from uncontrolled historical process data, which are far more prone to confounding.

Frequently asked questions

Can correlation ever prove causation on its own?

No. Correlation from observational data alone can never prove causation, because it cannot rule out reverse causation or a confounding variable. It can, however, be combined with a well-established causal mechanism, controlled experiments, or statistical techniques that explicitly control for known confounders to build a credible causal case.

What is a confounding variable?

A confounding variable is a third factor that influences both variables being studied, creating an association between them that has no direct causal link. The classic textbook example is ice cream sales and drowning deaths, which correlate strongly because hot summer weather independently increases both — not because ice cream causes drownings.

Why do randomized controlled experiments help establish causation?

Random assignment of subjects to treatment and control groups spreads both known and unknown confounding factors roughly evenly across the groups purely by chance. If the outcome still differs meaningfully between groups afterward, confounding has been statistically ruled out as the explanation, leaving the treatment itself as the remaining explanation.

Is a negative correlation less meaningful than a positive one?

No — the sign only indicates direction (variables moving oppositely vs. together), not strength or causal validity. A correlation of −0.85 is just as strong an association as +0.85; both require the same scrutiny before any causal claim is drawn from them.

What does "correlation does not imply causation" actually mean in practice?

It means observing that two variables move together is not, by itself, sufficient justification to conclude that changing one will change the other. Before acting on a correlation — redesigning a process, changing a policy, adjusting a control input — an engineer should look for a plausible mechanism, rule out reverse causation and known confounders, and ideally test the relationship with a controlled experiment before treating it as causal.

🎓

Try our STEM Learning Studio

More calculators, simulators, and guides for this discipline.

📖

STEM Fundamentals Handbook (Full Access)

Premium Content

A zoomable interactive reader — free preview, then unlock the full set.

Related tools & guides

STEM Fundamentals HandbookNumerical Methods VisualizerEngineering Unit Converter