Skip to content
Calc.
← All articles
statisticsBy Emil BjörkOctober 3, 2026

Understanding Correlation: What the Coefficient Tells You

When you notice that taller people tend to weigh more, or that students who study longer tend to score higher, you're observing correlation. But how strong is that relationship—a tight, reliable pattern or just a loose tendency? A single number, the correlation coefficient, answers that question and turns vague hunches into precise, defensible statements about your data.

In this guide, you'll learn what correlation measures, how to interpret Pearson's r, and how to avoid the most dangerous mistake people make when reading it. By the end, you'll know not just what the number says, but what it can and cannot tell you.

What Correlation Actually Measures

Correlation measures the degree to which two variables move together. When one changes, does the other tend to change predictably? Correlation quantifies both the direction of that relationship (do they rise together or move in opposite directions?) and its strength (how consistently does the pattern hold?).

It's important to understand what correlation is not. It does not measure how much one variable changes for a given change in another—that's the job of regression slopes. It simply tells you how tightly the two variables track each other.

The most widely used measure is the Pearson correlation coefficient, written as r. It captures the strength of the linear relationship between two continuous variables—a crucial qualifier we'll return to later.

Pearson's r: A Scale From −1 to +1

Pearson's r always falls between −1 and +1. Those two endpoints and the midpoint anchor your interpretation:

  • r = +1 indicates a perfect positive linear relationship. Every data point sits exactly on an upward-sloping line.
  • r = −1 indicates a perfect negative linear relationship. Every point sits on a downward-sloping line.
  • r = 0 indicates no linear relationship at all. Knowing one variable tells you nothing linear about the other.
The sign tells you the direction. A positive r means the variables move in the same direction: as one increases, the other tends to increase. A negative r means they move in opposite directions: as one increases, the other tends to decrease.

The magnitude (how close the absolute value is to 1) tells you the strength. While exact thresholds vary by field, a common rule of thumb interprets r like this:

  • 0.0 to 0.3 (or 0.0 to −0.3): weak or negligible
  • 0.3 to 0.7 (or −0.3 to −0.7): moderate
  • 0.7 to 1.0 (or −0.7 to −1.0): strong
So an r of −0.85 describes a strong negative relationship, while an r of +0.20 describes a weak positive one. The sign and the magnitude are independent: −0.85 is a much stronger relationship than +0.20, even though one is negative.

A Worked Conceptual Example

Imagine a coffee shop tracking daily outdoor temperature against iced-coffee sales over two weeks. On colder days, sales hover low; on warmer days, sales climb steadily. When you plot temperature against sales, the points form a clear upward-sloping cloud.

Running the numbers might yield an r of +0.82. That single value tells you two things at once. First, the positive sign confirms the direction your eyes already saw: warmer days bring higher sales. Second, the magnitude of 0.82 tells you the relationship is strong—the points cluster tightly around an upward line, so temperature is a reliable indicator of iced-coffee demand.

Now suppose hot-coffee sales over the same period produced an r of −0.74. The negative sign reveals the opposite direction—hot coffee sells less as temperatures rise—and the magnitude tells you this inverse pattern is also strong.

You don't need to grind through the formula by hand. A correlation coefficient calculator takes your paired data points and returns r instantly, letting you focus on interpretation rather than arithmetic.

Correlation Is Not Causation

This is the single most important caution in all of statistics, and the most frequently ignored. A strong correlation tells you two variables move together—it does not prove that one causes the other.

Consider a classic example: across a city's summer months, ice cream sales and drowning incidents are strongly positively correlated. Does buying ice cream cause people to drown? Obviously not. A third variable—hot weather—drives both. Heat pushes people to buy ice cream and sends them swimming, which raises drowning incidents. The ice cream and the drownings are linked only through this hidden common cause.

Before concluding that X causes Y, you must rule out three alternatives: Y might cause X (reverse causation), a third variable might cause both (confounding), or the correlation might simply be coincidence in a small or cherry-picked sample. Establishing genuine causation requires controlled experiments or carefully designed studies—never a correlation coefficient alone.

Positive, Negative, and No Correlation

It helps to picture three scatterplots side by side:

  • Positive correlation (r > 0): points trend upward from lower-left to upper-right. Example: hours studied and exam scores.
  • Negative correlation (r < 0): points trend downward from upper-left to lower-right. Example: a car's age and its resale value.
  • No correlation (r ≈ 0): points scatter randomly with no discernible slope. Example: shoe size and intelligence.
A near-zero coefficient doesn't necessarily mean the variables are unrelated—only that they have no linear relationship, which leads directly to the limitations you must respect.

Limitations You Must Keep in Mind

Pearson's r is powerful but narrow. Two limitations deserve special attention.

First, it only detects linear relationships. Suppose plant growth increases with fertilizer up to a point, then declines as too much fertilizer burns the roots. This forms a strong, clear, upside-down-U pattern—yet Pearson's r might come out near zero because the upward and downward halves cancel out. The relationship is real and strong; it's just not linear. Always plot your data before trusting a coefficient.

Second, r is highly sensitive to outliers. A single extreme point can dramatically inflate or deflate the coefficient, making a weak relationship look strong or masking a strong one. Examining a scatterplot lets you spot outliers that the number alone would hide.

For these reasons, the coefficient should never be your only diagnostic. Visualize the data, watch for non-linear shapes, and investigate any points that sit far from the crowd.

Key Takeaways

  • Correlation measures direction and strength of how two variables move together, but not how much one changes relative to the other.
  • Pearson's r ranges from −1 to +1, where the sign indicates direction (positive or negative) and the absolute value indicates strength, with values near ±1 being strong and values near 0 being weak.
  • Correlation is never proof of causation—a third variable, reverse causation, or coincidence can all produce strong correlations, as the ice-cream-and-drownings example shows.
  • Pearson's r only captures linear relationships, so a curved pattern can yield a near-zero coefficient despite a real, strong association.
  • Outliers can distort r significantly, which is why you should always plot your data alongside computing the coefficient.
Mastering the correlation coefficient turns raw data into clear insight, but its real value comes from interpreting it responsibly. Read the sign for direction, the magnitude for strength, plot your data to confirm the shape, and resist the urge to leap from correlation to cause. Do that consistently, and a single number becomes one of the most useful tools in your analytical toolkit.

Related articles

Looking for a calculator?

Calculator Collection has 3,800+ free calculators. Browse all calculators →