Two variables may be in perfect sync for years before one has any impact on the other. Correlation analysis is the process of measuring this synchronization, and understanding its limitations is what distinguishes an accurate conclusion from a misleading one.
Correlation analysis is a statistical technique that quantifies the strength and direction of the relationship between two variables, expressed in terms of a single figure (r) ranging between -1 and +1. The closer r is to +1, the more the two variables increase in tandem. The closer it is to -1, the more one variable increases while the other decreases. The closer it is to 0, the less linear the relationship.
The single thing it does not do is establish a cause-and-effect relationship. This is where most conclusions drawn from correlation are faulty, and it is explained thoroughly below.
What Will I Learn?
What Correlation Actually Measures
Two characteristics of a correlation relationship among variables include the direction of the relationship and the strength of the relationship.
Direction refers to whether the two variables move together in the same direction or in opposite directions. Direction can take on any of the following values:
- Positive correlation: Both variables move together in the same direction. Example: Hours of study and test score.
- Negative correlation: Variables move in opposite directions. Example: Outside temperature and costs of heating.
- No correlation: Relationship between variables has no discernible pattern. Example: Shoe size and income earned per month.
The strength of the relationship refers to the closeness of the scatter points to a straight line. Strength is determined using a value from -1 to +1.
The Correlation Coefficient (r), Explained
Data Analyst Course
Average time: 6 months
Skills you’ll build: SQL, Python for Data Analysis, Power BI, Excel with AI, Data Storytelling, Stakeholder Reporting
Correlation Coefficient r is a numerical value resulting from a correlation analysis. It indicates the direction and degree of a relationship in one value from -1 to +1.
The correlation coefficient is dimensionless. Hence, the r-value is independent of the units of measure of the underlying variables. A correlation of height measured in centimeters with weight measured in kilograms will yield the same r-value as the same data but in inches and pounds.
Karl Pearson developed the popular correlation coefficient in the late 1890s. The Pearson correlation coefficient assesses linear correlation between two continuous variables.
The 3 Main Types of Correlation Coefficients
Three correlation coefficients are used most often in statistical analysis. Each applies to a different type of data.
| Coefficient | Named After | Measures | Data Type Required |
| Pearson correlation coefficient | Karl Pearson | Linear relationships | Continuous, normally distributed data |
| Spearman’s rank correlation coefficient | Charles Spearman | Monotonic relationships (not necessarily linear) | Ordinal data, or continuous data that is not normally distributed |
| Kendall’s Tau | Maurice Kendall | Rank agreement between two variables | Ordinal data, smaller sample sizes |
The Pearson’s correlation coefficient requires both the variables to have a normal distribution and the relation between them to be linear. The Spearman’s correlation coefficient involves first ranking the values and then calculating the coefficient on those ranked values. Thus it is useful for ordinal variables like customer satisfaction indices and for situations where there are outliers.
Which Coefficient Should You Use?
The correct coefficient depends on two factors: the type of data and whether the relationship is linear.
| Your Data | Relationship Type | Use This Coefficient |
| Continuous, normally distributed | Linear | Pearson |
| Ordinal, or continuous but not normally distributed | Monotonic (consistently increasing or decreasing, not necessarily straight) | Spearman |
| Ordinal, small sample size | Monotonic | Kendall’s Tau |
It is advisable to use Pearson’s coefficient where the assumptions about it hold since it gives accurate results. In case the data have outliers or are not normally distributed, then it is wise to use Spearman’s coefficient.
How to Interpret a Correlation Coefficient
The closer a correlation coefficient is to +1 or –1, the stronger the relationship. The closer a correlation coefficient is to 0, the weaker the relationship.
Jacob Cohen, a statistician, offered a commonly cited standard of interpretation of correlation strength in the social sciences:
| Absolute Value of r | Strength |
| 0.00 – 0.10 | Negligible |
| 0.10 – 0.39 | Weak |
| 0.40 – 0.69 | Moderate |
| 0.70 – 0.89 | Strong |
| 0.90 – 1.00 | Very strong |
They are conventions, not absolute laws. Correlation value 0.40 can be taken as strong in some areas like psychology but weak in other areas like physics, since there is less measurement error in physics.
A Worked Example: Correlation Analysis Step by Step
This example calculates the correlation between daily temperature and ice cream sales over 10 days.
Step 1: Collect the data.
| Day | Temperature (°F) | Ice Cream Units Sold |
| 1 | 68 | 120 |
| 2 | 71 | 135 |
| 3 | 74 | 150 |
| 4 | 79 | 170 |
| 5 | 82 | 190 |
| 6 | 85 | 205 |
| 7 | 88 | 230 |
| 8 | 90 | 255 |
| 9 | 93 | 270 |
| 10 | 95 | 290 |
Step 2: Calculate the mean of each variable.
The mean temperature is 82.5°F. The mean sales figure is 201.5 units.
Step 3: Calculate the deviations and their products.
For each day, subtract the mean from the actual value, then multiply the two deviations together. Sum these products across all 10 days. This sum equals 4,892.5.
Step 4: Calculate the denominator.
We will multiply the sum of squared temperature deviations by the sum of squared sales deviations and get the square root of it. This product is 4,941.9.
Step 5: Divide to find r.
4,892.5 ÷ 4,941.9 = 0.99
Step 6: Interpret the result.
If r = 0.99, it means there is a very strong positive correlation between temperature and ice cream sales. As one goes up, the other increases proportionally.
This result proves only the existence of correlation, which does not imply causation. The business may use this fact to manage its inventory based on weather forecasts, but further research is needed to confirm what makes customers buy more on hot days.
How to Calculate Correlation
There are three major ways to determine the correlation coefficient: spreadsheets, programming, and by hand, as shown above.
In Excel or Google Sheets
Microsoft Excel and Google Sheets both have a pre-programmed function called CORREL which determines the Pearson correlation coefficient automatically.
- Input the values for your first variable in one column.
- Input the values for your second variable in a column right next to it.
- In an empty cell, enter the formula =CORREL(range1, range2), replacing range1 and range2 with the two data ranges.
- Press Enter. The cell displays the correlation coefficient.
Using the worked example above, the formula =CORREL(A2:A11, B2:B11) returns 0.99.
In Python (NumPy and pandas)
Python calculates correlation coefficients through the NumPy and pandas libraries.
Using NumPy:
python
import numpy as np
temperature = np.array([68, 71, 74, 79, 82, 85, 88, 90, 93, 95])
sales = np.array([120, 135, 150, 170, 190, 205, 230, 255, 270, 290])
correlation_matrix = np.corrcoef(temperature, sales)
print("Correlation Coefficient:", correlation_matrix[0, 1])
Output: Correlation Coefficient: 0.99
Using pandas:
python
import pandas as pd
data = pd.DataFrame({
'Temperature': [68, 71, 74, 79, 82, 85, 88, 90, 93, 95],
'Sales': [120, 135, 150, 170, 190, 205, 230, 255, 270, 290]
})
correlation = data['Temperature'].corr(data['Sales'])
print("Correlation Coefficient:", correlation)
Output: Correlation Coefficient: 0.99
The pandas .corr() method also generates a correlation matrix when applied to a full DataFrame with more than two columns. A correlation matrix displays the correlation coefficient between every pair of variables in a dataset, arranged in a grid. This makes it possible to identify relationships across many variables at once, rather than testing pairs individually.
Correlation vs. Regression vs. Covariance
Correlation, regression, and covariance each describe relationships between variables, but they answer different questions.
| Method | What It Measures | Output Range | Primary Use |
| Correlation | Strength and direction of a linear relationship | -1 to +1 | Determine whether and how strongly two variables are related |
| Regression | The mathematical equation that predicts one variable from another | Not bounded | Predict the value of a dependent variable based on an independent variable |
| Covariance | The direction of the relationship between two variables | Not bounded, depends on measurement units | Determine whether two variables increase or decrease together, before standardizing to correlation |
Both covariance and correlation are measurements that point to the direction of the relationship. Covariance does not indicate strength since the measurement is relative, based on units of the variables measured. On the other hand, correlation normalizes the measurement by dividing the covariance by the standard deviation of both variables.
Regression analysis goes beyond correlation analysis through generating an equation for predicting one variable from the other. Correlation simply indicates the existence of a relationship and how strong it is. Regression on the other hand gives the change in one variable when another is changed.
Why Correlation Isn’t Causation
A relationship between two variables does not mean that one causes the other. There could be a third unknown variable causing changes in both variables.
Example: Sales of ice cream and cases of drowning have shown a high positive correlation in the summer season. The ice cream is not causing drowning. The reason is that there is a third variable which causes an increase in both variables, which is warm temperature leading to more swimming and more ice cream consumption. Such a relationship is referred to as a spurious correlation.
There are three requirements to prove that the relationship is a causal relationship, first being cause preceding the effect, second being correlation between the variables and the third being excluding other possibilities. Correlation analysis fulfills only the second requirement.
Common Mistakes When Interpreting Correlation
Data Analyst Course
Average time: 6 months
Skills you’ll build: SQL, Python for Data Analysis, Power BI, Excel with AI, Data Storytelling, Stakeholder Reporting
There are four common mistakes made by researchers when interpreting correlation coefficients.
- Linearity assumption. The Pearson correlation assumes linearity in the relationship between two variables. It is possible for two variables to have a high degree of curved relationship but a Pearson r close to 0.
- Outlier problem. One outlier can distort or enhance the correlation coefficient. Checking outliers will determine their effect.
- Sample size problem. Correlation estimates with small samples are unreliable. There is a lot of difference between a correlation estimate with five data points compared to one with 500 data points.
- Not testing for statistical significance. Correlation coefficient alone does not prove the existence of a relationship, just that there is a relationship. Statistical significance testing proves that there is a significant relationship that is not by chance. It usually involves a t-test on the correlation coefficient.
Real-World Applications of Correlation Analysis
Correlation is employed in a number of different areas in order to discover the association between different variables prior to undertaking deeper analysis.
- Healthcare: Scientists analyze the correlation between some treatment method and its outcome, for example, the relationship between the dose of blood pressure medicine and the results of blood pressure readings.
- Finance: Traders analyze the correlation between prices of different assets, for example, two shares, or a share and a rate of interest.
- Marketing: Specialists analyze the correlation between customer satisfaction and repurchase rates to understand which variables affect retention.
- Human resources: Human resources department analyzes the correlation between engagement scores of employees and their productivity in order to assess company policies.
- Machine learning: Specialists use correlation matrices when selecting features in order to understand which variables correlate best with target one and to diagnose multicollinearity, which means that two or more input variables are correlated.
Advantages and Limitations of Correlation Analysis
Advantages
- Generates one, consistent value (r) that is easy to compute and interpret.
- Does not require manipulation of variables, making it applicable in cases where experimentation is impractical or unethical.
- Acts as an initial step in uncovering relationships worth examining further.
- Aids in selecting features for machine learning through ranking of variables based on their relation to a target variable.
Limitations
- Does not prove cause-and-effect relationship.
- Can recognize only linear (Pearson) and monotonic (Spearman, Kendall) relations, ignoring other types of connections.
- Is vulnerable to the presence of outliers.
- Does not yield valid outcomes in case of small samples.
- Is not capable of taking into account the impact of additional variables.
FAQs
Q1. What is a good correlation coefficient?
Ans. The higher the absolute value of a correlation coefficient exceeds 0.70, the stronger it is; however, acceptable levels depend on the specific domain and must be viewed in combination with the sample size and level of statistical significance.
Q2. Can a correlation coefficient be negative?
Ans. Certainly, a correlation coefficient varies between -1 and +1 and a negative correlation coefficient means that the larger the first variable is, the smaller the second one is and vice versa.
Q3. What is the difference between correlation and covariance?
Ans. While covariance reveals the direction of the relation in the units of the variables, correlation makes that measure standardized and unitless within the range from -1 to +1.
Q4. Do I need a large sample size for correlation analysis?
Ans. Indeed, the larger the sample size, the more reliable the estimate is; in contrast, a small sample size entails the risk of obtaining a coefficient which does not reflect the real situation.
Q5. What tools can I use to run correlation analysis?
Ans. There are various methods to conduct correlation analysis including Excel/Google Sheets with the use of the CORREL function, Python programming language through NumPy/pandas or specialized software, such as R/SPSS.