What Is Correlation Analysis? 

|
9 min read
|
62 views
What Is Correlation Analysis

Two variables may be in perfect sync for years before one has any impact on the other. Correlation analysis is the process of measuring this synchronization, and understanding its limitations is what distinguishes an accurate conclusion from a misleading one.

Correlation analysis is a statistical technique that quantifies the strength and direction of the relationship between two variables, expressed in terms of a single figure (r) ranging between -1 and +1. The closer r is to +1, the more the two variables increase in tandem. The closer it is to -1, the more one variable increases while the other decreases. The closer it is to 0, the less linear the relationship.

The single thing it does not do is establish a cause-and-effect relationship. This is where most conclusions drawn from correlation are faulty, and it is explained thoroughly below.

What Correlation Actually Measures

Two characteristics of a correlation relationship among variables include the direction of the relationship and the strength of the relationship.

Direction refers to whether the two variables move together in the same direction or in opposite directions. Direction can take on any of the following values:

  • Positive correlation: Both variables move together in the same direction. Example: Hours of study and test score.
  • Negative correlation: Variables move in opposite directions. Example: Outside temperature and costs of heating.
  • No correlation: Relationship between variables has no discernible pattern. Example: Shoe size and income earned per month.

The strength of the relationship refers to the closeness of the scatter points to a straight line. Strength is determined using a value from -1 to +1.

The Correlation Coefficient (r), Explained

Professional Certificate

Data Analyst Course

Get on the fast track to a career in data analytics. Learn SQL, Python, Power BI and Excel with AI-assisted workflows, and build the skills employers screen for — no degree or prior coding experience required.

4.8 (18,340 ratings)  •  46,210 already enrolled  •  Beginner level

Class Starts on 3 Oct, 2026 — SAT & SUN (Weekend Batch)

Average time: 6 months  

Skills you’ll build: SQL, Python for Data Analysis, Power BI, Excel with AI, Data Storytelling, Stakeholder Reporting

Correlation Coefficient r is a numerical value resulting from a correlation analysis. It indicates the direction and degree of a relationship in one value from -1 to +1.

The correlation coefficient is dimensionless. Hence, the r-value is independent of the units of measure of the underlying variables. A correlation of height measured in centimeters with weight measured in kilograms will yield the same r-value as the same data but in inches and pounds.

Karl Pearson developed the popular correlation coefficient in the late 1890s. The Pearson correlation coefficient assesses linear correlation between two continuous variables.

The 3 Main Types of Correlation Coefficients

Three correlation coefficients are used most often in statistical analysis. Each applies to a different type of data.

CoefficientNamed AfterMeasuresData Type Required
Pearson correlation coefficientKarl PearsonLinear relationshipsContinuous, normally distributed data
Spearman’s rank correlation coefficientCharles SpearmanMonotonic relationships (not necessarily linear)Ordinal data, or continuous data that is not normally distributed
Kendall’s TauMaurice KendallRank agreement between two variablesOrdinal data, smaller sample sizes

The Pearson’s correlation coefficient requires both the variables to have a normal distribution and the relation between them to be linear. The Spearman’s correlation coefficient involves first ranking the values and then calculating the coefficient on those ranked values. Thus it is useful for ordinal variables like customer satisfaction indices and for situations where there are outliers.

Which Coefficient Should You Use?

The correct coefficient depends on two factors: the type of data and whether the relationship is linear.

Your DataRelationship TypeUse This Coefficient
Continuous, normally distributedLinearPearson
Ordinal, or continuous but not normally distributedMonotonic (consistently increasing or decreasing, not necessarily straight)Spearman
Ordinal, small sample sizeMonotonicKendall’s Tau

It is advisable to use Pearson’s coefficient where the assumptions about it hold since it gives accurate results. In case the data have outliers or are not normally distributed, then it is wise to use Spearman’s coefficient.

How to Interpret a Correlation Coefficient

The closer a correlation coefficient is to +1 or –1, the stronger the relationship. The closer a correlation coefficient is to 0, the weaker the relationship.

Jacob Cohen, a statistician, offered a commonly cited standard of interpretation of correlation strength in the social sciences:

Absolute Value of rStrength
0.00 – 0.10Negligible
0.10 – 0.39Weak
0.40 – 0.69Moderate
0.70 – 0.89Strong
0.90 – 1.00Very strong

They are conventions, not absolute laws. Correlation value 0.40 can be taken as strong in some areas like psychology but weak in other areas like physics, since there is less measurement error in physics.

A Worked Example: Correlation Analysis Step by Step

This example calculates the correlation between daily temperature and ice cream sales over 10 days.

Step 1: Collect the data.

DayTemperature (°F)Ice Cream Units Sold
168120
271135
374150
479170
582190
685205
788230
890255
993270
1095290

Step 2: Calculate the mean of each variable.

The mean temperature is 82.5°F. The mean sales figure is 201.5 units.

Step 3: Calculate the deviations and their products.

For each day, subtract the mean from the actual value, then multiply the two deviations together. Sum these products across all 10 days. This sum equals 4,892.5.

Step 4: Calculate the denominator.

We will multiply the sum of squared temperature deviations by the sum of squared sales deviations and get the square root of it. This product is 4,941.9.

Step 5: Divide to find r.

4,892.5 ÷ 4,941.9 = 0.99

Step 6: Interpret the result.

If r = 0.99, it means there is a very strong positive correlation between temperature and ice cream sales. As one goes up, the other increases proportionally.

This result proves only the existence of correlation, which does not imply causation. The business may use this fact to manage its inventory based on weather forecasts, but further research is needed to confirm what makes customers buy more on hot days.

How to Calculate Correlation

There are three major ways to determine the correlation coefficient: spreadsheets, programming, and by hand, as shown above.

In Excel or Google Sheets

Microsoft Excel and Google Sheets both have a pre-programmed function called CORREL which determines the Pearson correlation coefficient automatically.

  1. Input the values for your first variable in one column.
  2. Input the values for your second variable in a column right next to it.
  3. In an empty cell, enter the formula =CORREL(range1, range2), replacing range1 and range2 with the two data ranges.
  4. Press Enter. The cell displays the correlation coefficient.

Using the worked example above, the formula =CORREL(A2:A11, B2:B11) returns 0.99.

In Python (NumPy and pandas)

Python calculates correlation coefficients through the NumPy and pandas libraries.

Using NumPy:

python
import numpy as np

temperature = np.array([68, 71, 74, 79, 82, 85, 88, 90, 93, 95])
sales = np.array([120, 135, 150, 170, 190, 205, 230, 255, 270, 290])

correlation_matrix = np.corrcoef(temperature, sales)
print("Correlation Coefficient:", correlation_matrix[0, 1])
Output: Correlation Coefficient: 0.99
Using pandas:
python
import pandas as pd

data = pd.DataFrame({
    'Temperature': [68, 71, 74, 79, 82, 85, 88, 90, 93, 95],
    'Sales': [120, 135, 150, 170, 190, 205, 230, 255, 270, 290]
})

correlation = data['Temperature'].corr(data['Sales'])
print("Correlation Coefficient:", correlation)
Output: Correlation Coefficient: 0.99

The pandas .corr() method also generates a correlation matrix when applied to a full DataFrame with more than two columns. A correlation matrix displays the correlation coefficient between every pair of variables in a dataset, arranged in a grid. This makes it possible to identify relationships across many variables at once, rather than testing pairs individually.

Correlation vs. Regression vs. Covariance

Correlation, regression, and covariance each describe relationships between variables, but they answer different questions.

MethodWhat It MeasuresOutput RangePrimary Use
CorrelationStrength and direction of a linear relationship-1 to +1Determine whether and how strongly two variables are related
RegressionThe mathematical equation that predicts one variable from anotherNot boundedPredict the value of a dependent variable based on an independent variable
CovarianceThe direction of the relationship between two variablesNot bounded, depends on measurement unitsDetermine whether two variables increase or decrease together, before standardizing to correlation

Both covariance and correlation are measurements that point to the direction of the relationship. Covariance does not indicate strength since the measurement is relative, based on units of the variables measured. On the other hand, correlation normalizes the measurement by dividing the covariance by the standard deviation of both variables.

Regression analysis goes beyond correlation analysis through generating an equation for predicting one variable from the other. Correlation simply indicates the existence of a relationship and how strong it is. Regression on the other hand gives the change in one variable when another is changed.

Why Correlation Isn’t Causation

A relationship between two variables does not mean that one causes the other. There could be a third unknown variable causing changes in both variables.

Example: Sales of ice cream and cases of drowning have shown a high positive correlation in the summer season. The ice cream is not causing drowning. The reason is that there is a third variable which causes an increase in both variables, which is warm temperature leading to more swimming and more ice cream consumption. Such a relationship is referred to as a spurious correlation.

There are three requirements to prove that the relationship is a causal relationship, first being cause preceding the effect, second being correlation between the variables and the third being excluding other possibilities. Correlation analysis fulfills only the second requirement.

Common Mistakes When Interpreting Correlation

Professional Certificate

Data Analyst Course

Get on the fast track to a career in data analytics. Learn SQL, Python, Power BI and Excel with AI-assisted workflows, and build the skills employers screen for — no degree or prior coding experience required.

4.8 (18,340 ratings)  •  46,210 already enrolled  •  Beginner level

Class Starts on 3 Oct, 2026 — SAT & SUN (Weekend Batch)

Average time: 6 months  

Skills you’ll build: SQL, Python for Data Analysis, Power BI, Excel with AI, Data Storytelling, Stakeholder Reporting

There are four common mistakes made by researchers when interpreting correlation coefficients.

  1. Linearity assumption. The Pearson correlation assumes linearity in the relationship between two variables. It is possible for two variables to have a high degree of curved relationship but a Pearson r close to 0.
  2. Outlier problem. One outlier can distort or enhance the correlation coefficient. Checking outliers will determine their effect.
  3. Sample size problem. Correlation estimates with small samples are unreliable. There is a lot of difference between a correlation estimate with five data points compared to one with 500 data points.
  4. Not testing for statistical significance. Correlation coefficient alone does not prove the existence of a relationship, just that there is a relationship. Statistical significance testing proves that there is a significant relationship that is not by chance. It usually involves a t-test on the correlation coefficient.

Real-World Applications of Correlation Analysis

What Is Correlation Analysis

Correlation is employed in a number of different areas in order to discover the association between different variables prior to undertaking deeper analysis.

  • Healthcare: Scientists analyze the correlation between some treatment method and its outcome, for example, the relationship between the dose of blood pressure medicine and the results of blood pressure readings.
  • Finance: Traders analyze the correlation between prices of different assets, for example, two shares, or a share and a rate of interest.
  • Marketing: Specialists analyze the correlation between customer satisfaction and repurchase rates to understand which variables affect retention.
  • Human resources: Human resources department analyzes the correlation between engagement scores of employees and their productivity in order to assess company policies.
  • Machine learning: Specialists use correlation matrices when selecting features in order to understand which variables correlate best with target one and to diagnose multicollinearity, which means that two or more input variables are correlated.

Advantages and Limitations of Correlation Analysis

Advantages

  • Generates one, consistent value (r) that is easy to compute and interpret.
  • Does not require manipulation of variables, making it applicable in cases where experimentation is impractical or unethical.
  • Acts as an initial step in uncovering relationships worth examining further.
  • Aids in selecting features for machine learning through ranking of variables based on their relation to a target variable.

Limitations

  • Does not prove cause-and-effect relationship.
  • Can recognize only linear (Pearson) and monotonic (Spearman, Kendall) relations, ignoring other types of connections.
  • Is vulnerable to the presence of outliers.
  • Does not yield valid outcomes in case of small samples.
  • Is not capable of taking into account the impact of additional variables.

FAQs

Q1. What is a good correlation coefficient? 

Ans. The higher the absolute value of a correlation coefficient exceeds 0.70, the stronger it is; however, acceptable levels depend on the specific domain and must be viewed in combination with the sample size and level of statistical significance.

Q2. Can a correlation coefficient be negative? 

Ans. Certainly, a correlation coefficient varies between -1 and +1 and a negative correlation coefficient means that the larger the first variable is, the smaller the second one is and vice versa.

Q3. What is the difference between correlation and covariance? 

Ans. While covariance reveals the direction of the relation in the units of the variables, correlation makes that measure standardized and unitless within the range from -1 to +1.

Q4. Do I need a large sample size for correlation analysis? 

Ans. Indeed, the larger the sample size, the more reliable the estimate is; in contrast, a small sample size entails the risk of obtaining a coefficient which does not reflect the real situation.

Q5. What tools can I use to run correlation analysis? 

Ans. There are various methods to conduct correlation analysis including Excel/Google Sheets with the use of the CORREL function, Python programming language through NumPy/pandas or specialized software, such as R/SPSS.

Shalki Aggarwal is a Software Engineer II at Microsoft and an AI & Data Science expert specializing in Generative AI, Agentic AI, Python, LangChain, LangGraph, CrewAI, Deep Agents, and Loop Engineering. She is also a corporate trainer for leading organizations including L&T, Bharat Petroleum, Luminous, Denso, and Toshiba Midea, helping teams apply AI and emerging technologies to real-world business challenges.