What Is an Outlier? How Does It Affect Data Analysis and Should It Be Removed?
- Data Investigator Team

- 5 days ago
- 6 min read

When analysing research data, researchers may encounter observations that appear noticeably different from most of the dataset. For example, satisfaction scores may generally fall within a similar range, while one observation appears unusually high or low. In income data, most respondents may report comparable earnings, while a small number report values several times higher than the rest.
These unusually distant observations are commonly referred to as outliers. Identifying outliers is an important part of checking a dataset before statistical analysis.
However, finding an outlier does not automatically mean that the observation should be removed. Some outliers result from data entry or coding errors, while others are genuine observations that reflect real characteristics of the sample. The decision to retain, correct or remove an outlier should therefore be based on its cause and its potential effect on the analysis.
What Is an Outlier?
An outlier is an observation that differs substantially from most other observations in a dataset. It may occur within a single variable or arise from the relationship between multiple variables.
Consider the following ages:
25, 27, 29, 30, 31, 32, 98
The value 98 appears very different from the rest, but this alone does not prove that it is incorrect. If the target population can reasonably include a 98-year-old participant, the value may be completely valid.
On the other hand, if the study specifically includes university students aged 18–25, a recorded age of 98 would clearly require further investigation. It might have resulted from a data entry or coding error.
This is why checking data before statistical analysis should involve more than identifying values that simply look unusually high or low. Researchers should also consider whether those values are plausible within the context of the study.
What Causes Outliers?
Outliers can occur for several reasons, and understanding their origin is important when deciding how to handle them.
A common cause is a data entry error. For example, an age of 25 might accidentally be entered as 255, or a variable that should contain values between 1 and 5 might contain a value of 55.
Coding problems can produce similar issues. A researcher might use 99 to represent “no response” but fail to define 99 as a Missing Value in the statistical software. The program could then treat 99 as an actual observation and include it in calculations.
However, an outlier is not necessarily an error. Some observations are genuinely very different from the rest of the sample. Examples might include an exceptionally high income, an unusually long treatment duration, or a customer whose spending is substantially higher than that of other customers.
Therefore, the first question when an outlier is detected should not be “Should I delete it?” but rather:
Why is this observation different from the others?
How Do Outliers Affect Data Analysis?
Outliers can influence several statistical measures, particularly those that are sensitive to extremely high or low values.
A simple example is the mean.
Suppose the monthly incomes of five respondents are:
30,000 / 32,000 / 35,000 / 38,000 / 200,000
The value of 200,000 will substantially increase the mean even though most respondents earn between 30,000 and 38,000.
In this situation, the median may provide a very different picture of the centre of the distribution.
Outliers can also affect the standard deviation, correlation and regression, as well as other statistical analyses. In some cases, an influential observation can noticeably alter estimated relationships or the fitted regression line.
The extent of the effect depends on factors such as the location and magnitude of the outlier, sample size, the number of unusual observations and the statistical method being used.
The presence of an outlier therefore does not automatically invalidate an analysis. What matters is whether it meaningfully influences the results.
How Can You Detect Outliers?
There are several ways to detect outliers, and researchers should generally avoid relying on a single method without considering the context of the data.
A useful starting point is to examine the distribution visually using a Boxplot, Histogram or Scatterplot. These can help identify observations that appear noticeably separated from the majority of the data.
For quantitative variables, researchers may also examine standardised values such as Z-scores to identify observations located unusually far from the mean.
Another commonly used approach is the Interquartile Range (IQR).
The IQR is calculated as:
IQR = Q3 − Q1
Values below:
Q1 − 1.5(IQR)
or above:
Q3 + 1.5(IQR)
may be flagged as potential outliers requiring further examination.
Importantly, these thresholds should be treated as tools for identifying observations that warrant investigation, rather than automatic rules for deleting data.
Should Outliers Be Removed Before Data Analysis?
Not necessarily. Outliers should not be removed automatically.
This is one of the most important principles when handling outliers.
If an unusual value results from an obvious data entry error—for example, 25 being entered as 255—the appropriate response is to check the original data and correct the error where possible, rather than simply deleting the entire case.
If the outlier represents a genuine observation, removing it should have a defensible research or methodological justification.
Researchers should not remove observations simply because doing so improves a statistical result, increases a correlation, or changes a P-value to the desired level.
In some situations, it may be useful to compare the analysis with and without an influential outlier to understand how sensitive the findings are to that observation. Such comparisons can help determine whether a small number of cases are materially affecting the conclusions.
The more appropriate question is therefore not simply:
“Should this outlier be removed?”
but:
“Why is this observation an outlier, and how does retaining or removing it affect the analysis?”
Do Not Confuse an Outlier with Incorrect Data
An outlier and a data error are not the same thing.
Suppose a questionnaire item is measured on a scale from 1 to 5 but the dataset contains a value of 8. This value is outside the permitted range and the original response should be checked.
In contrast, if income has no predefined upper limit and one respondent reports an income substantially higher than everyone else, the value may be an outlier but still be entirely valid.
This distinction matters. If every unusual observation is treated as an error and removed, researchers may inadvertently eliminate genuine characteristics of the population being studied.
For larger datasets, checking data accuracy, out-of-range values and unusual observations before analysis can help distinguish between information that needs to be corrected and valid observations that require further statistical consideration.
How Can Outliers Be Checked in SPSS?
SPSS provides several ways to explore potential outliers, including Frequencies, Descriptives, Explore, Boxplots and standardised values or Z-scores.
For regression analysis, additional diagnostic measures can help identify observations that may have an unusually strong influence on the model. These can include standardised residuals, Cook’s Distance and leverage.
However, a case appearing as an outlier on a Boxplot or having a large Z-score does not automatically mean it should be deleted. The next step should still involve examining the observation, checking its validity and considering its relevance to the statistical analysis being performed.
This is why SPSS statistical data analysis should involve more than simply selecting a statistical test. Data checking and relevant statistical assumptions should also be considered before hypothesis testing begins.
How Should Outliers Be Handled?
There is no single method for handling every outlier. The appropriate approach depends on why the unusual observation occurred and the characteristics of the data.
If the outlier results from a data entry error, the original source should be checked and the value corrected where possible. If the problem results from an incorrectly defined Missing Value or coding scheme, the coding should be corrected.
If the observation is valid but unusually high or low, researchers may need to consider how sensitive the selected statistical method is to outliers. Depending on the research design and data characteristics, approaches such as data transformation or alternative statistical methods may sometimes be appropriate.
In some studies, excluding an outlier may be justified. However, there should be a clear and defensible criterion for doing so. Decisions should not be made simply after seeing whether removing particular observations produces more favourable statistical results.
Check Outliers Before Hypothesis Testing
Outliers are only one of several issues that should be examined before formal statistical analysis begins. A dataset may also contain Missing Data, coding errors, values outside the expected range or problems relating to the distribution of variables.
Identifying these issues early can reduce the risk that relatively small data problems influence later statistical results.
For research requiring support from data checking and statistical test selection through to hypothesis testing and interpretation, Data Investigator’s SPSS statistical data analysis service can help develop an analysis process appropriate to the research objectives, variables and characteristics of the dataset. Outliers can then be considered as part of the overall data-checking process rather than automatically being removed simply because they appear different.
Conclusion
An outlier is an observation that differs substantially from most of the data and can potentially affect the mean, variability, correlation, regression and other statistical results.
However, an outlier is not necessarily incorrect data, and detecting an outlier does not automatically mean that it should be removed before analysis.
Researchers should first determine whether the observation resulted from an error or represents a genuine value, assess how strongly it influences the analysis, and select an appropriate method for handling it.
The purpose of checking for outliers is not to make a dataset look cleaner or to produce the expected statistical result. It is to ensure that the analysis reflects the data appropriately and that decisions about unusual observations can be clearly justified.
For more information, please kindly contact:
E-mail: info@datainvestigatorth.com
Line Official Account: @datainvestigator Tel: 063-969-7944

_edited_ed.png)


Comments