Feature Scaling Before DBSCAN · Why Standardisation Changes Your Neighbourhoods
What eps measures
The scikit-learn API defines eps as the maximum distance between two samples for one to belong to the other’s neighbourhood. With the default metric='euclidean', that distance is computed directly from the feature array passed to DBSCAN. The API default is eps=0.5, but the user guide states that eps must be chosen for the dataset and distance function and usually should not be left at its default.12
For a point p, its neighbourhood can be written as:
N_eps(p) = { q : dist(p, q) <= eps }
A sample satisfying DBSCAN’s min_samples condition is a core sample. Every core sample belongs to a cluster. A non-core sample close to a core sample is assigned to that cluster; a sample outside every core sample’s eps-neighbourhood is labelled as noise with -1.2
The radius is expressed in the coordinates of the data being clustered. If preprocessing changes a feature from its raw value x to a transformed value z, the numerical distance between samples can change even when the rows and their order remain the same. A fixed eps can therefore include different neighbours before and after scaling.
That distinction is operational, not cosmetic:
- A small
epsleaves most samples outside sufficient neighbourhoods, producing mostly-1labels. - A large
epsjoins nearby clusters and can eventually return the entire dataset as one cluster.2 - Changing
min_sampleschanges the density requirement, but it does not giveepsunit-independent meaning. - Changing feature units or scaling changes the distance coordinate system in which
epsis evaluated.
Treat eps as coupled to the scaler, feature units, and distance metric. A numerical value tuned after one transformation is not automatically meaningful in another coordinate system.
A reproducible change in cluster count
The following results come from a local reproduction, not from a quoted documentation example. The run used scikit-learn 1.9.1 and NumPy 2.5.3 on 2026-10-05. It generated 300 observations with 3 ground-truth clusters, 3 features, and cluster_std=0.5, then multiplied the third feature by 1000.
from sklearn.cluster import DBSCAN
from sklearn.datasets import make_blobs
from sklearn.preprocessing import MinMaxScaler, RobustScaler, StandardScaler
X, labels_true = make_blobs(
n_samples=300,
centers=3,
n_features=3,
cluster_std=0.5,
random_state=0,
)
X[:, 2] *= 1000
def summarize(values, eps):
model = DBSCAN(eps=eps, min_samples=5).fit(values)
labels = model.labels_
n_clusters = len(set(labels)) - (1 if -1 in labels else 0)
n_noise = list(labels).count(-1)
return n_clusters, n_noise
cases = [
("unscaled", X, (0.5, 5, 100, 400)),
("StandardScaler", StandardScaler().fit_transform(X), (0.5, 1.0, 1.5)),
("MinMaxScaler", MinMaxScaler().fit_transform(X), (0.1, 0.2)),
("RobustScaler", RobustScaler().fit_transform(X), (0.5, 1.0)),
]
for name, values, eps_values in cases:
for eps in eps_values:
clusters, noise = summarize(values, eps)
print(name, eps, clusters, noise)
The exact locally reproduced results were:
| Input | eps |
min_samples |
Clusters found | Noise |
|---|---|---|---|---|
Unscaled, third feature × 1000 |
0.5 |
5 |
0 |
300 |
| Unscaled | 5 |
5 |
1 |
294 |
| Unscaled | 100 |
5 |
3 |
12 |
| Unscaled | 400 |
5 |
2 |
2 |
StandardScaler |
0.5 |
5 |
3 |
2 |
StandardScaler |
1.0 |
5 |
3 |
0 |
StandardScaler |
1.5 |
5 |
2 |
0 |
MinMaxScaler |
0.1 |
5 |
3 |
4 |
MinMaxScaler |
0.2 |
5 |
3 |
0 |
RobustScaler |
0.5 |
5 |
2 |
0 |
RobustScaler |
1.0 |
5 |
1 |
0 |
The direct comparison is the first StandardScaler row against the unscaled row with the same eps=0.5 and min_samples=5. Without scaling, DBSCAN finds 0 clusters and labels all 300 observations as noise. After standardisation, the same parameter values produce 3 clusters and 2 noise observations.
The generated dataset has not been replaced with another dataset; its coordinate representation has been transformed. Because core samples are the mechanism through which DBSCAN converts dense neighbourhoods into clusters, the 0-to-3 cluster-count change demonstrates a change in the density structure, not a change in execution speed.
The remaining rows also show why an eps value cannot be transferred between representations. eps=0.5 produces 3 clusters after StandardScaler but 2 after RobustScaler. These are different radius spaces, so the outputs are not controlled comparisons of the scalers themselves.
Choosing between the scalers
StandardScaler removes each feature’s training-set mean and scales it to unit variance:
z = (x - u) / s
Here, u is the mean and s is the standard deviation of the training samples. The defaults are with_mean=True and with_std=True.3
The API reference explicitly states that StandardScaler is sensitive to outliers and that features may be scaled differently from each other in the presence of outliers.3 Use it when unit-variance coordinates are the intended basis for distance and the mean and standard deviation are sufficiently stable. Its output makes eps a radius in standard-score coordinates, not in the original feature units.
The same API reference documents an important sparse-matrix constraint: with the default centering enabled, StandardScaler raises an exception for sparse matrices because centering would construct a dense matrix that may be too large to fit in memory.3 The runnable DBSCAN example therefore uses a dense feature array.
RobustScaler removes the median and scales according to a quantile range. Its default quantile_range=(25.0, 75.0) is the interquartile range: the range between the 1st and 3rd quantiles. The documented parameter constraint is 0.0 < q_min < q_max < 100.0, and unit_variance=False by default.4
The scikit-learn documentation states that outliers can negatively affect the mean and variance, while the median and interquartile range often provide better results in those cases.4 Use RobustScaler when outlier-sensitive mean and variance would produce an unsuitable neighbourhood geometry. After this transformation, eps is evaluated in median-centred, quantile-range-scaled coordinates.
MinMaxScaler transforms each feature independently to a configured range. Its documented transformation is:
X_std = (X - X.min(axis=0)) / (X.max(axis=0) - X.min(axis=0))
X_scaled = X_std * (max - min) + min
The default feature_range=(0, 1) maps the training-set extrema to the configured endpoints.5 Use MinMaxScaler when a fixed range is required and the extreme observations are acceptable as the endpoints that define that range.
MinMaxScaler does not reduce the effect of outliers. The largest observed value becomes the maximum endpoint and the smallest becomes the minimum endpoint.5 It should not be selected as a way to contain outlier influence while preserving their relative effect. The default is clip=False; with clip=True, values outside the range are clipped and inverse_transform may not recover the original values.5
The cited sparse-matrix disclosures are asymmetric. The StandardScaler API explicitly documents the centering exception. The supplied RobustScaler and MinMaxScaler API material does not state a corresponding sparse-matrix restriction, so the StandardScaler exception should not be presented as a documented constraint of those estimators.
Use the following selection rule:
- Use
StandardScalerwhen raw feature spreads should be replaced by training-set standard deviations, and inspect its sensitivity to outliers. - Use
RobustScalerwhen outliers could distort the mean or variance and median/IQR scaling is appropriate. - Use
MinMaxScalerwhen a fixed feature range is required and the observed extrema are meaningful. - Keep raw coordinates when their existing units already encode the intended distance geometry;
epsis then interpreted in those raw units.
Whichever transformation is selected, compute the DBSCAN distances and select eps in the transformed space.
Fit, tune, and interpret one pipeline
The required order is scaler first, DBSCAN second. The official scikit-learn example uses the same sequence.6
scaler = StandardScaler()
X_for_dbscan = scaler.fit_transform(X)
model = DBSCAN(eps=0.5, min_samples=5).fit(X_for_dbscan)
labels = model.labels_
n_clusters = len(set(labels)) - (1 if -1 in labels else 0)
n_noise = list(labels).count(-1)
To use RobustScaler or MinMaxScaler, substitute that preprocessing estimator while leaving the transformed output as DBSCAN’s input. Do not select eps against the raw matrix and then apply a different transformation during clustering.
Apply these rules during parameter selection:
- Hold the distance metric and
min_samplesfixed while comparing candidateepsvalues. - Select
epsagainst the exact transformed matrix passed toDBSCAN.fit. - Treat excessive
-1labels as evidence that the radius is too small for the current representation. - Treat merging of distinct structures into one cluster as evidence that the radius is too large for the current representation.
- Record the scaler, its configuration, the feature representation, the metric, and
epsas one experimental configuration. - Repeat fitting and tuning after any change to feature units, scaling, or distance metric.
In the local reproduction, eps=0.5 means different things before and after standardisation: it produces 0 clusters in the raw coordinate system and 3 clusters in the standard-score coordinate system. The value is therefore meaningless outside the chosen scaling unless it is explicitly re-expressed or retuned. Changing a unit conversion or scaler changes DBSCAN’s neighbourhoods even if the model class, metric, eps, and min_samples remain unchanged.
Frequently asked questions
Is the scaler passed to DBSCAN?
No. StandardScaler, RobustScaler, and MinMaxScaler are separate preprocessing estimators, not arguments to DBSCAN.fit. Apply the chosen scaler first, then pass its output to DBSCAN.
Does min_samples include the current point?
Yes. The API defines min_samples as the number of samples, or total weight, in a neighbourhood and explicitly includes the point itself. A higher value requires a denser neighbourhood, while a lower value permits sparser clusters.1
Can fitted DBSCAN label unseen observations?
The scikit-learn overview states that transductive clustering methods are not designed for new, unseen data.7 Treat DBSCAN labels as a partition of the rows supplied to fit, not as predictions from an inductive labelling model.
Is eps a maximum distance inside a cluster?
No. The API explicitly states that eps is not a maximum bound on distances between points within a cluster. It is the maximum distance for two samples to count in each other’s local neighbourhood.1