Feature Scaling Before DBSCAN · Why Standardisation Changes Your Neighbourhoods

What eps measures

The scikit-learn API defines eps as the maximum distance between two samples for one to belong to the other’s neighbourhood. With the default metric='euclidean', that distance is computed directly from the feature array passed to DBSCAN. The API default is eps=0.5, but the user guide states that eps must be chosen for the dataset and distance function and usually should not be left at its default.12

For a point p, its neighbourhood can be written as:

N_eps(p) = { q : dist(p, q) <= eps }

A sample satisfying DBSCAN’s min_samples condition is a core sample. Every core sample belongs to a cluster. A non-core sample close to a core sample is assigned to that cluster; a sample outside every core sample’s eps-neighbourhood is labelled as noise with -1.2

The radius is expressed in the coordinates of the data being clustered. If preprocessing changes a feature from its raw value x to a transformed value z, the numerical distance between samples can change even when the rows and their order remain the same. A fixed eps can therefore include different neighbours before and after scaling.

That distinction is operational, not cosmetic:

  • A small eps leaves most samples outside sufficient neighbourhoods, producing mostly -1 labels.
  • A large eps joins nearby clusters and can eventually return the entire dataset as one cluster.2
  • Changing min_samples changes the density requirement, but it does not give eps unit-independent meaning.
  • Changing feature units or scaling changes the distance coordinate system in which eps is evaluated.

Treat eps as coupled to the scaler, feature units, and distance metric. A numerical value tuned after one transformation is not automatically meaningful in another coordinate system.

A reproducible change in cluster count

The following results come from a local reproduction, not from a quoted documentation example. The run used scikit-learn 1.9.1 and NumPy 2.5.3 on 2026-10-05. It generated 300 observations with 3 ground-truth clusters, 3 features, and cluster_std=0.5, then multiplied the third feature by 1000.

from sklearn.cluster import DBSCAN
from sklearn.datasets import make_blobs
from sklearn.preprocessing import MinMaxScaler, RobustScaler, StandardScaler

X, labels_true = make_blobs(
    n_samples=300,
    centers=3,
    n_features=3,
    cluster_std=0.5,
    random_state=0,
)
X[:, 2] *= 1000

def summarize(values, eps):
    model = DBSCAN(eps=eps, min_samples=5).fit(values)
    labels = model.labels_
    n_clusters = len(set(labels)) - (1 if -1 in labels else 0)
    n_noise = list(labels).count(-1)
    return n_clusters, n_noise

cases = [
    ("unscaled", X, (0.5, 5, 100, 400)),
    ("StandardScaler", StandardScaler().fit_transform(X), (0.5, 1.0, 1.5)),
    ("MinMaxScaler", MinMaxScaler().fit_transform(X), (0.1, 0.2)),
    ("RobustScaler", RobustScaler().fit_transform(X), (0.5, 1.0)),
]

for name, values, eps_values in cases:
    for eps in eps_values:
        clusters, noise = summarize(values, eps)
        print(name, eps, clusters, noise)

The exact locally reproduced results were:

Input eps min_samples Clusters found Noise
Unscaled, third feature × 1000 0.5 5 0 300
Unscaled 5 5 1 294
Unscaled 100 5 3 12
Unscaled 400 5 2 2
StandardScaler 0.5 5 3 2
StandardScaler 1.0 5 3 0
StandardScaler 1.5 5 2 0
MinMaxScaler 0.1 5 3 4
MinMaxScaler 0.2 5 3 0
RobustScaler 0.5 5 2 0
RobustScaler 1.0 5 1 0

The direct comparison is the first StandardScaler row against the unscaled row with the same eps=0.5 and min_samples=5. Without scaling, DBSCAN finds 0 clusters and labels all 300 observations as noise. After standardisation, the same parameter values produce 3 clusters and 2 noise observations.

The generated dataset has not been replaced with another dataset; its coordinate representation has been transformed. Because core samples are the mechanism through which DBSCAN converts dense neighbourhoods into clusters, the 0-to-3 cluster-count change demonstrates a change in the density structure, not a change in execution speed.

The remaining rows also show why an eps value cannot be transferred between representations. eps=0.5 produces 3 clusters after StandardScaler but 2 after RobustScaler. These are different radius spaces, so the outputs are not controlled comparisons of the scalers themselves.

Choosing between the scalers

StandardScaler removes each feature’s training-set mean and scales it to unit variance:

z = (x - u) / s

Here, u is the mean and s is the standard deviation of the training samples. The defaults are with_mean=True and with_std=True.3

The API reference explicitly states that StandardScaler is sensitive to outliers and that features may be scaled differently from each other in the presence of outliers.3 Use it when unit-variance coordinates are the intended basis for distance and the mean and standard deviation are sufficiently stable. Its output makes eps a radius in standard-score coordinates, not in the original feature units.

The same API reference documents an important sparse-matrix constraint: with the default centering enabled, StandardScaler raises an exception for sparse matrices because centering would construct a dense matrix that may be too large to fit in memory.3 The runnable DBSCAN example therefore uses a dense feature array.

RobustScaler removes the median and scales according to a quantile range. Its default quantile_range=(25.0, 75.0) is the interquartile range: the range between the 1st and 3rd quantiles. The documented parameter constraint is 0.0 < q_min < q_max < 100.0, and unit_variance=False by default.4

The scikit-learn documentation states that outliers can negatively affect the mean and variance, while the median and interquartile range often provide better results in those cases.4 Use RobustScaler when outlier-sensitive mean and variance would produce an unsuitable neighbourhood geometry. After this transformation, eps is evaluated in median-centred, quantile-range-scaled coordinates.

MinMaxScaler transforms each feature independently to a configured range. Its documented transformation is:

X_std = (X - X.min(axis=0)) / (X.max(axis=0) - X.min(axis=0))

X_scaled = X_std * (max - min) + min

The default feature_range=(0, 1) maps the training-set extrema to the configured endpoints.5 Use MinMaxScaler when a fixed range is required and the extreme observations are acceptable as the endpoints that define that range.

MinMaxScaler does not reduce the effect of outliers. The largest observed value becomes the maximum endpoint and the smallest becomes the minimum endpoint.5 It should not be selected as a way to contain outlier influence while preserving their relative effect. The default is clip=False; with clip=True, values outside the range are clipped and inverse_transform may not recover the original values.5

The cited sparse-matrix disclosures are asymmetric. The StandardScaler API explicitly documents the centering exception. The supplied RobustScaler and MinMaxScaler API material does not state a corresponding sparse-matrix restriction, so the StandardScaler exception should not be presented as a documented constraint of those estimators.

Use the following selection rule:

  • Use StandardScaler when raw feature spreads should be replaced by training-set standard deviations, and inspect its sensitivity to outliers.
  • Use RobustScaler when outliers could distort the mean or variance and median/IQR scaling is appropriate.
  • Use MinMaxScaler when a fixed feature range is required and the observed extrema are meaningful.
  • Keep raw coordinates when their existing units already encode the intended distance geometry; eps is then interpreted in those raw units.

Whichever transformation is selected, compute the DBSCAN distances and select eps in the transformed space.

Fit, tune, and interpret one pipeline

The required order is scaler first, DBSCAN second. The official scikit-learn example uses the same sequence.6

scaler = StandardScaler()
X_for_dbscan = scaler.fit_transform(X)

model = DBSCAN(eps=0.5, min_samples=5).fit(X_for_dbscan)
labels = model.labels_

n_clusters = len(set(labels)) - (1 if -1 in labels else 0)
n_noise = list(labels).count(-1)

To use RobustScaler or MinMaxScaler, substitute that preprocessing estimator while leaving the transformed output as DBSCAN’s input. Do not select eps against the raw matrix and then apply a different transformation during clustering.

Apply these rules during parameter selection:

  • Hold the distance metric and min_samples fixed while comparing candidate eps values.
  • Select eps against the exact transformed matrix passed to DBSCAN.fit.
  • Treat excessive -1 labels as evidence that the radius is too small for the current representation.
  • Treat merging of distinct structures into one cluster as evidence that the radius is too large for the current representation.
  • Record the scaler, its configuration, the feature representation, the metric, and eps as one experimental configuration.
  • Repeat fitting and tuning after any change to feature units, scaling, or distance metric.

In the local reproduction, eps=0.5 means different things before and after standardisation: it produces 0 clusters in the raw coordinate system and 3 clusters in the standard-score coordinate system. The value is therefore meaningless outside the chosen scaling unless it is explicitly re-expressed or retuned. Changing a unit conversion or scaler changes DBSCAN’s neighbourhoods even if the model class, metric, eps, and min_samples remain unchanged.

Frequently asked questions

Is the scaler passed to DBSCAN?

No. StandardScaler, RobustScaler, and MinMaxScaler are separate preprocessing estimators, not arguments to DBSCAN.fit. Apply the chosen scaler first, then pass its output to DBSCAN.

Does min_samples include the current point?

Yes. The API defines min_samples as the number of samples, or total weight, in a neighbourhood and explicitly includes the point itself. A higher value requires a denser neighbourhood, while a lower value permits sparser clusters.1

Can fitted DBSCAN label unseen observations?

The scikit-learn overview states that transductive clustering methods are not designed for new, unseen data.7 Treat DBSCAN labels as a partition of the rows supplied to fit, not as predictions from an inductive labelling model.

Is eps a maximum distance inside a cluster?

No. The API explicitly states that eps is not a maximum bound on distances between points within a cluster. It is the maximum distance for two samples to count in each other’s local neighbourhood.1


  1. scikit-learn, the DBSCAN API reference. ↩↩↩

  2. scikit-learn, User Guide section 2.3.7, DBSCAN. ↩↩↩

  3. scikit-learn, the StandardScaler API reference. ↩↩↩

  4. scikit-learn, the RobustScaler API reference. ↩↩

  5. scikit-learn, the MinMaxScaler API reference. ↩↩↩

  6. scikit-learn, the official DBSCAN clustering demo example. ↩

  7. scikit-learn, User Guide section 2.3.1, the clustering methods overview. ↩