DBSCAN Benchmark Datasets · sklearn Clustering Datasets and What Each One Tests

eps is not a dataset-independent constant. The scikit-learn clustering comparison standardises each dataset before applying algorithm parameters, then deliberately includes cases where the parameter values do not match the data structure.5 A useful DBSCAN benchmark therefore records both successful configurations and the settings at which clusters fragment, merge, or become meaningless.

Built-in generators and canonical benchmark cases

The primary built-in scikit-learn generators are make_circles, make_moons, and make_blobs. The canonical comparison extends them with a varied-density transformation, a linear anisotropy transformation, and a null dataset.5

  • noisy_circles — make_circles generates a large circle containing a smaller circle in 2D. factor controls the scale of the inner relative to the outer circle and is documented in the range [0, 1); noise is the standard deviation of Gaussian noise. The benchmark uses n_samples=500, factor=0.5, noise=0.05, and random_state=30.2

  • noisy_moons — make_moons generates two interleaving half-circles. Here, noise is the standard deviation of Gaussian noise, and the benchmark uses n_samples=500, noise=0.05, and random_state=30.1

  • blobs — make_blobs generates isotropic Gaussian blobs. With an integer n_samples, the samples are divided equally among clusters; when centers=None, the API generates three centers. The baseline benchmark uses n_samples=500 and random_state=30.3

  • varied — This remains a make_blobs dataset but assigns cluster_std=[1.0, 2.5, 0.5] with random_state=170. It tests whether one eps can recover blobs with substantially different spreads rather than treating every Gaussian component as having the same density.35

  • aniso — The benchmark first generates blobs with n_samples=500 and random_state=170, then applies the linear transformation [[0.6, -0.6], [-0.4, 0.8]]. It tests whether an apparently Euclidean neighbourhood remains meaningful after the geometry has been made anisotropic.5

  • no_structure — The comparison generates RandomState(30).rand(500, 2) and assigns no labels. The scikit-learn comparison defines this as the null clustering case: the data is homogeneous, and there is no good clustering.5

Every case should be standardised independently with StandardScaler().fit_transform before the eps sweep. That protocol makes the resulting parameter values comparable across coordinate systems without changing the underlying constructions.

Measured failure boundaries

The scikit-learn user guide gives the basic parameter failure pattern: when eps is too small, most points remain unclustered and receive the noise label -1; when it is too large, nearby clusters merge and can eventually become one cluster.4 min_samples is a separate control. The local run below changes it while holding eps and the dataset fixed.

These results were reproduced locally with scikit-learn 1.9.1 and NumPy 2.5.3 on 2026-10-05; they are measured local results, not quotations from scikit-learn.

Dataset eps min_samples Clusters Noise Core
noisy_circles 0.3 5 2 0 498
noisy_circles 0.08 5 16 379 60
noisy_moons 0.3 5 2 0 500
noisy_moons 0.1 5 21 48 360
varied 0.3 5 3 16 464
varied 0.18 5 6 81 387
aniso 0.3 5 1 1 493
blobs 0.3 5 3 7 277
blobs 0.3 7 2 11 457
no_structure 0.3 5 1 0 498
no_structure 0.3 7 1 0 471

The cases expose distinct failures:

  • Non-flat geometry: At eps=0.3, both curved datasets recover two clusters with no noise. Reducing eps to 0.08 fragments circles into 16 clusters with 379 noise points; reducing it to 0.1 fragments moons into 21 clusters with 48 noise points. These datasets test whether DBSCAN follows curved generating components instead of imposing a linear partition, but they also show the low-eps failure documented by the user guide.45

  • Density variation: At eps=0.3, varied returns three clusters, 16 noise points, and 464 core points. At the comparison’s eps=0.18 override, the same setting over-segments the three blobs into six clusters with 81 noise points. A single global eps therefore does not represent the three specified spreads equally well.

  • Anisotropy: At eps=0.3, aniso collapses into one cluster with one noise point and 493 core points. Running the same configuration without standardisation produces the same cluster, noise, and core counts. At eps=0.15 the dataset fragments. With min_samples=7 it gives four clusters, 55 noise points and 391 core points. With min_samples=5 it gives six clusters, 27 noise points and 433 core points. No plain radius recovers the intended three blobs. The local interpretation is that no plain eps split recovers the transformed blobs: the geometry is an artefact of the linear transformation, and the reported remedy is dimensionality reduction rather than parameter tuning alone.

  • The min_samples axis: On baseline blobs, eps=0.3 with min_samples=5 produces three clusters, seven noise points, and 277 core points. Changing only min_samples to 7 produces two clusters and 11 noise points. The cluster count therefore changes along the min_samples axis even when eps and the input data remain fixed.

  • The null case: The homogeneous dataset has no good clustering, yet eps=0.3 returns one cluster and zero noise at both recorded min_samples settings. This is the parameter–structure mismatch highlighted by the canonical comparison: DBSCAN can return a confident cluster label even when the input has no meaningful cluster structure to recover.5

Runnable DBSCAN comparison

The following script constructs the same datasets, standardises each one, sweeps the measured eps values, varies min_samples, and reports cluster, noise, and core counts.

import numpy as np
from sklearn.cluster import DBSCAN
from sklearn.datasets import make_blobs, make_circles, make_moons
from sklearn.preprocessing import StandardScaler

n_samples = 500
seed = 30
aniso_seed = 170

noisy_circles = make_circles(
    n_samples=n_samples,
    factor=0.5,
    noise=0.05,
    random_state=seed,
)
noisy_moons = make_moons(
    n_samples=n_samples,
    noise=0.05,
    random_state=seed,
)
blobs = make_blobs(
    n_samples=n_samples,
    random_state=seed,
)

base_aniso, y_aniso = make_blobs(
    n_samples=n_samples,
    random_state=aniso_seed,
)
X_aniso = np.dot(
    base_aniso,
    [[0.6, -0.6], [-0.4, 0.8]],
)

varied, y_varied = make_blobs(
    n_samples=n_samples,
    cluster_std=[1.0, 2.5, 0.5],
    random_state=aniso_seed,
)

rng = np.random.RandomState(seed)
no_structure = rng.rand(n_samples, 2), None

datasets = {
    "noisy_circles": noisy_circles,
    "noisy_moons": noisy_moons,
    "varied": varied,
    "aniso": (X_aniso, y_aniso),
    "blobs": blobs,
    "no_structure": no_structure,
}

eps_grid = (0.08, 0.1, 0.15, 0.18, 0.3)
min_samples_grid = (5, 7)

print("dataset,eps,min_samples,clusters,noise,core")

for name, (X, _) in datasets.items():
    X_scaled = StandardScaler().fit_transform(X)

    for eps in eps_grid:
        for min_samples in min_samples_grid:
            model = DBSCAN(eps=eps, min_samples=min_samples)
            labels = model.fit_predict(X_scaled)

            clusters = len(set(labels) - {-1})
            noise = np.count_nonzero(labels == -1)
            core = len(model.core_sample_indices_)

            print(
                f"{name},{eps},{min_samples},"
                f"{clusters},{noise},{core}"
            )

The printed grid is broader than the selected checkpoints in the results table. This makes the two transitions visible: decreasing eps can replace one cluster with many small clusters and noise, while increasing eps can merge distinct components. On the null dataset, neither transition should be interpreted as evidence of a meaningful target partition.

Ground-truth and internal evaluation

The scikit-learn DBSCAN demonstration uses make_blobs because the generator exposes the true synthetic cluster labels as labels_true.6 When those labels are available, the demonstration evaluates the unsupervised result with homogeneity, completeness, V-measure, Rand index, Adjusted Rand Index, and Adjusted Mutual Information.

from sklearn.cluster import DBSCAN
from sklearn.datasets import make_blobs
from sklearn.preprocessing import StandardScaler

centers = [[1, 1], [-1, -1], [1, -1]]
X, labels_true = make_blobs(
    n_samples=750,
    centers=centers,
    cluster_std=0.4,
    random_state=0,
)
X = StandardScaler().fit_transform(X)

db = DBSCAN(eps=0.3, min_samples=10).fit(X)
labels = db.labels_

The demonstration documents the following output:6

Clusters Noise Homogeneity Completeness V-measure Adjusted Rand Index Adjusted Mutual Information Silhouette
3 18 0.953 0.883 0.917 0.952 0.916 0.626

Ground-truth metrics compare the generated labels with the fitted cluster labels; they do not turn clustering into a supervised fitting procedure. They are appropriate only when the generated membership is the intended benchmark target.

When ground truth is unavailable, the scikit-learn demonstration states that evaluation must use the model results themselves and identifies the Silhouette Coefficient as an option.6 The no_structure construction supplies y=None, so an Adjusted Rand Index or Adjusted Mutual Information cannot be computed against generator labels for that case. Its cluster and noise counts remain useful diagnostics, but a nonzero Silhouette value would not supply the missing cluster truth.

Frequently asked questions

What does n_samples mean for make_moons?

n_samples is either an integer giving the total number of points or a two-element tuple giving the number of points in each moon. The tuple form was added in version 0.23.1

How does make_circles allocate an odd sample count?

For an odd n_samples, the inner circle receives one more point than the outer circle. The benchmark construction above sets n_samples=500.2

Why does no_structure inherit eps=0.15 in the canonical example but use eps=0.3 in the local results?

The canonical comparison gives no_structure an empty per-dataset override, so it inherits the previous row’s eps=0.15. The local table comes from a separate DBSCAN-only run and reports no_structure at eps=0.3; those are distinct measured configurations.5

Do these 2D results generalise to very high-dimensional data?

No. The scikit-learn comparison explicitly warns that intuition from its 2D examples might not apply to very high-dimensional data.5


  1. scikit-learn, the sklearn.datasets.make_moons API reference. ↩↩

  2. scikit-learn, the sklearn.datasets.make_circles API reference. ↩↩

  3. scikit-learn, the sklearn.datasets.make_blobs API reference. ↩↩

  4. scikit-learn, User Guide section 2.3.7, DBSCAN. ↩↩

  5. scikit-learn, the official toy-dataset clustering comparison example. ↩↩↩↩↩↩↩↩↩

  6. scikit-learn, the official DBSCAN clustering demo example. ↩↩↩