Density Reachability · Core Points, Border Points and the ε-Neighbourhood Definition
The ε-neighbourhood and its distance function
Let X be the dataset, p a sample in X, and d(p, q) the distance between samples p and q. The ε-neighbourhood of p is:
N_eps(p) = { q ∈ X : d(p, q) <= eps }
The comparison is non-strict: a sample at distance exactly eps is in the neighbourhood. Nothing in this definition requires Euclidean distance. The function d can be replaced by another metric selected through DBSCAN’s metric parameter, provided the resulting distances are used consistently.2
This separation matters because eps is interpreted in the units of d. With one distance function, eps=0.3 defines one set of neighbours; changing the distance function while keeping eps unchanged can change those neighbours and therefore change which samples are core, border, or noise.
The scikit-learn API defines eps as the maximum distance within which two samples are considered neighbours. It explicitly does not impose a maximum distance between every pair of members of a cluster. A cluster can extend farther than eps when successive core samples provide a chain of local neighbourhoods.2
Core, border, and noise samples
For an unweighted dataset, p is a core sample when:
|N_eps(p)| >= min_samples
The neighbourhood count includes p itself. Thus, with min_samples=10, ten samples must be in N_eps(p), including p; the requirement is not for p plus ten other samples. The scikit-learn API makes this counting convention explicit, while the user guide describes the resulting condition as evidence that the sample lies in a dense region.12
A border sample, also called a non-core or fringe sample, is a sample that:
- is not core under the current
epsandmin_samples; and - belongs to the ε-neighbourhood of a core sample in the cluster.
Formally, for a cluster with core set C, its border set is:
B(C) = { x ∈ X \ C : x ∈ N_eps(c) for at least one c ∈ C }
Border membership therefore depends on direct membership in a core sample’s ε-neighbourhood. Merely being close to another non-core sample does not make a sample a border sample.
A noise sample is a non-core sample that is not in the ε-neighbourhood of any core sample:
Noise = { x ∈ X : x is not core and N_eps(x) contains no core sample }
The scikit-learn user guide states the rule in prose: “Any sample that is not a core sample, and is at least eps in distance from any core sample, is considered an outlier by the algorithm.”1 For boundary arithmetic, the neighbourhood expression above is unambiguous: because N_eps uses d(p, q) <= eps, a distance equal to eps is inside the closed neighbourhood.
The three tiers partition the dataset. Core samples provide the dense structure, border samples attach to that structure without satisfying the core test, and noise samples satisfy neither the core test nor attachment to a core sample.
Density-connected expansion
DBSCAN does not build a cluster by comparing every sample with every other sample. It starts from a core sample and expands only through core samples.
For cluster construction, the core backbone can be written as a density-connectivity relation:
p_start ↝ p_target
when a finite sequence of core samples connects them and every consecutive pair a, b satisfies d(a, b) <= eps. The sequence need not be a direct path in the input data; it is a path in the graph whose edges are ε-neighbourhood relationships and whose traversal is restricted to core samples.1
An expansion can be expressed as follows:
core_set = {seed}
frontier = [seed]
while frontier:
p = frontier.pop()
for q in N_eps(p):
if is_core(q) and q not in core_set:
core_set.add(q)
frontier.append(q)
The frontier contains core samples whose neighbourhoods still need to be processed. Expanding p discovers its core neighbours. Each newly discovered core sample is then expanded in turn, so the search follows the recursive core-to-core relation described in the scikit-learn user guide.1
After the frontier is empty, the cluster’s border set is the union of the ε-neighbourhoods of its core samples, restricted to non-core samples. Border samples are not added to the frontier and therefore do not extend the core chain. This is the operational distinction between a fringe sample and a core sample.
Termination occurs when the frontier contains no unprocessed core sample. Every discovered core sample is recorded, so it is not added again; expansion stops along that connected core backbone. Any core sample is part of a cluster by definition. A core sample not reached from the current seed can seed another cluster, while a non-core sample with no core neighbour remains noise.1
Because expansion is recursive, the diameter of the resulting cluster is not bounded by eps. Two distant samples can receive the same label even when their direct distance exceeds eps, provided a chain of core-sample neighbourhoods connects them.2
Parameter semantics and failure modes
The two parameters define the required local density:1
min_samplesis the minimum neighbourhood count, including the sample itself, required for core status.epsis the maximum distance for membership in a sample’s neighbourhood.
Increasing min_samples or decreasing eps makes the core requirement stricter. A sample that was core can become non-core, and a non-core sample can then become border if it still neighbours a remaining core sample, or noise if it does not. If no sample satisfies the core test, no cluster can be initialized because expansion requires a core seed.
Decreasing min_samples or increasing eps admits a less restrictive neighbourhood. More samples may satisfy the core test, and the API documents lower min_samples as producing sparser clusters and higher min_samples as finding denser clusters.2
Changing metric changes the neighbourhood relation itself. A sample’s core status is therefore conditional on the distance function, not only on eps and min_samples. A distance matrix or another supported metric can represent a domain in which coordinatewise Euclidean distance is unsuitable, but its scale and boundary semantics must remain consistent with the chosen eps.2
A common interpretation error is to treat eps as a cluster-diameter limit. It is not. It controls each local edge, while recursive core expansion determines cluster membership. Consequently, increasing eps can affect both local core status and the paths through which clusters are built; it does not impose a global radius on the completed clusters.2
Worked three-tier classification
The following example uses the dataset and parameters from the scikit-learn DBSCAN demo. It additionally counts border samples by subtracting noise and core samples from the total.
import numpy as np
from sklearn.cluster import DBSCAN
from sklearn.datasets import make_blobs
from sklearn.preprocessing import StandardScaler
centers = [[1, 1], [-1, -1], [1, -1]]
X, _ = make_blobs(
n_samples=750,
centers=centers,
cluster_std=0.4,
random_state=0,
)
X = StandardScaler().fit_transform(X)
db = DBSCAN(eps=0.3, min_samples=10).fit(X)
labels = db.labels_
n_clusters = len(set(labels)) - (1 if -1 in labels else 0)
n_noise = list(labels).count(-1)
n_core = len(db.core_sample_indices_)
n_border = len(labels) - n_noise - n_core
print("clusters =", n_clusters)
print("noise =", n_noise)
print("core =", n_core)
print("border =", n_border)
The output is:
clusters = 3
noise = 18
core = 679
border = 53
The documented demo reports an estimated cluster count of 3 and an estimated noise count of 18.3 The complete three-tier result above was reproduced locally on this machine with scikit-learn 1.9.1 on 2026-10-05; the values 679 core samples and 53 border samples are local results, not quotations from the demo.
The border count follows directly from the partition:
750 total samples - 18 noise samples - 679 core samples = 53 border samples
The 18 noise samples satisfy neither the core test nor border attachment. The 53 border samples are non-core samples attached to a cluster core. The remaining 679 samples satisfy the neighbourhood count of at least min_samples=10, including themselves, and supply the core backbones from which the 3 clusters are expanded.
Frequently asked questions
How does sample_weight affect core status?
sample_weight changes the core test from an unweighted sample count to a total-weight comparison. A sample with a weight of at least min_samples is core by itself, while a sample with a negative weight may inhibit its ε-neighbour from becoming core. Weights are absolute and default to 1.2
Which fitted attributes expose the DBSCAN result?
core_sample_indices_ contains the indices of the core samples. components_ contains a copy of each core sample found during fitting, and labels_ contains one cluster label per fitted sample. In labels_, -1 denotes noise and non-negative integers denote cluster membership.2
What does p mean for a Minkowski metric?
p specifies the power used for the Minkowski distance calculation. If p=None, Minkowski distance uses p=2, which is equivalent to Euclidean distance; p=1 is equivalent to Manhattan distance.2
What changes when metric="precomputed" is used?
With metric="precomputed", the input X must be a square distance matrix. It may be a sparse graph, but DBSCAN considers only its nonzero elements when determining neighbours.2