Math
4 articles
Clustering with High-Dimensional Data · The Curse of Dimensionality
Why clustering fails in high dimensions — distance concentration, the emptiness phenomenon, and strategies for clustering high-dimensional data with DBSCAN.
DBSCAN Time Complexity · Naive O(n²) vs O(n log n) with Spatial Indexing
Understanding DBSCAN's computational complexity — the naive algorithm, the role of spatial indexing, worst-case vs average-case, and practical runtime expectations.
Clustering Validation Metrics · Silhouette, Davies-Bouldin, and Calinski-Harabasz
How to evaluate clustering quality — internal validation metrics (silhouette score, Davies-Bouldin index, Calinski-Harabasz index), when each is appropriate, and their limitations.
Distance Metrics for Clustering · Euclidean, Manhattan, Cosine, and More
How the choice of distance metric affects clustering results — Euclidean, Manhattan, Cosine, Mahalanobis, Haversine, and when to use each.