Unsupervised Learning

Unsupervised learning fits structure from without observed target labels. The output may be clusters, components, density estimates, embeddings, or anomaly scores. Because there is no direct , validation depends more heavily on assumptions and downstream utility than in supervised learning.

Structure without labels

Without labels, the method’s definition of “structure” is the whole game. Variance, Euclidean distance, density, and neighborhood preservation can all disagree, and dimensionality reduction that helps visualization may discard information a later classifier needs. The main families differ in what they output and what they optimize:

TaskOutputTypical objective
Clusteringgroup assignmentswithin-cluster distance (k-means)
Dimensionality reductionlow-dimensional coordinatesretained variance (PCA)
Density estimationa probability modeldata likelihood
Anomaly detectionoutlier scoresa score threshold

Common objectives

Many unsupervised methods optimize a reconstruction, partition, or likelihood objective. Clustering with k-means partitions the points into clusters to minimize the within-cluster squared distance,

where is a data point and is the centroid (mean) of the points assigned to cluster . PCA solves for an orthonormal projection , and anomaly methods learn a score and flag points whose score crosses a threshold.

When reference labels are available after clustering, the adjusted Rand index compares how pairs of points are grouped while correcting for chance:

where is the Rand index — the fraction of point pairs that are either together in both the clustering and the reference, or apart in both — and is its expected value under random labeling.

It is an external validation score, not an unsupervised objective.

Worked example

This snippet clusters standardized Iris measurements without labels, then compares internal silhouette quality with adjusted Rand agreement against the hidden species labels.

from sklearn.cluster import KMeans
from sklearn.datasets import load_iris
from sklearn.metrics import adjusted_rand_score, silhouette_score
from sklearn.preprocessing import StandardScaler
 
X, y = load_iris(return_X_y=True)
X2 = StandardScaler().fit_transform(X)
labels = KMeans(n_clusters=3, random_state=19, n_init=10).fit_predict(X2)
print("silhouette", round(silhouette_score(X2, labels), 3))
print("adjusted_rand_vs_species", round(adjusted_rand_score(y, labels), 3))

Observed output:

silhouette 0.46
adjusted_rand_vs_species 0.62

The silhouette score uses only geometry. Adjusted Rand uses species labels after the fact, showing that geometric clusters partially but imperfectly align with botanical classes.

Caveats

Scaling can dominate unsupervised results. Internal metrics can reward artificial structure even when clusters are not actionable. If labels are later used to choose the unsupervised representation, that choice becomes supervised model selection.

References