अनुपर्वेक्षित अधिगम
उत्तर न होने पर संरचना को स्वयं डेटा से उभरना होता है — और ‘अच्छा’ की परिभाषा फिर से तय करनी पड़ती है
यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।
परिभाषा
Unsupervised learning finds structure in data that carries no labels. Instead of predicting a correct answer, it uncovers the data’s own regularities — which samples resemble each other (clustering), how many dimensions the data really occupies (dimensionality reduction), what distribution it follows (density estimation), and which samples deviate from the norm (anomaly detection).
सहज समझ
Picture tipping a crate of unlabelled groceries onto the floor and shelving them by resemblance: bottles with bottles, packets with packets. Nobody tells you how many categories there should be or what each is called, and the groups you form only mean something once you assign them meaning. That is exactly what makes unsupervised learning hard: with no answer key, what counts as a good grouping has to be redefined from scratch.
The geometry of dimensionality reduction: flatten high-dimensional data onto a plane while preserving who is near whom
A 2D projection after reduction (e.g. t-SNE / PCA): the data falls into tight clusters, with a few samples stranded in the ambiguous space between them
कार्यप्रणाली
- 01
Choose a representation
First decide what features or embeddings represent each sample. A distance metric — Euclidean, cosine — is meaningless before the representation is fixed, because “similar” depends entirely on how the coordinates are defined.
- 02
Define what “structure” means
Clustering minimises within-cluster distance and maximises between-cluster distance; dimensionality reduction preserves variance or neighbourhood relations; density estimation maximises the data likelihood. Different objectives give different answers.
- 03
Optimise and pick hyperparameters
k-means needs the number of clusters k, t-SNE needs a perplexity, PCA needs the retained dimension. There is no label to validate these “how many groups” choices, so they are judged indirectly through silhouette scores, the elbow method, or downstream performance.
- 04
Assign meaning
The algorithm hands back groups or coordinates, not meaning. What a cluster should be called, or what a reduced axis stands for, still needs domain knowledge to interpret — which is exactly why unsupervised conclusions are the hardest to verify.
उपयोग के क्षेत्र
- Customer segmentation: group users by behaviour for targeted operations
- Visualisation and exploration: compress high-dimensional embeddings to 2D with t-SNE/UMAP
- Anomaly detection: surface the few samples that deviate from the mainstream in fraud, faults or intrusions
- Representation pretraining: learn features without labels, then fine-tune downstream (the seed of self-supervision)
सामान्य भ्रांतियाँ
- Clusters are not “true classes”. k-means will always return k clusters even when the data has none, and a different seed or k can rearrange the grouping beyond recognition.
- Inter-cluster distances in t-SNE/UMAP are not trustworthy. They preserve local neighbourhoods well, but how far apart two clusters appear is often meaningless — do not conclude that two groups are “more alike”.
- No labels does not mean no bias. What counts as “normal” is defined by the training data, so groups that are rare there get systematically flagged as anomalous.
मुख्य शब्द
- Clustering
- Grouping samples by similarity (k-means, hierarchical clustering)
- Dimensionality reduction
- Compressing high-dimensional data to fewer dimensions while preserving structure (PCA, t-SNE, UMAP)
- Density estimation
- Estimating the probability distribution the data follows
- Anomaly detection
- Finding the few samples that deviate from the bulk distribution