HomeSubjectsUniversityBlogAbout

Unsupervised Learning

Topic in AI / Machine Learning & Data Analytics

210 total MCQsShowing 30 with explanations10 Easy10 Medium10 Hard

About This Topic

Unsupervised learning finds hidden structure in data that has no labels, such as natural groups, compact representations or unusual points. Clustering dominates: K-Means and its centroids, the within-cluster sum of squares objective it minimizes, the elbow method and silhouette score for choosing k, K-Medoids, hierarchical clustering with dendrograms, and density-based DBSCAN. Dimensionality reduction questions cover PCA and explained variance, and t-SNE or UMAP for visualizing high-dimensional data. Other items touch Gaussian mixture models trained with expectation-maximization, association rule mining with Apriori using support and confidence, anomaly detection and variational inference for latent variable models.

Below are 30 practice questions from a pool of 210 Unsupervised Learning MCQs, one of 17 topics in AI / Machine Learning & Data Analytics. Each shows the correct answer with an explanation; when you are ready, take a timed quiz to test recall under exam conditions.

Practice Questions

Each question below shows the correct answer with a full explanation. Use these to build conceptual understanding before attempting a timed quiz.

Unsupervised LearningEasy

Q1. What does clustering do?

  1. A.Classifies data with known labels
  2. B.Generates entirely new data points
  3. C.Predicts continuous numerical values
  4. D.Groups similar data points together✓ Correct

Explanation

Clustering groups data points with similar characteristics together without predefined labels.

Report an error in this question

Unsupervised LearningEasy

Q2. PCA stands for:

  1. A.Principal Component Analysis✓ Correct
  2. B.Primary Cluster Algorithm
  3. C.Probability Calculation Approach
  4. D.Pattern Classification Analysis

Explanation

PCA (Principal Component Analysis) is a dimensionality reduction technique that finds the directions of maximum variance.

Report an error in this question

Unsupervised LearningEasy

Q3. In K-Means, what is a centroid?

  1. A.A class target label
  2. B.An extreme outlier point
  3. C.The center point of a cluster✓ Correct
  4. D.A single input feature

Explanation

A centroid is the mean position of all points in a cluster, serving as the cluster's center.

Report an error in this question

Unsupervised LearningMedium

Q4. In PCA, principal components are:

  1. A.Orthogonal directions of maximum variance
  2. B.Cluster center points✓ Correct
  3. C.Randomly chosen directions
  4. D.Features that contain missing values

Explanation

Principal components are orthogonal (perpendicular) directions that capture the maximum variance in the data, ordered by variance explained.

Report an error in this question

Unsupervised LearningEasy

Q5. What is an unsupervised learning application?

  1. A.Labeled prediction
  2. B.Anomaly detection✓ Correct
  3. C.Supervised classification
  4. D.Regression

Explanation

Anomaly detection identifies unusual patterns without labeled examples of anomalies, making it an unsupervised task.

Report an error in this question

Unsupervised LearningEasy

Q6. Which is an example of unsupervised learning?

  1. A.Predicting house prices
  2. B.Labeled image classification
  3. C.Customer segmentation✓ Correct
  4. D.Email spam detection

Explanation

Customer segmentation groups customers based on behavior without predefined labels, making it an unsupervised task.

Report an error in this question

Unsupervised LearningMedium

Q7. How does the elbow method work for K-Means?

  1. A.Select the value of K completely at random from a uniform range
  2. B.Use the largest possible K equal to the number of data points
  3. C.Plot inertia vs K and find where adding clusters gives diminishing returns✓ Correct
  4. D.Always use K equal to three regardless of the dataset size

Explanation

The elbow method plots the within-cluster sum of squares against K, and the optimal K is at the 'elbow' where improvement begins to diminish significantly.

Report an error in this question

Unsupervised LearningMedium

Q8. DBSCAN stands for:

  1. A.Data-Based Statistical Clustering Algorithm
  2. B.Database Scan Algorithm Notation
  3. C.Density-Based Spatial Clustering of Applications with Noise
  4. D.Deep Binary Sorting Classification Algorithm✓ Correct

Explanation

DBSCAN groups together densely packed points and marks points in low-density regions as outliers/noise.

Report an error in this question

Unsupervised LearningEasy

Q9. What does the 'K' in K-Means represent?

  1. A.The number of iterations
  2. B.The number of features
  3. C.The learning rate
  4. D.The number of clusters✓ Correct

Explanation

K represents the predetermined number of clusters that the algorithm will partition the data into.

Report an error in this question

Unsupervised LearningMedium

Q10. What advantage does DBSCAN have over K-Means?

  1. A.It only works with fully labeled and supervised training data
  2. B.It can find arbitrarily shaped clusters and doesn't require specifying K✓ Correct
  3. C.It always gives strictly better results than all alternatives
  4. D.It is always computationally faster than every other algorithm

Explanation

DBSCAN discovers clusters of arbitrary shape, handles noise/outliers naturally, and doesn't require prespecifying the number of clusters.

Report an error in this question

Unsupervised LearningEasy

Q11. Unsupervised learning works with:

  1. A.Unlabeled data✓ Correct
  2. B.Numerical data only
  3. C.Textual data only
  4. D.Labeled output data

Explanation

Unsupervised learning discovers patterns in data without labeled outputs or predefined categories.

Report an error in this question

Unsupervised LearningEasy

Q12. What is dimensionality reduction?

  1. A.Adding more derived features to expand the feature space
  2. B.Removing all data from the storage system entirely
  3. C.Reducing the number of features while preserving important information✓ Correct
  4. D.Increasing the overall size of the raw training dataset

Explanation

Dimensionality reduction reduces the number of input variables while retaining the essential information in the data.

Report an error in this question

Unsupervised LearningEasy

Q13. K-Means is a type of:

  1. A.Regression algorithm
  2. B.Clustering algorithm✓ Correct
  3. C.Classification algorithm
  4. D.Sorting algorithm

Explanation

K-Means is a clustering algorithm that partitions data into K groups based on similarity.

Report an error in this question

Unsupervised LearningEasy

Q14. Association rule mining is used to:

  1. A.Find relationships between items in transactions✓ Correct
  2. B.Classify documents into predefined categories
  3. C.Predict future numerical stock price values
  4. D.Train multi-layer deep neural network models

Explanation

Association rule mining discovers interesting relationships between items in large datasets, such as market basket analysis.

Report an error in this question

Unsupervised LearningMedium

Q15. What is the silhouette score?

  1. A.The Euclidean distance between every pair of cluster centroids
  2. B.A measure of how similar a point is to its own cluster vs neighboring clusters✓ Correct
  3. C.The total number of clusters found by the clustering algorithm
  4. D.The total number of iterations before the algorithm converges

Explanation

Silhouette score ranges from -1 to 1, measuring how similar a point is to its own cluster compared to other clusters. Higher values indicate better clustering.

Report an error in this question

Unsupervised LearningHard

Q16. In DBSCAN, what are core points, border points, and noise points?

  1. A.Noise points are the ones that form the densest most cohesive clusters
  2. B.Core points have enough neighbors; border points are near core points; noise points are isolated✓ Correct
  3. C.Core points are the extreme outliers that are furthest from all clusters
  4. D.All three point types are treated exactly the same by the clustering algorithm

Explanation

Core points have at least MinPts neighbors within epsilon; border points are within epsilon of a core point but have fewer neighbors; noise points are neither.

Report an error in this question

Unsupervised LearningHard

Q17. What is the Information Criterion (BIC/AIC) used for in clustering?

  1. A.Performing feature selection for supervised model training
  2. B.Measuring the internal purity of each individual data cluster
  3. C.Calculating the pairwise distances between all data points
  4. D.Selecting the optimal number of clusters by balancing fit and complexity✓ Correct

Explanation

BIC and AIC balance model fit against complexity to select the optimal number of clusters, penalizing models with more parameters.

Report an error in this question

Unsupervised LearningHard

Q18. What is the UMAP algorithm?

  1. A.A dimensionality reduction technique based on manifold learning and topology✓ Correct
  2. B.A gradient-based numerical regression training technique
  3. C.A supervised multi-class classification prediction method
  4. D.An unsupervised density-based spatial clustering algorithm

Explanation

UMAP (Uniform Manifold Approximation and Projection) is a manifold learning technique for dimensionality reduction that often preserves more global structure than t-SNE.

Report an error in this question

Unsupervised LearningHard

Q19. What is contrastive learning in the context of unsupervised representation learning?

  1. A.A linear regression method for predicting numerical outputs
  2. B.A purely supervised classification technique using labels
  3. C.A density-based spatial clustering approach for grouping data
  4. D.Learning representations by contrasting similar and dissimilar pairs✓ Correct

Explanation

Contrastive learning learns representations by pulling similar (positive) pairs together and pushing dissimilar (negative) pairs apart in the embedding space.

Report an error in this question

Unsupervised LearningMedium

Q20. What is hierarchical clustering?

  1. A.A supervised regression technique for numerical prediction
  2. B.A deep neural network architecture for image recognition
  3. C.Building a tree of clusters by merging or splitting iteratively✓ Correct
  4. D.Using K-Means clustering multiple times in succession

Explanation

Hierarchical clustering creates a tree (dendrogram) of clusters by either merging (agglomerative) or splitting (divisive) iteratively.

Report an error in this question

Unsupervised LearningMedium

Q21. What is the Apriori algorithm used for?

  1. A.Frequent itemset mining and association rules✓ Correct
  2. B.Dimensionality reduction projections
  3. C.Numerical regression for predictions
  4. D.Supervised classification with labeled data

Explanation

The Apriori algorithm finds frequent itemsets in transaction data and generates association rules (e.g., if A then B).

Report an error in this question

Unsupervised LearningHard

Q22. What is Self-Organizing Map (SOM)?

  1. A.An unsupervised neural network that produces a low-dimensional representation of data✓ Correct
  2. B.A supervised classification model that predicts categorical labels
  3. C.A gradient boosting algorithm that builds sequential tree models
  4. D.A linear regression model that predicts continuous target values

Explanation

SOM is an unsupervised neural network that maps high-dimensional data onto a low-dimensional grid while preserving topological properties.

Report an error in this question

Unsupervised LearningHard

Q23. What is the manifold hypothesis in unsupervised learning?

  1. A.All real-world data is always uniformly distributed in space
  2. B.Data always naturally forms spherical well-separated clusters
  3. C.All feature dimensions are always statistically independent
  4. D.High-dimensional data often lies on a lower-dimensional manifold✓ Correct

Explanation

The manifold hypothesis states that real-world high-dimensional data tends to concentrate near a lower-dimensional manifold, justifying dimensionality reduction techniques.

Report an error in this question

Unsupervised LearningMedium

Q24. What is t-SNE used for?

  1. A.Visualizing high-dimensional data in 2D or 3D✓ Correct
  2. B.Imputing missing data point values
  3. C.Training deep neural network models
  4. D.Selecting the most relevant features

Explanation

t-SNE (t-distributed Stochastic Neighbor Embedding) reduces high-dimensional data to 2D or 3D for visualization while preserving local structure.

Report an error in this question

Unsupervised LearningHard

Q25. What is the difference between hard and soft clustering?

  1. A.Hard assigns each point to one cluster; soft assigns probabilities of belonging to each cluster✓ Correct
  2. B.Soft clustering does not exist as a valid machine learning technique
  3. C.Hard clustering is always the strictly better approach for all problems
  4. D.They are completely identical approaches with no meaningful differences

Explanation

Hard clustering (e.g., K-Means) assigns each point to exactly one cluster, while soft clustering (e.g., GMM) assigns probabilities of belonging to each cluster.

Report an error in this question

Unsupervised LearningHard

Q26. What is the Gaussian Mixture Model (GMM)?

  1. A.A tree-based decision boundary model for supervised classification prediction
  2. B.A probabilistic model assuming data comes from a mixture of Gaussian distributions✓ Correct
  3. C.A linear regression method for predicting continuous numerical target values
  4. D.A filter-based feature selection technique using statistical hypothesis tests

Explanation

GMM models data as arising from a mixture of multiple Gaussian distributions, using the EM algorithm to estimate parameters, providing soft cluster assignments.

Report an error in this question

Unsupervised LearningHard

Q27. How does the OPTICS algorithm improve upon DBSCAN?

  1. A.It requires explicitly specifying the number K of clusters
  2. B.It only works with fully labeled and supervised training data
  3. C.It is always computationally faster than every other algorithm
  4. D.It handles clusters of varying density without requiring a fixed epsilon✓ Correct

Explanation

OPTICS (Ordering Points To Identify Clustering Structure) creates a reachability plot that allows identifying clusters at different density levels, unlike DBSCAN's fixed epsilon.

Report an error in this question

Unsupervised LearningMedium

Q28. In association rules, what does 'support' measure?

  1. A.The conditional confidence level of the rule
  2. B.The overall predictive accuracy of the rule
  3. C.The proportion of transactions containing an itemset✓ Correct
  4. D.The total number of features in the dataset

Explanation

Support measures how frequently an itemset appears in the dataset, calculated as the proportion of transactions containing that itemset.

Report an error in this question

Unsupervised LearningHard

Q29. How does spectral clustering work?

  1. A.Uses deep neural networks to learn cluster assignment mappings
  2. B.Uses decision tree ensembles to group data into discrete clusters
  3. C.Uses eigenvalues of a similarity matrix to reduce dimensionality before clustering✓ Correct
  4. D.Applies K-Means clustering directly without any preprocessing transformation

Explanation

Spectral clustering constructs a similarity graph, computes its Laplacian's eigenvectors, and clusters in the reduced eigenspace, handling non-convex shapes.

Report an error in this question

Unsupervised LearningMedium

Q30. What is the difference between agglomerative and divisive hierarchical clustering?

  1. A.Neither method uses a hierarchy
  2. B.Agglomerative is bottom-up; divisive is top-down
  3. C.Agglomerative is top-down; divisive is bottom-up✓ Correct
  4. D.The two methods are completely identical

Explanation

Agglomerative clustering starts with individual points and merges clusters bottom-up, while divisive starts with one cluster and splits top-down.

Report an error in this question

Ready to test yourself on Unsupervised Learning?

Take a timed quiz drawn from 210+ questions on this topic. No signup required — your progress saves in your browser.

Start Unsupervised Learning Quiz