HomeSubjectsUniversityBlogAbout

Supervised Learning

Topic in AI / Machine Learning & Data Analytics

210 total MCQsShowing 30 with explanations10 Easy10 Medium10 Hard

About This Topic

Supervised learning trains a model on labelled examples, input features paired with known outputs, so it can predict outputs for new unseen inputs. MCQs separate regression from classification and test linear and logistic regression, k-nearest neighbours (KNN), decision trees, Naive Bayes and support vector machines, including the role of support vectors, margins and kernels. Regularization is a frequent theme: L1 (Lasso), L2 (Ridge) and elastic net, which mixes the two. Expect questions on overfitting, underfitting and the bias-variance trade-off, plus theory items on VC dimension, loss functions such as cross-entropy and hinge loss, and label smoothing.

Below are 30 practice questions from a pool of 210 Supervised Learning MCQs, one of 17 topics in AI / Machine Learning & Data Analytics. Each shows the correct answer with an explanation; when you are ready, take a timed quiz to test recall under exam conditions.

Practice Questions

Each question below shows the correct answer with a full explanation. Use these to build conceptual understanding before attempting a timed quiz.

Supervised LearningEasy

Q1. In supervised learning, the model learns from:

  1. A.Labeled data with known outputs✓ Correct
  2. B.Randomly generated fake data
  3. C.Unlabeled unstructured data only
  4. D.No data of any kind at all

Explanation

Supervised learning uses labeled training data where both inputs and correct outputs are provided.

Report an error in this question

Supervised LearningEasy

Q2. Linear regression is used for:

  1. A.Classification tasks only
  2. B.Reducing dataset dimensionality
  3. C.Predicting continuous numerical values
  4. D.Clustering data points✓ Correct

Explanation

Linear regression predicts a continuous target variable based on a linear relationship with input features.

Report an error in this question

Supervised LearningEasy

Q3. What is classification in supervised learning?

  1. A.Grouping similar data together
  2. B.Predicting a discrete category or class
  3. C.Predicting a continuous output value✓ Correct
  4. D.Reducing feature dimensions

Explanation

Classification assigns input data to predefined discrete categories or classes.

Report an error in this question

Supervised LearningEasy

Q4. Which algorithm is commonly used for classification?

  1. A.K-Means method
  2. B.Decision Tree
  3. C.DBSCAN method
  4. D.PCA method✓ Correct

Explanation

Decision Trees are widely used classification algorithms that split data based on feature values to make predictions.

Report an error in this question

Supervised LearningEasy

Q5. What is the target variable in supervised learning?

  1. A.The output the model predicts✓ Correct
  2. B.An internal training weight
  3. C.A tunable hyperparameter
  4. D.An individual input feature

Explanation

The target variable (label) is the output that the model is trained to predict from the input features.

Report an error in this question

Supervised LearningEasy

Q6. Logistic regression is used for:

  1. A.Dimensionality reduction
  2. B.Linear regression
  3. C.Binary classification✓ Correct
  4. D.Clustering

Explanation

Despite its name, logistic regression is used for binary classification by predicting probabilities using the sigmoid function.

Report an error in this question

Supervised LearningEasy

Q7. What is overfitting?

  1. A.Model performs well on training data but poorly on new data✓ Correct
  2. B.Model is too simple and has very few learned parameters
  3. C.Model performs poorly on both training and testing data
  4. D.Model has absolutely no trainable weight parameters

Explanation

Overfitting occurs when a model learns the noise in training data and fails to generalize to unseen data.

Report an error in this question

Supervised LearningEasy

Q8. K-Nearest Neighbors (KNN) classifies based on:

  1. A.A low-rank matrix decomposition factorization
  2. B.The steepest gradient descent optimization step
  3. C.An entirely random selection of a class label
  4. D.The majority class among the K nearest data points✓ Correct

Explanation

KNN classifies a new point by finding the K nearest training examples and assigning the majority class among them.

Report an error in this question

Supervised LearningEasy

Q9. What is a feature in machine learning?

  1. A.The model architecture itself
  2. B.The optimization training algorithm
  3. C.An input variable used for prediction✓ Correct
  4. D.The output target variable label

Explanation

A feature is an individual measurable property or characteristic of the data used as input for model training.

Report an error in this question

Supervised LearningEasy

Q10. What is underfitting?

  1. A.Model has an excessive number of parameters
  2. B.Model has been trained for far too many epochs
  3. C.Model is too simple to capture patterns in data✓ Correct
  4. D.Model has memorized all the training data points

Explanation

Underfitting occurs when a model is too simple to capture the underlying patterns in the data.

Report an error in this question

Supervised LearningMedium

Q11. What is the cost function in linear regression?

  1. A.Hinge loss function
  2. B.Log loss function✓ Correct
  3. C.Cross-entropy loss
  4. D.Mean Squared Error (MSE)

Explanation

Linear regression typically minimizes the Mean Squared Error, which is the average of squared differences between predicted and actual values.

Report an error in this question

Supervised LearningMedium

Q12. The sigmoid function maps values to:

  1. A.Range between 0 and 1✓ Correct
  2. B.Only integers
  3. C.Range between -1 and 1
  4. D.Any real number

Explanation

The sigmoid function maps any real number to a value between 0 and 1, useful for probability estimation.

Report an error in this question

Supervised LearningMedium

Q13. What is regularization in supervised learning?

  1. A.Increasing overall model complexity further
  2. B.Adding more training data to the model
  3. C.Adding a penalty term to prevent overfitting✓ Correct
  4. D.Removing unnecessary features from the model

Explanation

Regularization adds a penalty term to the loss function to discourage overly complex models and prevent overfitting.

Report an error in this question

Supervised LearningMedium

Q14. What is the difference between L1 and L2 regularization?

  1. A.L1 uses absolute values and can produce sparse models; L2 uses squared values
  2. B.L1 and L2 are completely identical techniques
  3. C.L2 regularization always produces sparse model weights
  4. D.L1 regularization is always the better choice overall✓ Correct

Explanation

L1 (Lasso) adds absolute values of coefficients and can zero out features; L2 (Ridge) adds squared coefficients and shrinks them towards zero.

Report an error in this question

Supervised LearningMedium

Q15. What is a Support Vector Machine (SVM)?

  1. A.An algorithm that finds the optimal hyperplane to separate classes✓ Correct
  2. B.A linear dimensionality reduction projection method
  3. C.A type of deep feedforward artificial neural network
  4. D.An unsupervised density-based clustering algorithm

Explanation

SVM finds the hyperplane that maximizes the margin between different classes in the feature space.

Report an error in this question

Supervised LearningMedium

Q16. What is the bias-variance tradeoff?

  1. A.Balancing model simplicity (bias) against sensitivity to training data (variance)✓ Correct
  2. B.Choosing between running computation on a CPU versus a GPU device
  3. C.Selecting the right proportions for training and test data splits
  4. D.Choosing an appropriate learning rate for gradient optimization

Explanation

The bias-variance tradeoff balances underfitting (high bias) and overfitting (high variance) to achieve optimal model performance.

Report an error in this question

Supervised LearningMedium

Q17. In decision trees, what is information gain?

  1. A.The reduction in entropy after splitting on a feature
  2. B.The overall accuracy of the trained model✓ Correct
  3. C.The total number of leaf nodes in the tree
  4. D.The current depth of the decision tree

Explanation

Information gain measures the reduction in entropy (uncertainty) achieved by splitting the data on a particular feature.

Report an error in this question

Supervised LearningMedium

Q18. What is cross-validation?

  1. A.Evaluating using only the training data subset
  2. B.Splitting data into multiple folds to train and validate the model
  3. C.Training the model on all available data at once
  4. D.Evaluating using only the test data subset✓ Correct

Explanation

Cross-validation divides data into K folds, training on K-1 and validating on the remaining fold, repeating K times for robust evaluation.

Report an error in this question

Supervised LearningMedium

Q19. What is the purpose of the softmax function?

  1. A.Converting raw scores into probabilities for multi-class classification✓ Correct
  2. B.Performing binary classification tasks on labeled data only
  3. C.Performing unsupervised clustering of unlabeled data points
  4. D.Performing regression to predict continuous numeric values

Explanation

Softmax converts a vector of raw scores into a probability distribution over multiple classes, where all probabilities sum to 1.

Report an error in this question

Supervised LearningMedium

Q20. What is Naive Bayes based on?

  1. A.A deep learning based neural network model
  2. B.Bayes' theorem with an assumption of feature independence✓ Correct
  3. C.An iterative gradient descent optimization method
  4. D.An ensemble of many decision tree classifiers

Explanation

Naive Bayes applies Bayes' theorem assuming features are conditionally independent given the class, making computation efficient.

Report an error in this question

Supervised LearningHard

Q21. What is the kernel trick in SVM?

  1. A.Mapping data to a higher-dimensional space without explicit computation✓ Correct
  2. B.A dimensionality reduction visualization projection technique
  3. C.A wrapper-based method for iterative feature selection
  4. D.A preprocessing data cleaning and normalization technique

Explanation

The kernel trick allows SVM to compute the dot product in a higher-dimensional space without explicitly transforming the data, enabling non-linear classification.

Report an error in this question

Supervised LearningHard

Q22. What is the VC dimension?

  1. A.The total count of input features in the dataset
  2. B.The step size or magnitude of the learning rate
  3. C.A measure of the capacity and complexity of a classification model✓ Correct
  4. D.The total number of training samples available

Explanation

The Vapnik-Chervonenkis dimension measures the capacity of a model class - the largest set of points it can shatter (classify in all possible ways).

Report an error in this question

Supervised LearningHard

Q23. In gradient boosting, what does each subsequent tree learn?

  1. A.The residuals (errors) of the previous ensemble
  2. B.Random patterns found within the noise
  3. C.The original target values from the training data
  4. D.Nothing new beyond the original input features✓ Correct

Explanation

In gradient boosting, each new tree is trained on the residual errors of the current ensemble, progressively reducing the overall error.

Report an error in this question

Supervised LearningHard

Q24. What is the difference between hard and soft margin SVM?

  1. A.Hard margin and soft margin are completely identical approaches
  2. B.Hard margin is always the preferred choice over soft margin
  3. C.Soft margin SVM never works well on real-world datasets✓ Correct
  4. D.Hard margin allows no misclassification; soft margin allows some with a penalty

Explanation

Hard margin SVM requires perfect separation (no misclassifications), while soft margin allows some violations with a penalty parameter C controlling the tradeoff.

Report an error in this question

Supervised LearningHard

Q25. What is the Representer Theorem?

  1. A.A foundational theorem about efficient compressed data representation and storage formats
  2. B.The optimal solution in kernel methods can be expressed as a linear combination of kernel evaluations at training points✓ Correct
  3. C.A graphical visualization principle for plotting high-dimensional feature space relationships
  4. D.A data warehousing method for efficiently storing columnar records in structured databases

Explanation

The Representer Theorem states that the solution to regularized empirical risk minimization in a reproducing kernel Hilbert space is a linear combination of kernel evaluations at training points.

Report an error in this question

Supervised LearningHard

Q26. What is Platt scaling?

  1. A.A method to calibrate SVM outputs into probabilities using a sigmoid function✓ Correct
  2. B.A standardization feature scaling method using z-score transforms of values
  3. C.A model combination ensemble technique using weighted averaging of outputs
  4. D.A min-max data normalization technique for scaling feature ranges to a bound

Explanation

Platt scaling fits a sigmoid function to the SVM decision values to produce calibrated probability estimates.

Report an error in this question

Supervised LearningHard

Q27. What is the effect of increasing the C parameter in SVM?

  1. A.The model becomes much simpler and heavily regularized overall
  2. B.The decision boundary margin becomes significantly wider
  3. C.The model becomes more sensitive to individual points, risking overfitting✓ Correct
  4. D.The amount of L2 regularization penalty increases significantly

Explanation

A larger C penalizes misclassifications more heavily, leading to a narrower margin and potentially overfitting to training data.

Report an error in this question

Supervised LearningHard

Q28. What is the difference between Gini impurity and entropy in decision trees?

  1. A.Gini impurity is always the better choice for all tree problems
  2. B.Both measure node impurity but Gini is computationally simpler; entropy uses logarithms✓ Correct
  3. C.Entropy is always the better choice for all tree problems
  4. D.They produce completely different tree structures every single time

Explanation

Both measure impurity at a node. Gini impurity uses squared probabilities while entropy uses logarithms. They usually yield similar trees, but Gini is slightly faster to compute.

Report an error in this question

Supervised LearningHard

Q29. What is the Structural Risk Minimization principle?

  1. A.Always using the simplest available model regardless of performance
  2. B.Maximizing model complexity to fit all training data points exactly✓ Correct
  3. C.Balancing empirical risk and model complexity to minimize generalization error
  4. D.Minimizing only the training error and ignoring test performance

Explanation

Structural Risk Minimization minimizes an upper bound on generalization error by balancing training error against model complexity, providing theoretical justification for regularization.

Report an error in this question

Supervised LearningHard

Q30. In logistic regression, what does multicollinearity cause?

  1. A.More accurate and well-calibrated probability estimates
  2. B.Significantly faster convergence during the training process
  3. C.Unstable and unreliable coefficient estimates with large standard errors✓ Correct
  4. D.Consistently better and more accurate model predictions overall

Explanation

Multicollinearity causes coefficient estimates to become unstable with inflated standard errors, making it difficult to determine individual feature contributions.

Report an error in this question

Ready to test yourself on Supervised Learning?

Take a timed quiz drawn from 210+ questions on this topic. No signup required — your progress saves in your browser.

Start Supervised Learning Quiz