HomeSubjectsUniversityBlogAbout

Model Evaluation & Validation

Topic in AI / Machine Learning & Data Analytics

210 total MCQsShowing 30 with explanations10 Easy10 Medium10 Hard

About This Topic

Model evaluation measures how well a trained model performs on unseen data, using suitable metrics and validation schemes to estimate real-world accuracy. Questions define accuracy, precision, recall, F1 score and specificity from a confusion matrix, and ask when recall matters more, namely when false negatives are costly. Expect items on ROC curves and AUC, the precision-recall trade-off, and regression metrics such as MAE, MSE, RMSE and R-squared. Validation questions cover train, validation and test splits, k-fold and stratified cross-validation, and data leakage. Advanced items mention Cohen's kappa, calibration and statistical power in model comparison.

Below are 30 practice questions from a pool of 210 Model Evaluation & Validation MCQs, one of 17 topics in AI / Machine Learning & Data Analytics. Each shows the correct answer with an explanation; when you are ready, take a timed quiz to test recall under exam conditions.

Practice Questions

Each question below shows the correct answer with a full explanation. Use these to build conceptual understanding before attempting a timed quiz.

Model Evaluation & ValidationEasy

Q1. Accuracy is defined as:

  1. A.Number of correct predictions divided by total predictions✓ Correct
  2. B.Only the count of true negative classification results
  3. C.The total number of features used in the model
  4. D.Only the count of true positive classification results

Explanation

Accuracy = (correct predictions) / (total predictions), measuring the overall proportion of correct classifications.

Report an error in this question

Model Evaluation & ValidationEasy

Q2. What is a confusion matrix?

  1. A.A Python data visualization chart library
  2. B.A type of multi-layer deep neural network
  3. C.A preprocessing data cleaning utility tool
  4. D.A table showing predicted vs actual classifications✓ Correct

Explanation

A confusion matrix is a table that summarizes the performance of a classification model by showing true positives, true negatives, false positives, and false negatives.

Report an error in this question

Model Evaluation & ValidationEasy

Q3. What is a training set used for?

  1. A.Deploying the model to production
  2. B.Training the model to learn patterns✓ Correct
  3. C.Evaluating the final performance
  4. D.Tuning the model hyperparameters

Explanation

The training set is used to train the model by allowing it to learn patterns and relationships in the data.

Report an error in this question

Model Evaluation & ValidationEasy

Q4. What is a test set used for?

  1. A.Training the model on labeled training samples
  2. B.Tuning the model's hyperparameter settings
  3. C.Cleaning and preprocessing the raw data
  4. D.Evaluating the model's performance on unseen data✓ Correct

Explanation

The test set evaluates how well the trained model generalizes to new, unseen data.

Report an error in this question

Model Evaluation & ValidationEasy

Q5. Mean Squared Error (MSE) is used to evaluate:

  1. A.Regression models✓ Correct
  2. B.Classification models only
  3. C.Clustering algorithms
  4. D.Association rule mining

Explanation

MSE measures the average squared difference between predicted and actual values, commonly used to evaluate regression models.

Report an error in this question

Model Evaluation & ValidationEasy

Q6. What is a validation set?

  1. A.The same set as the independent held-out test set
  2. B.The entire full dataset including all partitions
  3. C.A preprocessing data cleaning utility pipeline
  4. D.A subset used to tune hyperparameters during training✓ Correct

Explanation

A validation set is separate from training and test sets, used to tune hyperparameters and make model selection decisions during development.

Report an error in this question

Model Evaluation & ValidationEasy

Q7. What is precision in classification?

  1. A.The recall evaluation metric✓ Correct
  2. B.Overall model accuracy metric
  3. C.True Positives divided by the total count
  4. D.True Positives / (True Positives + False Positives)

Explanation

Precision is the proportion of positive predictions that are actually correct: TP / (TP + FP).

Report an error in this question

Model Evaluation & ValidationEasy

Q8. What is a false positive?

  1. A.Making a correct and accurate model prediction
  2. B.Predicting positive when the actual value is negative✓ Correct
  3. C.Encountering a missing value in the raw data
  4. D.Predicting negative when the actual value is positive

Explanation

A false positive (Type I error) occurs when the model incorrectly predicts the positive class.

Report an error in this question

Model Evaluation & ValidationEasy

Q9. R-squared measures:

  1. A.The proportion of variance in the target explained by the model✓ Correct
  2. B.The total number of data samples in the dataset
  3. C.The total number of input features in the training dataset
  4. D.The total elapsed time required for model training

Explanation

R-squared (coefficient of determination) measures how much of the variance in the dependent variable is explained by the model, ranging from 0 to 1.

Report an error in this question

Model Evaluation & ValidationEasy

Q10. What is recall (sensitivity)?

  1. A.True Positives / (True Positives + False Positives)
  2. B.True Positives / (True Positives + False Negatives)
  3. C.The specificity metric
  4. D.The overall accuracy metric✓ Correct

Explanation

Recall is the proportion of actual positives correctly identified: TP / (TP + FN).

Report an error in this question

Model Evaluation & ValidationMedium

Q11. What does AUC-ROC measure?

  1. A.Only the precision metric of the trained classification model✓ Correct
  2. B.Only the accuracy metric of the trained classification model
  3. C.Only the recall metric of the trained classification model
  4. D.The overall ability of the model to discriminate between classes

Explanation

AUC-ROC (Area Under the ROC Curve) measures the model's ability to distinguish between classes across all thresholds. A value of 1.0 is perfect.

Report an error in this question

Model Evaluation & ValidationMedium

Q12. What does the ROC curve plot?

  1. A.Training error values plotted against validation error values
  2. B.True Positive Rate vs False Positive Rate at various thresholds✓ Correct
  3. C.Precision values plotted against corresponding recall values
  4. D.Training accuracy values plotted against the loss values

Explanation

The ROC (Receiver Operating Characteristic) curve plots TPR against FPR at various classification thresholds.

Report an error in this question

Model Evaluation & ValidationMedium

Q13. What is the F1-score?

  1. A.The arithmetic mean of precision and recall✓ Correct
  2. B.The product of precision and recall values
  3. C.The harmonic mean of precision and recall
  4. D.The square of the model accuracy

Explanation

F1-score = 2 * (precision * recall) / (precision + recall), providing a balanced measure when precision and recall are both important.

Report an error in this question

Model Evaluation & ValidationMedium

Q14. What is the difference between micro and macro averaging?

  1. A.Micro averaging is always a better choice than macro averaging✓ Correct
  2. B.Micro aggregates all instances globally; macro averages per-class metrics equally
  3. C.Micro and macro averaging are completely identical approaches
  4. D.Macro averaging ignores the overall class distribution entirely

Explanation

Micro-averaging aggregates all instances globally, while macro-averaging computes metrics per class and averages them, giving equal weight to each class.

Report an error in this question

Model Evaluation & ValidationMedium

Q15. When is recall more important than precision?

  1. A.When simple accuracy alone is a sufficient metric
  2. B.When the dataset classes are perfectly balanced overall
  3. C.When false positives are very costly to the outcome✓ Correct
  4. D.When missing positive cases is costly, like disease detection

Explanation

Recall is prioritized when false negatives are costly, such as disease screening where missing a positive case could be life-threatening.

Report an error in this question

Model Evaluation & ValidationMedium

Q16. What is stratified cross-validation?

  1. A.Performing completely random splitting of the data
  2. B.Ignoring overall class balance during data splitting✓ Correct
  3. C.Using only the majority class for model training purposes
  4. D.K-fold CV that preserves the class distribution in each fold

Explanation

Stratified CV ensures each fold has approximately the same proportion of each class as the original dataset, important for imbalanced data.

Report an error in this question

Model Evaluation & ValidationHard

Q17. What is the difference between Type I and Type II errors?

  1. A.The two types are completely identical errors
  2. B.Type I is a false positive; Type II is a false negative✓ Correct
  3. C.Type I is a false negative; Type II is a false positive
  4. D.Both are actually true positive outcomes

Explanation

Type I error (false positive) incorrectly rejects a true null hypothesis. Type II error (false negative) fails to reject a false null hypothesis.

Report an error in this question

Model Evaluation & ValidationMedium

Q18. What is the log loss (cross-entropy loss)?

  1. A.A feature selection method based on mutual information scores
  2. B.The exact same metric as mean squared error for regression
  3. C.A metric specifically designed for evaluating cluster quality
  4. D.A loss function that penalizes confident wrong predictions heavily✓ Correct

Explanation

Log loss measures the performance of a classification model by penalizing predictions based on how far they are from the actual labels, with confident wrong predictions penalized heavily.

Report an error in this question

Model Evaluation & ValidationHard

Q19. What is the Matthews Correlation Coefficient (MCC)?

  1. A.The exact same metric as standard classification accuracy for balanced data
  2. B.A metric specifically designed for evaluating unsupervised clustering
  3. C.A metric that is only applicable to multi-class classification problems
  4. D.A balanced metric using all four confusion matrix values, ranging from -1 to 1✓ Correct

Explanation

MCC considers TP, TN, FP, and FN, producing a balanced measure even on imbalanced datasets. +1 is perfect, 0 is random, -1 is complete disagreement.

Report an error in this question

Model Evaluation & ValidationMedium

Q20. What is the purpose of a learning curve?

  1. A.To select the most important features from the training data
  2. B.To diagnose overfitting or underfitting by plotting performance vs training size✓ Correct
  3. C.To clean and preprocess the raw input data before training
  4. D.To learn about new machine learning algorithms from documentation

Explanation

A learning curve plots training and validation performance as a function of training set size, helping diagnose whether the model suffers from high bias or high variance.

Report an error in this question

Model Evaluation & ValidationMedium

Q21. What is K-fold cross-validation?

  1. A.Dividing data into K folds and using each fold as validation once
  2. B.Training the model completely without any validation step✓ Correct
  3. C.Using all available data solely for the training process
  4. D.Using only two simple train-test splits for evaluation

Explanation

K-fold CV splits data into K equal folds, trains K times using K-1 folds for training and the remaining fold for validation, averaging the results.

Report an error in this question

Model Evaluation & ValidationMedium

Q22. What is Mean Absolute Error (MAE)?

  1. A.The single maximum prediction error value✓ Correct
  2. B.The median of all prediction errors
  3. C.The average of absolute differences between predicted and actual values
  4. D.The average of all squared differences between predictions

Explanation

MAE = mean(|predicted - actual|). It measures the average magnitude of errors without considering direction, less sensitive to outliers than MSE.

Report an error in this question

Model Evaluation & ValidationHard

Q23. What is the difference between leave-one-out CV and K-fold CV?

  1. A.They are exactly identical approaches with no meaningful differences whatsoever
  2. B.LOOCV is always significantly faster computationally than K-fold approaches
  3. C.K-fold always produces strictly more accurate results than LOOCV methods
  4. D.LOOCV uses N-1 samples for training and 1 for testing each time; K-fold uses larger folds✓ Correct

Explanation

LOOCV is a special case where K=N (number of samples), using each sample once as validation. It has low bias but high variance and is computationally expensive.

Report an error in this question

Model Evaluation & ValidationHard

Q24. What is the purpose of bootstrapping in model evaluation?

  1. A.Performing automated feature selection and dimensionality reduction
  2. B.Estimating confidence intervals for metrics by resampling with replacement✓ Correct
  3. C.Running hyperparameter tuning using grid or random search
  4. D.Increasing the total amount of available training data samples

Explanation

Bootstrapping creates multiple resampled datasets to estimate the distribution of a statistic, providing confidence intervals for model performance metrics.

Report an error in this question

Model Evaluation & ValidationHard

Q25. When should you use the PR-AUC instead of ROC-AUC?

  1. A.For evaluating unsupervised clustering partition quality
  2. B.When dealing with highly imbalanced datasets where positive class is rare✓ Correct
  3. C.When the class distribution is perfectly balanced and equal
  4. D.For evaluating regression models with continuous outputs

Explanation

PR-AUC is preferred for highly imbalanced datasets because ROC-AUC can be misleadingly high when the negative class dominates, while PR-AUC focuses on the minority positive class.

Report an error in this question

Model Evaluation & ValidationHard

Q26. What is nested cross-validation?

  1. A.An outer CV loop for evaluation with an inner CV loop for hyperparameter tuning✓ Correct
  2. B.A single random train-test split without any cross-validation folds
  3. C.Cross-validation performed entirely without any held-out validation data
  4. D.The same standard approach as regular single-loop K-fold cross-validation

Explanation

Nested CV uses an outer loop for unbiased performance estimation and an inner loop for hyperparameter tuning, preventing optimistic bias from using the same data for both.

Report an error in this question

Model Evaluation & ValidationHard

Q27. What is the Brier score?

  1. A.A feature importance measure for trees
  2. B.The mean squared error of probabilistic predictions✓ Correct
  3. C.A ranking metric for ordering search results
  4. D.A clustering score for partition quality

Explanation

The Brier score measures the accuracy of probabilistic predictions by computing the mean squared difference between predicted probabilities and actual outcomes (0 or 1).

Report an error in this question

Model Evaluation & ValidationHard

Q28. What is the Cohen's Kappa statistic?

  1. A.A metric that accounts for agreement occurring by chance between predicted and actual✓ Correct
  2. B.The exact same calculation as standard overall classification accuracy
  3. C.A regression-only metric for evaluating continuous predictions
  4. D.A feature importance measure for ranking input variables

Explanation

Cohen's Kappa measures inter-rater agreement adjusted for chance, ranging from -1 to 1, more informative than accuracy for imbalanced datasets.

Report an error in this question

Model Evaluation & ValidationHard

Q29. What is calibration in the context of model evaluation?

  1. A.Whether predicted probabilities reflect true likelihoods of outcomes✓ Correct
  2. B.The overall classification accuracy of the trained model
  3. C.A data preprocessing step for cleaning raw data
  4. D.The process of selecting the most relevant features

Explanation

Calibration measures whether a model's predicted probabilities match actual frequencies. A well-calibrated model predicting 70% probability should be correct about 70% of the time.

Report an error in this question

Model Evaluation & ValidationHard

Q30. What is the expected calibration error (ECE)?

  1. A.A regression metric measuring the average absolute difference in predictions
  2. B.The exact same calculation as the standard log loss or cross-entropy loss metric
  3. C.The weighted average of difference between predicted confidence and actual accuracy across bins✓ Correct
  4. D.A clustering metric measuring the silhouette score across all partitions

Explanation

ECE divides predictions into bins by confidence, computes the gap between average confidence and accuracy in each bin, and takes the weighted average, measuring calibration quality.

Report an error in this question

Ready to test yourself on Model Evaluation & Validation?

Take a timed quiz drawn from 210+ questions on this topic. No signup required — your progress saves in your browser.

Start Model Evaluation & Validation Quiz