Each question below shows the correct answer with a full explanation. Use these to build conceptual understanding before attempting a timed quiz.
Model Evaluation & ValidationEasy
Q1. Accuracy is defined as:
- A.Number of correct predictions divided by total predictions✓ Correct
- B.Only the count of true negative classification results
- C.The total number of features used in the model
- D.Only the count of true positive classification results
Explanation
Accuracy = (correct predictions) / (total predictions), measuring the overall proportion of correct classifications.
Report an error in this question
Model Evaluation & ValidationEasy
Q2. What is a confusion matrix?
- A.A Python data visualization chart library
- B.A type of multi-layer deep neural network
- C.A preprocessing data cleaning utility tool
- D.A table showing predicted vs actual classifications✓ Correct
Explanation
A confusion matrix is a table that summarizes the performance of a classification model by showing true positives, true negatives, false positives, and false negatives.
Report an error in this question
Model Evaluation & ValidationEasy
Q3. What is a training set used for?
- A.Deploying the model to production
- B.Training the model to learn patterns✓ Correct
- C.Evaluating the final performance
- D.Tuning the model hyperparameters
Explanation
The training set is used to train the model by allowing it to learn patterns and relationships in the data.
Report an error in this question
Model Evaluation & ValidationEasy
Q4. What is a test set used for?
- A.Training the model on labeled training samples
- B.Tuning the model's hyperparameter settings
- C.Cleaning and preprocessing the raw data
- D.Evaluating the model's performance on unseen data✓ Correct
Explanation
The test set evaluates how well the trained model generalizes to new, unseen data.
Report an error in this question
Model Evaluation & ValidationEasy
Q5. Mean Squared Error (MSE) is used to evaluate:
- A.Regression models✓ Correct
- B.Classification models only
- C.Clustering algorithms
- D.Association rule mining
Explanation
MSE measures the average squared difference between predicted and actual values, commonly used to evaluate regression models.
Report an error in this question
Model Evaluation & ValidationEasy
Q6. What is a validation set?
- A.The same set as the independent held-out test set
- B.The entire full dataset including all partitions
- C.A preprocessing data cleaning utility pipeline
- D.A subset used to tune hyperparameters during training✓ Correct
Explanation
A validation set is separate from training and test sets, used to tune hyperparameters and make model selection decisions during development.
Report an error in this question
Model Evaluation & ValidationEasy
Q7. What is precision in classification?
- A.The recall evaluation metric✓ Correct
- B.Overall model accuracy metric
- C.True Positives divided by the total count
- D.True Positives / (True Positives + False Positives)
Explanation
Precision is the proportion of positive predictions that are actually correct: TP / (TP + FP).
Report an error in this question
Model Evaluation & ValidationEasy
Q8. What is a false positive?
- A.Making a correct and accurate model prediction
- B.Predicting positive when the actual value is negative✓ Correct
- C.Encountering a missing value in the raw data
- D.Predicting negative when the actual value is positive
Explanation
A false positive (Type I error) occurs when the model incorrectly predicts the positive class.
Report an error in this question
Model Evaluation & ValidationEasy
Q9. R-squared measures:
- A.The proportion of variance in the target explained by the model✓ Correct
- B.The total number of data samples in the dataset
- C.The total number of input features in the training dataset
- D.The total elapsed time required for model training
Explanation
R-squared (coefficient of determination) measures how much of the variance in the dependent variable is explained by the model, ranging from 0 to 1.
Report an error in this question
Model Evaluation & ValidationEasy
Q10. What is recall (sensitivity)?
- A.True Positives / (True Positives + False Positives)
- B.True Positives / (True Positives + False Negatives)
- C.The specificity metric
- D.The overall accuracy metric✓ Correct
Explanation
Recall is the proportion of actual positives correctly identified: TP / (TP + FN).
Report an error in this question
Model Evaluation & ValidationMedium
Q11. What does AUC-ROC measure?
- A.Only the precision metric of the trained classification model✓ Correct
- B.Only the accuracy metric of the trained classification model
- C.Only the recall metric of the trained classification model
- D.The overall ability of the model to discriminate between classes
Explanation
AUC-ROC (Area Under the ROC Curve) measures the model's ability to distinguish between classes across all thresholds. A value of 1.0 is perfect.
Report an error in this question
Model Evaluation & ValidationMedium
Q12. What does the ROC curve plot?
- A.Training error values plotted against validation error values
- B.True Positive Rate vs False Positive Rate at various thresholds✓ Correct
- C.Precision values plotted against corresponding recall values
- D.Training accuracy values plotted against the loss values
Explanation
The ROC (Receiver Operating Characteristic) curve plots TPR against FPR at various classification thresholds.
Report an error in this question
Model Evaluation & ValidationMedium
Q13. What is the F1-score?
- A.The arithmetic mean of precision and recall✓ Correct
- B.The product of precision and recall values
- C.The harmonic mean of precision and recall
- D.The square of the model accuracy
Explanation
F1-score = 2 * (precision * recall) / (precision + recall), providing a balanced measure when precision and recall are both important.
Report an error in this question
Model Evaluation & ValidationMedium
Q14. What is the difference between micro and macro averaging?
- A.Micro averaging is always a better choice than macro averaging✓ Correct
- B.Micro aggregates all instances globally; macro averages per-class metrics equally
- C.Micro and macro averaging are completely identical approaches
- D.Macro averaging ignores the overall class distribution entirely
Explanation
Micro-averaging aggregates all instances globally, while macro-averaging computes metrics per class and averages them, giving equal weight to each class.
Report an error in this question
Model Evaluation & ValidationMedium
Q15. When is recall more important than precision?
- A.When simple accuracy alone is a sufficient metric
- B.When the dataset classes are perfectly balanced overall
- C.When false positives are very costly to the outcome✓ Correct
- D.When missing positive cases is costly, like disease detection
Explanation
Recall is prioritized when false negatives are costly, such as disease screening where missing a positive case could be life-threatening.
Report an error in this question
Model Evaluation & ValidationMedium
Q16. What is stratified cross-validation?
- A.Performing completely random splitting of the data
- B.Ignoring overall class balance during data splitting✓ Correct
- C.Using only the majority class for model training purposes
- D.K-fold CV that preserves the class distribution in each fold
Explanation
Stratified CV ensures each fold has approximately the same proportion of each class as the original dataset, important for imbalanced data.
Report an error in this question
Model Evaluation & ValidationHard
Q17. What is the difference between Type I and Type II errors?
- A.The two types are completely identical errors
- B.Type I is a false positive; Type II is a false negative✓ Correct
- C.Type I is a false negative; Type II is a false positive
- D.Both are actually true positive outcomes
Explanation
Type I error (false positive) incorrectly rejects a true null hypothesis. Type II error (false negative) fails to reject a false null hypothesis.
Report an error in this question
Model Evaluation & ValidationMedium
Q18. What is the log loss (cross-entropy loss)?
- A.A feature selection method based on mutual information scores
- B.The exact same metric as mean squared error for regression
- C.A metric specifically designed for evaluating cluster quality
- D.A loss function that penalizes confident wrong predictions heavily✓ Correct
Explanation
Log loss measures the performance of a classification model by penalizing predictions based on how far they are from the actual labels, with confident wrong predictions penalized heavily.
Report an error in this question
Model Evaluation & ValidationHard
Q19. What is the Matthews Correlation Coefficient (MCC)?
- A.The exact same metric as standard classification accuracy for balanced data
- B.A metric specifically designed for evaluating unsupervised clustering
- C.A metric that is only applicable to multi-class classification problems
- D.A balanced metric using all four confusion matrix values, ranging from -1 to 1✓ Correct
Explanation
MCC considers TP, TN, FP, and FN, producing a balanced measure even on imbalanced datasets. +1 is perfect, 0 is random, -1 is complete disagreement.
Report an error in this question
Model Evaluation & ValidationMedium
Q20. What is the purpose of a learning curve?
- A.To select the most important features from the training data
- B.To diagnose overfitting or underfitting by plotting performance vs training size✓ Correct
- C.To clean and preprocess the raw input data before training
- D.To learn about new machine learning algorithms from documentation
Explanation
A learning curve plots training and validation performance as a function of training set size, helping diagnose whether the model suffers from high bias or high variance.
Report an error in this question
Model Evaluation & ValidationMedium
Q21. What is K-fold cross-validation?
- A.Dividing data into K folds and using each fold as validation once
- B.Training the model completely without any validation step✓ Correct
- C.Using all available data solely for the training process
- D.Using only two simple train-test splits for evaluation
Explanation
K-fold CV splits data into K equal folds, trains K times using K-1 folds for training and the remaining fold for validation, averaging the results.
Report an error in this question
Model Evaluation & ValidationMedium
Q22. What is Mean Absolute Error (MAE)?
- A.The single maximum prediction error value✓ Correct
- B.The median of all prediction errors
- C.The average of absolute differences between predicted and actual values
- D.The average of all squared differences between predictions
Explanation
MAE = mean(|predicted - actual|). It measures the average magnitude of errors without considering direction, less sensitive to outliers than MSE.
Report an error in this question
Model Evaluation & ValidationHard
Q23. What is the difference between leave-one-out CV and K-fold CV?
- A.They are exactly identical approaches with no meaningful differences whatsoever
- B.LOOCV is always significantly faster computationally than K-fold approaches
- C.K-fold always produces strictly more accurate results than LOOCV methods
- D.LOOCV uses N-1 samples for training and 1 for testing each time; K-fold uses larger folds✓ Correct
Explanation
LOOCV is a special case where K=N (number of samples), using each sample once as validation. It has low bias but high variance and is computationally expensive.
Report an error in this question
Model Evaluation & ValidationHard
Q24. What is the purpose of bootstrapping in model evaluation?
- A.Performing automated feature selection and dimensionality reduction
- B.Estimating confidence intervals for metrics by resampling with replacement✓ Correct
- C.Running hyperparameter tuning using grid or random search
- D.Increasing the total amount of available training data samples
Explanation
Bootstrapping creates multiple resampled datasets to estimate the distribution of a statistic, providing confidence intervals for model performance metrics.
Report an error in this question
Model Evaluation & ValidationHard
Q25. When should you use the PR-AUC instead of ROC-AUC?
- A.For evaluating unsupervised clustering partition quality
- B.When dealing with highly imbalanced datasets where positive class is rare✓ Correct
- C.When the class distribution is perfectly balanced and equal
- D.For evaluating regression models with continuous outputs
Explanation
PR-AUC is preferred for highly imbalanced datasets because ROC-AUC can be misleadingly high when the negative class dominates, while PR-AUC focuses on the minority positive class.
Report an error in this question
Model Evaluation & ValidationHard
Q26. What is nested cross-validation?
- A.An outer CV loop for evaluation with an inner CV loop for hyperparameter tuning✓ Correct
- B.A single random train-test split without any cross-validation folds
- C.Cross-validation performed entirely without any held-out validation data
- D.The same standard approach as regular single-loop K-fold cross-validation
Explanation
Nested CV uses an outer loop for unbiased performance estimation and an inner loop for hyperparameter tuning, preventing optimistic bias from using the same data for both.
Report an error in this question
Model Evaluation & ValidationHard
Q27. What is the Brier score?
- A.A feature importance measure for trees
- B.The mean squared error of probabilistic predictions✓ Correct
- C.A ranking metric for ordering search results
- D.A clustering score for partition quality
Explanation
The Brier score measures the accuracy of probabilistic predictions by computing the mean squared difference between predicted probabilities and actual outcomes (0 or 1).
Report an error in this question
Model Evaluation & ValidationHard
Q28. What is the Cohen's Kappa statistic?
- A.A metric that accounts for agreement occurring by chance between predicted and actual✓ Correct
- B.The exact same calculation as standard overall classification accuracy
- C.A regression-only metric for evaluating continuous predictions
- D.A feature importance measure for ranking input variables
Explanation
Cohen's Kappa measures inter-rater agreement adjusted for chance, ranging from -1 to 1, more informative than accuracy for imbalanced datasets.
Report an error in this question
Model Evaluation & ValidationHard
Q29. What is calibration in the context of model evaluation?
- A.Whether predicted probabilities reflect true likelihoods of outcomes✓ Correct
- B.The overall classification accuracy of the trained model
- C.A data preprocessing step for cleaning raw data
- D.The process of selecting the most relevant features
Explanation
Calibration measures whether a model's predicted probabilities match actual frequencies. A well-calibrated model predicting 70% probability should be correct about 70% of the time.
Report an error in this question
Model Evaluation & ValidationHard
Q30. What is the expected calibration error (ECE)?
- A.A regression metric measuring the average absolute difference in predictions
- B.The exact same calculation as the standard log loss or cross-entropy loss metric
- C.The weighted average of difference between predicted confidence and actual accuracy across bins✓ Correct
- D.A clustering metric measuring the silhouette score across all partitions
Explanation
ECE divides predictions into bins by confidence, computes the gap between average confidence and accuracy in each bin, and takes the weighted average, measuring calibration quality.
Report an error in this question