Each question below shows the correct answer with a full explanation. Use these to build conceptual understanding before attempting a timed quiz.
Ensemble LearningEasy
Q1. What is ensemble learning?
- A.Using a single standalone model for output
- B.Selecting the most relevant input features
- C.Combining multiple models to improve prediction✓ Correct
- D.Cleaning and preprocessing raw input data
Explanation
Ensemble learning combines predictions from multiple models to produce more accurate and robust results than any single model.
Report an error in this question
Ensemble LearningEasy
Q2. Random Forest is an ensemble of:
- A.Neural networks
- B.Linear models
- C.Decision trees✓ Correct
- D.SVMs
Explanation
Random Forest builds multiple decision trees and combines their predictions through majority voting (classification) or averaging (regression).
Report an error in this question
Ensemble LearningEasy
Q3. What is bagging in ensemble learning?
- A.Training models on random subsets of data with replacement✓ Correct
- B.Cleaning and imputing missing data point values
- C.Training a single model on the complete full dataset
- D.Removing irrelevant features from the training data
Explanation
Bagging (Bootstrap Aggregating) trains multiple models on different random subsets of the training data drawn with replacement.
Report an error in this question
Ensemble LearningEasy
Q4. What is boosting?
- A.Sequentially training models where each focuses on previous errors✓ Correct
- B.Reducing the overall size of the training dataset
- C.Engineering new features from the existing raw data
- D.Training all individual models simultaneously in parallel
Explanation
Boosting trains models sequentially, with each new model focusing on correcting the errors made by previous models.
Report an error in this question
Ensemble LearningEasy
Q5. The final prediction in a classification ensemble is typically made by:
- A.Using only the first model
- B.Random class selection
- C.Majority voting✓ Correct
- D.Using the weakest model
Explanation
In classification ensembles, the final prediction is typically determined by majority voting among all the individual models.
Report an error in this question
Ensemble LearningEasy
Q6. Which of these is a popular boosting algorithm?
- A.DBSCAN
- B.K-Means
- C.PCA
- D.XGBoost✓ Correct
Explanation
XGBoost (Extreme Gradient Boosting) is one of the most popular and effective boosting algorithms.
Report an error in this question
Ensemble LearningEasy
Q7. How does Random Forest handle overfitting?
- A.By removing all features from the training input
- B.By using a single very deep decision tree only
- C.By only adding more raw data samples to train
- D.By averaging predictions from many diverse trees✓ Correct
Explanation
Random Forest reduces overfitting by averaging predictions from many trees, each trained on different data subsets and feature subsets.
Report an error in this question
Ensemble LearningEasy
Q8. What is a base learner in ensemble methods?
- A.A tunable model hyperparameter value
- B.An individual model within the ensemble✓ Correct
- C.A data preprocessing pipeline step
- D.The final aggregated combined model output
Explanation
A base learner is an individual model (often a weak learner) that is combined with others in an ensemble.
Report an error in this question
Ensemble LearningEasy
Q9. AdaBoost stands for:
- A.Automated Boosting
- B.Adaptive Boosting✓ Correct
- C.Advanced Boosting
- D.Additional Boosting
Explanation
AdaBoost (Adaptive Boosting) adjusts sample weights to focus subsequent classifiers on previously misclassified examples.
Report an error in this question
Ensemble LearningEasy
Q10. In Random Forest, what is feature randomness?
- A.Features are randomly generated from scratch each time✓ Correct
- B.Each tree considers a random subset of features at each split
- C.Features are removed permanently from the entire dataset
- D.All decision trees use all available features every time
Explanation
Feature randomness means each tree in the forest considers only a random subset of features when making splits, increasing diversity among trees.
Report an error in this question
Ensemble LearningMedium
Q11. What is the difference between bagging and boosting?
- A.Bagging and boosting are completely identical approaches
- B.Boosting is always faster than bagging in every scenario
- C.Bagging always performs better than boosting overall✓ Correct
- D.Bagging trains models independently in parallel; boosting trains them sequentially on errors
Explanation
Bagging trains independent models in parallel on bootstrap samples, while boosting trains sequentially with each model correcting previous errors.
Report an error in this question
Ensemble LearningMedium
Q12. In AdaBoost, how are misclassified samples handled?
- A.They are completely removed from the data
- B.Their weights are increased for the next iteration
- C.Their weights are decreased in the next round✓ Correct
- D.They are duplicated in the training dataset
Explanation
AdaBoost increases the weights of misclassified samples so subsequent learners focus more on the difficult examples.
Report an error in this question
Ensemble LearningMedium
Q13. What is stacking in ensemble learning?
- A.Using only a single standalone model for all predictions
- B.Random feature selection from the available input columns
- C.Using a meta-model to combine predictions from multiple base models✓ Correct
- D.Data augmentation to increase the training set size
Explanation
Stacking trains a meta-learner on the predictions of base models, learning the optimal way to combine their outputs.
Report an error in this question
Ensemble LearningMedium
Q14. What is the out-of-bag (OOB) error in Random Forest?
- A.The overall training error computed on the full training dataset
- B.The test error computed on a held-out independent test set
- C.The validation error from a separate cross-validation fold
- D.Error estimated using samples not included in each tree's bootstrap sample✓ Correct
Explanation
OOB error uses samples not selected in each tree's bootstrap sample as a natural validation set, providing an unbiased error estimate without a separate test set.
Report an error in this question
Ensemble LearningMedium
Q15. Gradient Boosting minimizes the loss function by:
- A.Removing the least informative features iteratively
- B.Increasing the overall size of the training dataset
- C.Adding trees that fit the negative gradient of the loss✓ Correct
- D.Random bootstrap sampling of the original dataset
Explanation
Gradient Boosting fits each new tree to the negative gradient (pseudo-residuals) of the loss function, performing gradient descent in function space.
Report an error in this question
Ensemble LearningMedium
Q16. What is the max_features parameter in Random Forest?
- A.The minimum number of samples per leaf node
- B.The number of features to consider at each split
- C.The maximum depth of each individual tree✓ Correct
- D.The maximum total number of trees in the forest
Explanation
max_features controls how many features each tree considers when looking for the best split, affecting model diversity and performance.
Report an error in this question
Ensemble LearningMedium
Q17. Why does ensemble learning generally outperform single models?
- A.It reduces variance and/or bias by combining diverse models✓ Correct
- B.It always uses significantly less training data overall
- C.It is always computationally faster than a single model
- D.It always requires fewer input features to train on
Explanation
Ensembles reduce variance (bagging), bias (boosting), or both (stacking) by combining diverse models that make different types of errors.
Report an error in this question
Ensemble LearningMedium
Q18. What is a weak learner?
- A.A model that performs slightly better than random guessing✓ Correct
- B.A model that consistently achieves 100% accuracy
- C.A model that never makes any classification errors
- D.A model with absolutely no trainable parameters
Explanation
A weak learner is a model that performs only slightly better than random chance, yet can be combined in ensembles to create strong learners.
Report an error in this question
Ensemble LearningMedium
Q19. In XGBoost, what is the purpose of the learning rate?
- A.It controls the contribution of each tree to shrink step size✓ Correct
- B.It directly determines the maximum depth of each tree
- C.It selects which features are used at each tree split
- D.It explicitly sets the total number of trees in the forest
Explanation
The learning rate (eta) shrinks the contribution of each tree, requiring more trees but providing better generalization and preventing overfitting.
Report an error in this question
Ensemble LearningMedium
Q20. What is voting in ensemble methods?
- A.Training a single standalone model on the complete dataset
- B.Engineering new derived features from the raw input data
- C.Combining predictions by having each model vote on the outcome✓ Correct
- D.Removing outlier data points from the training samples
Explanation
Voting combines predictions from multiple models. Hard voting uses majority class; soft voting averages predicted probabilities.
Report an error in this question
Ensemble LearningHard
Q21. What is the bias-variance decomposition of ensemble methods?
- A.Bagging primarily reduces variance; boosting primarily reduces bias✓ Correct
- B.Neither technique affects the bias or variance components
- C.Both techniques only reduce the variance component of error
- D.Both techniques only reduce the bias component of error
Explanation
Bagging reduces variance by averaging independent models. Boosting reduces bias by sequentially fitting residuals. Both can improve overall performance.
Report an error in this question
Ensemble LearningHard
Q22. How does XGBoost handle regularization differently from traditional Gradient Boosting?
- A.XGBoost uses no regularization at all in its training objective
- B.Traditional Gradient Boosting has stronger regularization overall
- C.They handle regularization in the exact same identical manner
- D.XGBoost includes L1 and L2 regularization terms in its objective function✓ Correct
Explanation
XGBoost explicitly adds L1 and L2 regularization terms to its objective function, controlling model complexity and reducing overfitting more effectively.
Report an error in this question
Ensemble LearningHard
Q23. What is the difference between XGBoost and LightGBM's tree growing strategy?
- A.Neither of them actually uses tree models
- B.XGBoost grows level-wise; LightGBM grows leaf-wise
- C.They grow trees in an identical manner always
- D.XGBoost uses leaf-wise; LightGBM uses level-wise✓ Correct
Explanation
XGBoost grows trees level-wise (breadth-first), while LightGBM grows leaf-wise (choosing the leaf with max loss reduction), making LightGBM faster on large datasets.
Report an error in this question
Ensemble LearningHard
Q24. In CatBoost, how are categorical features handled?
- A.They are silently ignored during the training and inference steps
- B.They are completely removed from the feature set before training
- C.They must be manually one-hot encoded before model training begins
- D.Using ordered target statistics with random permutations to avoid target leakage✓ Correct
Explanation
CatBoost uses ordered target statistics with random permutations to encode categorical features, preventing target leakage that traditional target encoding can cause.
Report an error in this question
Ensemble LearningHard
Q25. What is the effect of increasing the number of trees in a Random Forest?
- A.It always causes the model to underfit the training data
- B.Performance improves then plateaus; it generally does not overfit✓ Correct
- C.It always causes the model to overfit the training data
- D.Performance always decreases with each additional tree
Explanation
Unlike boosting, adding more trees to a Random Forest improves performance up to a point and then plateaus without overfitting, due to the law of large numbers.
Report an error in this question
Ensemble LearningHard
Q26. What is the role of the subsample parameter in Gradient Boosting?
- A.It sets the step size or learning rate of the optimization
- B.It selects which features are considered at each individual tree split
- C.It introduces stochastic gradient boosting by using a fraction of samples per tree✓ Correct
- D.It controls the maximum allowed depth of each individual tree
Explanation
Subsample controls the fraction of training data used for each tree, introducing randomness (stochastic gradient boosting) that reduces overfitting and can improve generalization.
Report an error in this question
Ensemble LearningHard
Q27. How does the isolation mechanism work in Isolation Forest for ensemble anomaly detection?
- A.Anomalies require significantly more random splits to isolate
- B.The splits are entirely random and have no diagnostic meaning
- C.Anomalies are isolated in fewer splits because they are rare and different✓ Correct
- D.All data points require an exactly equal number of splits
Explanation
Isolation Forest isolates anomalies faster (fewer splits) because outliers are few and have attribute values that differ significantly from normal instances.
Report an error in this question
Ensemble LearningHard
Q28. What is Bayesian Model Averaging?
- A.Random model selection from the available ensemble
- B.Simple majority voting across all base model predictions
- C.Weighting ensemble models by their posterior probabilities✓ Correct
- D.Using only the single highest-accuracy base model
Explanation
Bayesian Model Averaging weights each model's predictions by its posterior probability given the data, providing a principled approach to ensemble combination.
Report an error in this question
Ensemble LearningHard
Q29. What is negative correlation learning in ensemble methods?
- A.Using the exact same training data for every individual model
- B.Removing highly correlated input features from the dataset
- C.Training all base models to produce completely identical outputs
- D.Encouraging base learners to make diverse errors through a penalty term✓ Correct
Explanation
Negative correlation learning adds a penalty term during training that encourages base learners to make diverse, uncorrelated errors, improving ensemble performance.
Report an error in this question
Ensemble LearningHard
Q30. How does DART (Dropouts meet Multiple Additive Regression Trees) improve boosting?
- A.By using significantly deeper trees for each individual boosting iteration
- B.By removing the least important features entirely from the training data✓ Correct
- C.By randomly dropping trees during boosting iterations to prevent over-specialization
- D.By adding many more trees to the overall boosted ensemble of the model
Explanation
DART applies dropout to boosting by randomly dropping previously built trees during each iteration, preventing individual trees from dominating and reducing overfitting.
Report an error in this question