Each question below shows the correct answer with a full explanation. Use these to build conceptual understanding before attempting a timed quiz.
Data Collection & Pre-processingEasy
Q1. What is data preprocessing?
- A.Deploying trained models into production systems
- B.Writing detailed summary reports on raw findings
- C.Transforming raw data into a clean and usable format✓ Correct
- D.Creating interactive charts to visualize raw data
Explanation
Data preprocessing involves cleaning, transforming, and organizing raw data before analysis or model training.
Report an error in this question
Data Collection & Pre-processingEasy
Q2. What is a missing value in a dataset?
- A.A data point that equals exactly zero
- B.A data point that is absent or null✓ Correct
- C.An extreme statistical outlier point
- D.A strongly negative numerical value
Explanation
Missing values are absent or null entries in a dataset that need to be handled before analysis.
Report an error in this question
Data Collection & Pre-processingEasy
Q3. What is data normalization?
- A.Duplicating data across tables
- B.Sorting data in alphabetical order
- C.Scaling data to a standard range✓ Correct
- D.Removing all data from storage
Explanation
Normalization scales data values to a common range, typically 0 to 1, to ensure features contribute equally.
Report an error in this question
Data Collection & Pre-processingEasy
Q4. Which of these is a common data format?
- A.SYS
- B.EXE
- C.CSV✓ Correct
- D.DLL
Explanation
CSV (Comma-Separated Values) is one of the most common formats for storing and exchanging tabular data.
Report an error in this question
Data Collection & Pre-processingEasy
Q5. What does data cleaning involve?
- A.Making the dataset physically larger in size
- B.Adding more synthetic records to the data
- C.Encrypting data columns for secure storage
- D.Removing errors and inconsistencies from data✓ Correct
Explanation
Data cleaning involves identifying and correcting errors, handling missing values, and removing inconsistencies.
Report an error in this question
Data Collection & Pre-processingEasy
Q6. What is a categorical variable?
- A.A variable that can only hold binary data
- B.A variable with randomly assigned values
- C.A variable with continuous numerical values
- D.A variable with discrete categories or labels✓ Correct
Explanation
Categorical variables represent discrete categories or labels like colors, genders, or product types.
Report an error in this question
Data Collection & Pre-processingEasy
Q7. What is a data source?
- A.A trained machine learning model file
- B.An interactive data visualization chart
- C.A computational optimization algorithm
- D.The origin from which data is collected✓ Correct
Explanation
A data source is the origin or location from which data is gathered, such as databases, APIs, or files.
Report an error in this question
Data Collection & Pre-processingEasy
Q8. What is the purpose of removing duplicates?
- A.To add new input features
- B.To change column data types
- C.To eliminate redundant records✓ Correct
- D.To increase the dataset size
Explanation
Removing duplicates eliminates redundant records that can bias analysis and waste computational resources.
Report an error in this question
Data Collection & Pre-processingEasy
Q9. What is a structured dataset?
- A.Continuous raw audio waveforms
- B.Data organized in rows and columns✓ Correct
- C.Unorganized raw free-form text
- D.Unstructured pixel image data
Explanation
Structured data is organized in a tabular format with rows and columns, like spreadsheets or database tables.
Report an error in this question
Data Collection & Pre-processingEasy
Q10. Which pandas method checks for null values?
- A.fillna()
- B.isnull()✓ Correct
- C.notnull()
- D.dropna()
Explanation
isnull() or isna() returns a boolean mask indicating which values in a DataFrame are null/missing.
Report an error in this question
Data Collection & Pre-processingMedium
Q11. What is one-hot encoding?
- A.Encrypting data columns for secure transmission
- B.Compressing data files for reduced disk usage
- C.Converting categorical variables into binary vectors✓ Correct
- D.Normalizing continuous numerical data columns
Explanation
One-hot encoding creates binary columns for each category, with 1 indicating the presence of that category and 0 otherwise.
Report an error in this question
Data Collection & Pre-processingMedium
Q12. What is the difference between normalization and standardization?
- A.Normalization scales to [0,1], standardization scales to zero mean and unit variance✓ Correct
- B.Normalization is always the superior approach for every ML algorithm
- C.They are identical processes producing the exact same transformed outputs
- D.Standardization is only applicable to integer-valued feature columns
Explanation
Normalization scales data to [0,1] range, while standardization transforms data to have mean=0 and standard deviation=1 (z-score).
Report an error in this question
Data Collection & Pre-processingMedium
Q13. What is label encoding?
- A.Assigning numerical values to categorical labels✓ Correct
- B.Encrypting labels for secure data storage
- C.Adding new labels to the feature space
- D.Removing labels from a dataset entirely
Explanation
Label encoding converts categorical labels into integers (e.g., Red=0, Blue=1, Green=2).
Report an error in this question
Data Collection & Pre-processingMedium
Q14. What is imputation?
- A.Adding new columns to expand the feature set
- B.Replacing missing values with estimated values✓ Correct
- C.Deleting every row containing missing values
- D.Sorting data by ascending numerical order
Explanation
Imputation replaces missing values with substituted values like mean, median, mode, or predicted values.
Report an error in this question
Data Collection & Pre-processingMedium
Q15. Why is feature scaling important for ML algorithms?
- A.It automatically removes all outlier data point values
- B.It adds new synthetic features to expand the model
- C.It makes the overall dataset physically larger in storage
- D.It ensures all features contribute equally regardless of their scale✓ Correct
Explanation
Feature scaling prevents features with larger ranges from dominating the learning process, especially for distance-based algorithms.
Report an error in this question
Data Collection & Pre-processingMedium
Q16. What is data augmentation?
- A.Encrypting data columns for privacy and security
- B.Deleting redundant records from the training set
- C.Compressing data files to reduce storage overhead
- D.Creating new training data by modifying existing data✓ Correct
Explanation
Data augmentation generates new training examples by applying transformations like rotation, flipping, or adding noise to existing data.
Report an error in this question
Data Collection & Pre-processingMedium
Q17. What is the purpose of the train-test split?
- A.To evaluate model performance on unseen data✓ Correct
- B.To increase the total size of the full dataset
- C.To clean and preprocess the raw input data
- D.To visualize the distribution of data points
Explanation
Train-test split separates data into training and testing sets to evaluate how well a model generalizes to unseen data.
Report an error in this question
Data Collection & Pre-processingMedium
Q18. What is an outlier?
- A.A value that is completely missing or absent
- B.A feature that stores categorical text labels
- C.A data point significantly different from others✓ Correct
- D.A record that has been duplicated in error
Explanation
An outlier is a data point that lies far from other observations, potentially due to errors or rare events.
Report an error in this question
Data Collection & Pre-processingMedium
Q19. What does the z-score standardization formula compute?
- A.(x - min) / (max - min) range
- B.x divided by the max value
- C.(x - mean) / standard deviation✓ Correct
- D.log of the value x
Explanation
Z-score standardization transforms each value to (x - mean) / std, centering data at 0 with unit variance.
Report an error in this question
Data Collection & Pre-processingMedium
Q20. What is web scraping in data collection?
- A.Designing interactive web page layouts
- B.Testing web application load performance
- C.Building responsive websites from scratch
- D.Extracting data from websites programmatically✓ Correct
Explanation
Web scraping uses automated scripts to extract data from websites, often using libraries like BeautifulSoup or Scrapy.
Report an error in this question
Data Collection & Pre-processingHard
Q21. What is the SMOTE technique used for?
- A.Reducing the dimensionality of the feature space with PCA
- B.Oversampling the minority class by creating synthetic examples✓ Correct
- C.Undersampling the majority class by randomly removing records
- D.Selecting only the most informative features from the dataset
Explanation
SMOTE (Synthetic Minority Over-sampling Technique) creates synthetic examples of the minority class to address class imbalance.
Report an error in this question
Data Collection & Pre-processingHard
Q22. When should you use target encoding instead of one-hot encoding?
- A.When the overall dataset is small
- B.When there are only two categories present✓ Correct
- C.When categorical variables have high cardinality
- D.When the data is purely numerical
Explanation
Target encoding is preferred for high-cardinality categorical features to avoid creating too many columns, as one-hot encoding would.
Report an error in this question
Data Collection & Pre-processingHard
Q23. What is data leakage in preprocessing?
- A.When data columns are encrypted for secure storage
- B.When data records are accidentally duplicated twice
- C.When data is physically lost during the cleaning step
- D.When information from the test set influences training✓ Correct
Explanation
Data leakage occurs when information from outside the training set improperly influences the model, leading to overly optimistic performance estimates.
Report an error in this question
Data Collection & Pre-processingHard
Q24. What is the purpose of Yeo-Johnson transformation?
- A.To detect and remove extreme outlier values from the numerical columns
- B.To encode high-cardinality categorical features into integer representations
- C.To make data more normally distributed, handling both positive and negative values✓ Correct
- D.To fill in missing values using the mean or median of each column
Explanation
Yeo-Johnson transformation is a power transformation that makes data more Gaussian-like and works with both positive and negative values, unlike Box-Cox.
Report an error in this question
Data Collection & Pre-processingHard
Q25. In handling missing data, what is MCAR?
- A.Missing data that follows a clear and observable systematic pattern
- B.Missing Completely At Random - missingness is unrelated to any variable✓ Correct
- C.Missing data that occurs in only one specific column of data
- D.Missing data that was caused by systematic data entry errors
Explanation
MCAR means the probability of a value being missing is completely random and unrelated to observed or unobserved data.
Report an error in this question
Data Collection & Pre-processingHard
Q26. What is the KNN imputation method?
- A.Replacing missing values with randomly generated numerical estimates
- B.Deleting every row that contains any missing values from the dataset
- C.Filling all missing entries with the single overall global mean value
- D.Using K nearest neighbors to estimate missing values based on similar records✓ Correct
Explanation
KNN imputation estimates missing values by taking the weighted average of the K nearest neighbors in the feature space.
Report an error in this question
Data Collection & Pre-processingHard
Q27. Why should standardization be fit only on training data?
- A.Because training data is inherently more important
- B.Because the computation is significantly faster
- C.To prevent data leakage from test set statistics✓ Correct
- D.Because the test data is a much smaller sample
Explanation
Fitting standardization only on training data prevents test set information from leaking into the preprocessing pipeline, ensuring unbiased evaluation.
Report an error in this question
Data Collection & Pre-processingHard
Q28. What is the Isolation Forest algorithm used for in preprocessing?
- A.Automated feature selection
- B.Anomaly and outlier detection✓ Correct
- C.Data range normalization
- D.Missing value imputation
Explanation
Isolation Forest detects anomalies by randomly partitioning data; outliers are isolated in fewer partitions than normal points.
Report an error in this question
Data Collection & Pre-processingHard
Q29. What is ordinal encoding and when is it preferred over one-hot encoding?
- A.It assigns ordered integers to categories with natural ordering✓ Correct
- B.It is exactly the same approach as basic label encoding
- C.It removes categorical variables from the feature set entirely
- D.It always creates separate binary columns for each category
Explanation
Ordinal encoding assigns integers preserving the natural order of categories (e.g., low=1, medium=2, high=3), preferred when categories have meaningful ordering.
Report an error in this question
Data Collection & Pre-processingHard
Q30. What is the effect of multicollinearity on preprocessing?
- A.It has absolutely no effect on preprocessing or model training
- B.Highly correlated features can cause instability in model coefficients✓ Correct
- C.It always improves the overall accuracy of the trained model
- D.It consistently reduces the total required model training time
Explanation
Multicollinearity means features are highly correlated, causing unstable coefficient estimates in linear models and requiring techniques like VIF analysis or PCA.
Report an error in this question