HomeSubjectsUniversityBlogAbout

Data Collection & Pre-processing

Topic in AI / Machine Learning & Data Analytics

210 total MCQsShowing 30 with explanations10 Easy10 Medium10 Hard

About This Topic

Data preprocessing is the cleaning and transformation of raw data into a consistent, numeric form that machine learning algorithms can learn from reliably. Questions begin with data sources and formats like CSV, JSON and SQL tables, then move to handling missing values through deletion, mean or median imputation, KNN imputation and multivariate or Bayesian imputation. Expect items on outlier detection, removing duplicates, scaling with min-max normalization or z-score standardization, and encoding categories with label or one-hot encoding. Class imbalance is also common, with techniques such as undersampling, oversampling and SMOTE, which synthesizes new minority-class examples by interpolation.

Below are 30 practice questions from a pool of 210 Data Collection & Pre-processing MCQs, one of 17 topics in AI / Machine Learning & Data Analytics. Each shows the correct answer with an explanation; when you are ready, take a timed quiz to test recall under exam conditions.

Practice Questions

Each question below shows the correct answer with a full explanation. Use these to build conceptual understanding before attempting a timed quiz.

Data Collection & Pre-processingEasy

Q1. What is data preprocessing?

  1. A.Deploying trained models into production systems
  2. B.Writing detailed summary reports on raw findings
  3. C.Transforming raw data into a clean and usable format✓ Correct
  4. D.Creating interactive charts to visualize raw data

Explanation

Data preprocessing involves cleaning, transforming, and organizing raw data before analysis or model training.

Report an error in this question

Data Collection & Pre-processingEasy

Q2. What is a missing value in a dataset?

  1. A.A data point that equals exactly zero
  2. B.A data point that is absent or null✓ Correct
  3. C.An extreme statistical outlier point
  4. D.A strongly negative numerical value

Explanation

Missing values are absent or null entries in a dataset that need to be handled before analysis.

Report an error in this question

Data Collection & Pre-processingEasy

Q3. What is data normalization?

  1. A.Duplicating data across tables
  2. B.Sorting data in alphabetical order
  3. C.Scaling data to a standard range✓ Correct
  4. D.Removing all data from storage

Explanation

Normalization scales data values to a common range, typically 0 to 1, to ensure features contribute equally.

Report an error in this question

Data Collection & Pre-processingEasy

Q4. Which of these is a common data format?

  1. A.SYS
  2. B.EXE
  3. C.CSV✓ Correct
  4. D.DLL

Explanation

CSV (Comma-Separated Values) is one of the most common formats for storing and exchanging tabular data.

Report an error in this question

Data Collection & Pre-processingEasy

Q5. What does data cleaning involve?

  1. A.Making the dataset physically larger in size
  2. B.Adding more synthetic records to the data
  3. C.Encrypting data columns for secure storage
  4. D.Removing errors and inconsistencies from data✓ Correct

Explanation

Data cleaning involves identifying and correcting errors, handling missing values, and removing inconsistencies.

Report an error in this question

Data Collection & Pre-processingEasy

Q6. What is a categorical variable?

  1. A.A variable that can only hold binary data
  2. B.A variable with randomly assigned values
  3. C.A variable with continuous numerical values
  4. D.A variable with discrete categories or labels✓ Correct

Explanation

Categorical variables represent discrete categories or labels like colors, genders, or product types.

Report an error in this question

Data Collection & Pre-processingEasy

Q7. What is a data source?

  1. A.A trained machine learning model file
  2. B.An interactive data visualization chart
  3. C.A computational optimization algorithm
  4. D.The origin from which data is collected✓ Correct

Explanation

A data source is the origin or location from which data is gathered, such as databases, APIs, or files.

Report an error in this question

Data Collection & Pre-processingEasy

Q8. What is the purpose of removing duplicates?

  1. A.To add new input features
  2. B.To change column data types
  3. C.To eliminate redundant records✓ Correct
  4. D.To increase the dataset size

Explanation

Removing duplicates eliminates redundant records that can bias analysis and waste computational resources.

Report an error in this question

Data Collection & Pre-processingEasy

Q9. What is a structured dataset?

  1. A.Continuous raw audio waveforms
  2. B.Data organized in rows and columns✓ Correct
  3. C.Unorganized raw free-form text
  4. D.Unstructured pixel image data

Explanation

Structured data is organized in a tabular format with rows and columns, like spreadsheets or database tables.

Report an error in this question

Data Collection & Pre-processingEasy

Q10. Which pandas method checks for null values?

  1. A.fillna()
  2. B.isnull()✓ Correct
  3. C.notnull()
  4. D.dropna()

Explanation

isnull() or isna() returns a boolean mask indicating which values in a DataFrame are null/missing.

Report an error in this question

Data Collection & Pre-processingMedium

Q11. What is one-hot encoding?

  1. A.Encrypting data columns for secure transmission
  2. B.Compressing data files for reduced disk usage
  3. C.Converting categorical variables into binary vectors✓ Correct
  4. D.Normalizing continuous numerical data columns

Explanation

One-hot encoding creates binary columns for each category, with 1 indicating the presence of that category and 0 otherwise.

Report an error in this question

Data Collection & Pre-processingMedium

Q12. What is the difference between normalization and standardization?

  1. A.Normalization scales to [0,1], standardization scales to zero mean and unit variance✓ Correct
  2. B.Normalization is always the superior approach for every ML algorithm
  3. C.They are identical processes producing the exact same transformed outputs
  4. D.Standardization is only applicable to integer-valued feature columns

Explanation

Normalization scales data to [0,1] range, while standardization transforms data to have mean=0 and standard deviation=1 (z-score).

Report an error in this question

Data Collection & Pre-processingMedium

Q13. What is label encoding?

  1. A.Assigning numerical values to categorical labels✓ Correct
  2. B.Encrypting labels for secure data storage
  3. C.Adding new labels to the feature space
  4. D.Removing labels from a dataset entirely

Explanation

Label encoding converts categorical labels into integers (e.g., Red=0, Blue=1, Green=2).

Report an error in this question

Data Collection & Pre-processingMedium

Q14. What is imputation?

  1. A.Adding new columns to expand the feature set
  2. B.Replacing missing values with estimated values✓ Correct
  3. C.Deleting every row containing missing values
  4. D.Sorting data by ascending numerical order

Explanation

Imputation replaces missing values with substituted values like mean, median, mode, or predicted values.

Report an error in this question

Data Collection & Pre-processingMedium

Q15. Why is feature scaling important for ML algorithms?

  1. A.It automatically removes all outlier data point values
  2. B.It adds new synthetic features to expand the model
  3. C.It makes the overall dataset physically larger in storage
  4. D.It ensures all features contribute equally regardless of their scale✓ Correct

Explanation

Feature scaling prevents features with larger ranges from dominating the learning process, especially for distance-based algorithms.

Report an error in this question

Data Collection & Pre-processingMedium

Q16. What is data augmentation?

  1. A.Encrypting data columns for privacy and security
  2. B.Deleting redundant records from the training set
  3. C.Compressing data files to reduce storage overhead
  4. D.Creating new training data by modifying existing data✓ Correct

Explanation

Data augmentation generates new training examples by applying transformations like rotation, flipping, or adding noise to existing data.

Report an error in this question

Data Collection & Pre-processingMedium

Q17. What is the purpose of the train-test split?

  1. A.To evaluate model performance on unseen data✓ Correct
  2. B.To increase the total size of the full dataset
  3. C.To clean and preprocess the raw input data
  4. D.To visualize the distribution of data points

Explanation

Train-test split separates data into training and testing sets to evaluate how well a model generalizes to unseen data.

Report an error in this question

Data Collection & Pre-processingMedium

Q18. What is an outlier?

  1. A.A value that is completely missing or absent
  2. B.A feature that stores categorical text labels
  3. C.A data point significantly different from others✓ Correct
  4. D.A record that has been duplicated in error

Explanation

An outlier is a data point that lies far from other observations, potentially due to errors or rare events.

Report an error in this question

Data Collection & Pre-processingMedium

Q19. What does the z-score standardization formula compute?

  1. A.(x - min) / (max - min) range
  2. B.x divided by the max value
  3. C.(x - mean) / standard deviation✓ Correct
  4. D.log of the value x

Explanation

Z-score standardization transforms each value to (x - mean) / std, centering data at 0 with unit variance.

Report an error in this question

Data Collection & Pre-processingMedium

Q20. What is web scraping in data collection?

  1. A.Designing interactive web page layouts
  2. B.Testing web application load performance
  3. C.Building responsive websites from scratch
  4. D.Extracting data from websites programmatically✓ Correct

Explanation

Web scraping uses automated scripts to extract data from websites, often using libraries like BeautifulSoup or Scrapy.

Report an error in this question

Data Collection & Pre-processingHard

Q21. What is the SMOTE technique used for?

  1. A.Reducing the dimensionality of the feature space with PCA
  2. B.Oversampling the minority class by creating synthetic examples✓ Correct
  3. C.Undersampling the majority class by randomly removing records
  4. D.Selecting only the most informative features from the dataset

Explanation

SMOTE (Synthetic Minority Over-sampling Technique) creates synthetic examples of the minority class to address class imbalance.

Report an error in this question

Data Collection & Pre-processingHard

Q22. When should you use target encoding instead of one-hot encoding?

  1. A.When the overall dataset is small
  2. B.When there are only two categories present✓ Correct
  3. C.When categorical variables have high cardinality
  4. D.When the data is purely numerical

Explanation

Target encoding is preferred for high-cardinality categorical features to avoid creating too many columns, as one-hot encoding would.

Report an error in this question

Data Collection & Pre-processingHard

Q23. What is data leakage in preprocessing?

  1. A.When data columns are encrypted for secure storage
  2. B.When data records are accidentally duplicated twice
  3. C.When data is physically lost during the cleaning step
  4. D.When information from the test set influences training✓ Correct

Explanation

Data leakage occurs when information from outside the training set improperly influences the model, leading to overly optimistic performance estimates.

Report an error in this question

Data Collection & Pre-processingHard

Q24. What is the purpose of Yeo-Johnson transformation?

  1. A.To detect and remove extreme outlier values from the numerical columns
  2. B.To encode high-cardinality categorical features into integer representations
  3. C.To make data more normally distributed, handling both positive and negative values✓ Correct
  4. D.To fill in missing values using the mean or median of each column

Explanation

Yeo-Johnson transformation is a power transformation that makes data more Gaussian-like and works with both positive and negative values, unlike Box-Cox.

Report an error in this question

Data Collection & Pre-processingHard

Q25. In handling missing data, what is MCAR?

  1. A.Missing data that follows a clear and observable systematic pattern
  2. B.Missing Completely At Random - missingness is unrelated to any variable✓ Correct
  3. C.Missing data that occurs in only one specific column of data
  4. D.Missing data that was caused by systematic data entry errors

Explanation

MCAR means the probability of a value being missing is completely random and unrelated to observed or unobserved data.

Report an error in this question

Data Collection & Pre-processingHard

Q26. What is the KNN imputation method?

  1. A.Replacing missing values with randomly generated numerical estimates
  2. B.Deleting every row that contains any missing values from the dataset
  3. C.Filling all missing entries with the single overall global mean value
  4. D.Using K nearest neighbors to estimate missing values based on similar records✓ Correct

Explanation

KNN imputation estimates missing values by taking the weighted average of the K nearest neighbors in the feature space.

Report an error in this question

Data Collection & Pre-processingHard

Q27. Why should standardization be fit only on training data?

  1. A.Because training data is inherently more important
  2. B.Because the computation is significantly faster
  3. C.To prevent data leakage from test set statistics✓ Correct
  4. D.Because the test data is a much smaller sample

Explanation

Fitting standardization only on training data prevents test set information from leaking into the preprocessing pipeline, ensuring unbiased evaluation.

Report an error in this question

Data Collection & Pre-processingHard

Q28. What is the Isolation Forest algorithm used for in preprocessing?

  1. A.Automated feature selection
  2. B.Anomaly and outlier detection✓ Correct
  3. C.Data range normalization
  4. D.Missing value imputation

Explanation

Isolation Forest detects anomalies by randomly partitioning data; outliers are isolated in fewer partitions than normal points.

Report an error in this question

Data Collection & Pre-processingHard

Q29. What is ordinal encoding and when is it preferred over one-hot encoding?

  1. A.It assigns ordered integers to categories with natural ordering✓ Correct
  2. B.It is exactly the same approach as basic label encoding
  3. C.It removes categorical variables from the feature set entirely
  4. D.It always creates separate binary columns for each category

Explanation

Ordinal encoding assigns integers preserving the natural order of categories (e.g., low=1, medium=2, high=3), preferred when categories have meaningful ordering.

Report an error in this question

Data Collection & Pre-processingHard

Q30. What is the effect of multicollinearity on preprocessing?

  1. A.It has absolutely no effect on preprocessing or model training
  2. B.Highly correlated features can cause instability in model coefficients✓ Correct
  3. C.It always improves the overall accuracy of the trained model
  4. D.It consistently reduces the total required model training time

Explanation

Multicollinearity means features are highly correlated, causing unstable coefficient estimates in linear models and requiring techniques like VIF analysis or PCA.

Report an error in this question

Ready to test yourself on Data Collection & Pre-processing?

Take a timed quiz drawn from 210+ questions on this topic. No signup required — your progress saves in your browser.

Start Data Collection & Pre-processing Quiz