HomeSubjectsUniversityBlogAbout

Exploratory Data Analysis

Topic in AI / Machine Learning & Data Analytics

210 total MCQsShowing 30 with explanations10 Easy10 Medium10 Hard

About This Topic

Exploratory data analysis (EDA) is the first look at a dataset through summary statistics and plots to learn its structure, spot anomalies and form hypotheses. Questions cover measures of central tendency and spread, skewness, kurtosis and quartiles, along with chart choices: histograms, box plots, scatter plots, heatmaps and kernel density plots. Correlation items ask what a coefficient of +1, 0 or -1 means, why Pearson correlation assumes a linear relationship, and when Spearman is better. You should also recognize correlation-versus-causation errors, Simpson's paradox, detecting outliers with the IQR rule, and using bootstrapped confidence intervals to judge how stable a statistic is.

Below are 30 practice questions from a pool of 210 Exploratory Data Analysis MCQs, one of 17 topics in AI / Machine Learning & Data Analytics. Each shows the correct answer with an explanation; when you are ready, take a timed quiz to test recall under exam conditions.

Practice Questions

Each question below shows the correct answer with a full explanation. Use these to build conceptual understanding before attempting a timed quiz.

Exploratory Data AnalysisEasy

Q1. What does EDA stand for?

  1. A.Extended Data Algorithm
  2. B.External Data Analysis
  3. C.Effective Data Assessment
  4. D.Exploratory Data Analysis✓ Correct

Explanation

EDA stands for Exploratory Data Analysis, the process of analyzing data sets to summarize their main characteristics.

Report an error in this question

Exploratory Data AnalysisEasy

Q2. Which chart is used to show the distribution of a single variable?

  1. A.Line chart
  2. B.Scatter plot
  3. C.Histogram✓ Correct
  4. D.Pie chart

Explanation

A histogram displays the frequency distribution of a single continuous variable by dividing data into bins.

Report an error in this question

Exploratory Data AnalysisEasy

Q3. What does a box plot display?

  1. A.The five-number summary: min, Q1, median, Q3, max✓ Correct
  2. B.Only the arithmetic mean of the distribution values
  3. C.Only the overall range from minimum to maximum
  4. D.Only the mode or most frequent value observed

Explanation

A box plot visualizes the five-number summary (minimum, Q1, median, Q3, maximum) and helps identify outliers.

Report an error in this question

Exploratory Data AnalysisEasy

Q4. Which pandas method gives summary statistics?

  1. A.describe()✓ Correct
  2. B.info()
  3. C.head()
  4. D.shape

Explanation

df.describe() provides summary statistics including count, mean, std, min, quartiles, and max for numerical columns.

Report an error in this question

Exploratory Data AnalysisEasy

Q5. A scatter plot shows the relationship between:

  1. A.Two numerical variables
  2. B.Categories data only
  3. C.Time series data only
  4. D.Textual data only✓ Correct

Explanation

A scatter plot displays the relationship between two numerical variables, with each point representing one observation.

Report an error in this question

Exploratory Data AnalysisEasy

Q6. What is the purpose of a correlation matrix?

  1. A.To clean and preprocess the raw input data
  2. B.To deploy and manage software applications
  3. C.To train and evaluate machine learning models
  4. D.To show relationships between multiple variables✓ Correct

Explanation

A correlation matrix shows the pairwise correlation coefficients between multiple variables, indicating linear relationships.

Report an error in this question

Exploratory Data AnalysisEasy

Q7. What does df.info() display in pandas?

  1. A.Only the numerical summary statistics values
  2. B.Only the count of missing or null values
  3. C.Column names, data types, and non-null counts✓ Correct
  4. D.Only the first five rows of the DataFrame

Explanation

df.info() displays column names, data types, non-null counts, and memory usage of a DataFrame.

Report an error in this question

Exploratory Data AnalysisEasy

Q8. A bar chart is best used for:

  1. A.Showing data distributions
  2. B.Comparing quantities across categories
  3. C.Displaying feature correlations
  4. D.Plotting time series trends✓ Correct

Explanation

Bar charts are ideal for comparing quantities across different categories using rectangular bars.

Report an error in this question

Exploratory Data AnalysisEasy

Q9. What is the mode of a dataset?

  1. A.The full range of the values
  2. B.The most frequently occurring value
  3. C.The middle value of the data✓ Correct
  4. D.The average value of the data

Explanation

The mode is the value that appears most frequently in a dataset.

Report an error in this question

Exploratory Data AnalysisEasy

Q10. Which plot is best for showing proportions of a whole?

  1. A.Histogram
  2. B.Pie chart✓ Correct
  3. C.Scatter plot
  4. D.Box plot

Explanation

A pie chart displays proportions of a whole, where each slice represents a category's percentage of the total.

Report an error in this question

Exploratory Data AnalysisMedium

Q11. What is a heatmap used for in EDA?

  1. A.Showing temperature measurements from a sensor only
  2. B.Visualizing the magnitude of values in a matrix using colors✓ Correct
  3. C.Building multi-layer neural network architectures
  4. D.Creating three-dimensional surface projection plots

Explanation

A heatmap uses color intensity to represent the magnitude of values in a matrix, commonly used for correlation matrices.

Report an error in this question

Exploratory Data AnalysisMedium

Q12. Skewness in a distribution refers to:

  1. A.Asymmetry of the distribution around the mean
  2. B.The total number of distribution peaks
  3. C.The height of the data distribution✓ Correct
  4. D.The overall width of the distribution

Explanation

Skewness measures the asymmetry of a probability distribution. Positive skew has a long right tail, negative skew has a long left tail.

Report an error in this question

Exploratory Data AnalysisMedium

Q13. What is a pair plot (pairplot) in Seaborn?

  1. A.A single scatter plot between exactly two variables
  2. B.A grid of plots showing pairwise relationships between variables✓ Correct
  3. C.A grouped bar chart comparing category counts
  4. D.A circular pie chart showing value proportions

Explanation

A pair plot creates a grid of scatter plots for each pair of variables, with histograms or KDE plots on the diagonal.

Report an error in this question

Exploratory Data AnalysisMedium

Q14. Kurtosis measures:

  1. A.The arithmetic mean
  2. B.The tailedness of a distribution
  3. C.The center of the distribution✓ Correct
  4. D.The measure of skewness

Explanation

Kurtosis measures how heavy or light the tails of a distribution are compared to a normal distribution.

Report an error in this question

Exploratory Data AnalysisMedium

Q15. What is the Interquartile Range (IQR)?

  1. A.Mean minus Median
  2. B.Max - Min range
  3. C.Standard deviation * 2✓ Correct
  4. D.Q3 - Q1 difference

Explanation

IQR = Q3 - Q1, representing the middle 50% of the data. It is used to identify outliers (values below Q1-1.5*IQR or above Q3+1.5*IQR).

Report an error in this question

Exploratory Data AnalysisMedium

Q16. A violin plot combines which two visualizations?

  1. A.Box plot and KDE plot✓ Correct
  2. B.Pie chart and histogram
  3. C.Scatter plot and bar chart
  4. D.Line chart and area chart

Explanation

A violin plot combines a box plot with a kernel density estimation (KDE) plot, showing both the summary statistics and distribution shape.

Report an error in this question

Exploratory Data AnalysisMedium

Q17. What does the value_counts() method do in pandas?

  1. A.Returns the frequency of unique values in a Series
  2. B.Sums together all values stored in a Series✓ Correct
  3. C.Counts the total number of rows in a Series
  4. D.Counts the number of null values in a Series

Explanation

value_counts() returns a Series with the count of each unique value, sorted in descending order of frequency.

Report an error in this question

Exploratory Data AnalysisMedium

Q18. A QQ plot is used to:

  1. A.Build linear regression prediction models
  2. B.Plot quarterly financial data over time
  3. C.Create grouped and stacked bar charts
  4. D.Check if data follows a particular distribution✓ Correct

Explanation

A QQ (Quantile-Quantile) plot compares the quantiles of the data to the quantiles of a theoretical distribution to assess distributional fit.

Report an error in this question

Exploratory Data AnalysisMedium

Q19. What is a kernel density estimation (KDE) plot?

  1. A.A circular chart showing proportions
  2. B.A variant of the standard scatter plot
  3. C.A smoothed continuous version of a histogram✓ Correct
  4. D.A type of stacked or grouped bar chart

Explanation

KDE creates a smooth, continuous curve estimating the probability density function of a variable, unlike the discrete bins of a histogram.

Report an error in this question

Exploratory Data AnalysisMedium

Q20. What is the purpose of a log transformation in EDA?

  1. A.To reduce skewness and make data more normally distributed✓ Correct
  2. B.To change the data types of columns in the table
  3. C.To encrypt data columns for compliance and security
  4. D.To completely delete all outlier values from data

Explanation

Log transformation compresses large values and expands small ones, reducing right skewness and making data more suitable for analysis.

Report an error in this question

Exploratory Data AnalysisHard

Q21. Simpson's paradox in EDA refers to:

  1. A.A trend that appears in groups but reverses when groups are combined✓ Correct
  2. B.A visualization error caused by incorrect axis scale calibration
  3. C.A specific type of missing data pattern observed in large datasets
  4. D.A random sampling method for selecting representative subsets

Explanation

Simpson's paradox occurs when a trend in separate groups of data reverses when the groups are combined, highlighting the importance of subgroup analysis.

Report an error in this question

Exploratory Data AnalysisHard

Q22. What is the purpose of the Kolmogorov-Smirnov test?

  1. A.To calculate the arithmetic mean of a data sample
  2. B.To test if a sample comes from a specific distribution✓ Correct
  3. C.To encode categorical data into numerical integers
  4. D.To remove extreme outlier values from the dataset

Explanation

The KS test compares the empirical distribution of a sample to a reference distribution (or two samples) to test distributional assumptions.

Report an error in this question

Exploratory Data AnalysisHard

Q23. What is the curse of dimensionality in EDA?

  1. A.There are too few features for the model to learn
  2. B.There is simply too much data volume to process
  3. C.Data becomes increasingly sparse as dimensions increase✓ Correct
  4. D.The data is already too clean to analyze further

Explanation

The curse of dimensionality means that as the number of features grows, data becomes sparse and distances between points become less meaningful.

Report an error in this question

Exploratory Data AnalysisHard

Q24. Cramér's V statistic measures:

  1. A.The skewness of a single distribution
  2. B.Correlation between two numerical variables✓ Correct
  3. C.The arithmetic mean of a given dataset
  4. D.Association between two categorical variables

Explanation

Cramér's V measures the strength of association between two categorical variables, ranging from 0 (no association) to 1 (complete association).

Report an error in this question

Exploratory Data AnalysisHard

Q25. What is the difference between correlation and causation in EDA?

  1. A.Correlation measures association; causation means one variable directly affects another
  2. B.Causation is a weaker relationship than correlation✓ Correct
  3. C.Correlation always directly implies a causal link
  4. D.They are completely identical concepts in statistics

Explanation

Correlation indicates a statistical association between variables, but does not prove that one causes the other. Causation requires controlled experiments or causal inference methods.

Report an error in this question

Exploratory Data AnalysisHard

Q26. What is the Shapiro-Wilk test used for?

  1. A.Testing for the presence of outlier values
  2. B.Testing if a dataset is normally distributed✓ Correct
  3. C.Testing for the existence of duplicate rows
  4. D.Testing for the count of missing null values

Explanation

The Shapiro-Wilk test is a statistical test of the null hypothesis that a sample comes from a normally distributed population.

Report an error in this question

Exploratory Data AnalysisHard

Q27. In EDA, what is a bimodal distribution?

  1. A.A completely uniform distribution
  2. B.A distribution with two distinct peaks✓ Correct
  3. C.A distribution with no peaks at all
  4. D.A distribution with one single peak

Explanation

A bimodal distribution has two distinct peaks (modes), potentially indicating that the data comes from two different underlying populations.

Report an error in this question

Exploratory Data AnalysisHard

Q28. What does the Durbin-Watson statistic test for?

  1. A.Normality of the data
  2. B.Homoscedasticity✓ Correct
  3. C.Multicollinearity issues
  4. D.Autocorrelation in residuals

Explanation

The Durbin-Watson statistic tests for autocorrelation in the residuals of a regression analysis, with values near 2 indicating no autocorrelation.

Report an error in this question

Exploratory Data AnalysisHard

Q29. What is heteroscedasticity and why is it problematic?

  1. A.Non-constant variance of residuals, violating regression assumptions
  2. B.A specific type of variable correlation
  3. C.A pattern of missing data values✓ Correct
  4. D.Perfectly constant variance of all residuals

Explanation

Heteroscedasticity means the variance of residuals changes across the range of predictions, violating OLS assumptions and making standard errors unreliable.

Report an error in this question

Exploratory Data AnalysisHard

Q30. What is the purpose of the Andrews curves visualization?

  1. A.To display proportional breakdowns in a circular pie format
  2. B.To visualize multivariate data as curves, each representing one observation✓ Correct
  3. C.To create standard grouped and stacked bar charts for categories
  4. D.To plot time series trend lines with seasonal decomposition

Explanation

Andrews curves map each observation to a Fourier-like function curve, allowing visual identification of clusters and outliers in high-dimensional data.

Report an error in this question

Ready to test yourself on Exploratory Data Analysis?

Take a timed quiz drawn from 210+ questions on this topic. No signup required — your progress saves in your browser.

Start Exploratory Data Analysis Quiz