Each question below shows the correct answer with a full explanation. Use these to build conceptual understanding before attempting a timed quiz.
Exploratory Data AnalysisEasy
Q1. What does EDA stand for?
- A.Extended Data Algorithm
- B.External Data Analysis
- C.Effective Data Assessment
- D.Exploratory Data Analysis✓ Correct
Explanation
EDA stands for Exploratory Data Analysis, the process of analyzing data sets to summarize their main characteristics.
Report an error in this question
Exploratory Data AnalysisEasy
Q2. Which chart is used to show the distribution of a single variable?
- A.Line chart
- B.Scatter plot
- C.Histogram✓ Correct
- D.Pie chart
Explanation
A histogram displays the frequency distribution of a single continuous variable by dividing data into bins.
Report an error in this question
Exploratory Data AnalysisEasy
Q3. What does a box plot display?
- A.The five-number summary: min, Q1, median, Q3, max✓ Correct
- B.Only the arithmetic mean of the distribution values
- C.Only the overall range from minimum to maximum
- D.Only the mode or most frequent value observed
Explanation
A box plot visualizes the five-number summary (minimum, Q1, median, Q3, maximum) and helps identify outliers.
Report an error in this question
Exploratory Data AnalysisEasy
Q4. Which pandas method gives summary statistics?
- A.describe()✓ Correct
- B.info()
- C.head()
- D.shape
Explanation
df.describe() provides summary statistics including count, mean, std, min, quartiles, and max for numerical columns.
Report an error in this question
Exploratory Data AnalysisEasy
Q5. A scatter plot shows the relationship between:
- A.Two numerical variables
- B.Categories data only
- C.Time series data only
- D.Textual data only✓ Correct
Explanation
A scatter plot displays the relationship between two numerical variables, with each point representing one observation.
Report an error in this question
Exploratory Data AnalysisEasy
Q6. What is the purpose of a correlation matrix?
- A.To clean and preprocess the raw input data
- B.To deploy and manage software applications
- C.To train and evaluate machine learning models
- D.To show relationships between multiple variables✓ Correct
Explanation
A correlation matrix shows the pairwise correlation coefficients between multiple variables, indicating linear relationships.
Report an error in this question
Exploratory Data AnalysisEasy
Q7. What does df.info() display in pandas?
- A.Only the numerical summary statistics values
- B.Only the count of missing or null values
- C.Column names, data types, and non-null counts✓ Correct
- D.Only the first five rows of the DataFrame
Explanation
df.info() displays column names, data types, non-null counts, and memory usage of a DataFrame.
Report an error in this question
Exploratory Data AnalysisEasy
Q8. A bar chart is best used for:
- A.Showing data distributions
- B.Comparing quantities across categories
- C.Displaying feature correlations
- D.Plotting time series trends✓ Correct
Explanation
Bar charts are ideal for comparing quantities across different categories using rectangular bars.
Report an error in this question
Exploratory Data AnalysisEasy
Q9. What is the mode of a dataset?
- A.The full range of the values
- B.The most frequently occurring value
- C.The middle value of the data✓ Correct
- D.The average value of the data
Explanation
The mode is the value that appears most frequently in a dataset.
Report an error in this question
Exploratory Data AnalysisEasy
Q10. Which plot is best for showing proportions of a whole?
- A.Histogram
- B.Pie chart✓ Correct
- C.Scatter plot
- D.Box plot
Explanation
A pie chart displays proportions of a whole, where each slice represents a category's percentage of the total.
Report an error in this question
Exploratory Data AnalysisMedium
Q11. What is a heatmap used for in EDA?
- A.Showing temperature measurements from a sensor only
- B.Visualizing the magnitude of values in a matrix using colors✓ Correct
- C.Building multi-layer neural network architectures
- D.Creating three-dimensional surface projection plots
Explanation
A heatmap uses color intensity to represent the magnitude of values in a matrix, commonly used for correlation matrices.
Report an error in this question
Exploratory Data AnalysisMedium
Q12. Skewness in a distribution refers to:
- A.Asymmetry of the distribution around the mean
- B.The total number of distribution peaks
- C.The height of the data distribution✓ Correct
- D.The overall width of the distribution
Explanation
Skewness measures the asymmetry of a probability distribution. Positive skew has a long right tail, negative skew has a long left tail.
Report an error in this question
Exploratory Data AnalysisMedium
Q13. What is a pair plot (pairplot) in Seaborn?
- A.A single scatter plot between exactly two variables
- B.A grid of plots showing pairwise relationships between variables✓ Correct
- C.A grouped bar chart comparing category counts
- D.A circular pie chart showing value proportions
Explanation
A pair plot creates a grid of scatter plots for each pair of variables, with histograms or KDE plots on the diagonal.
Report an error in this question
Exploratory Data AnalysisMedium
Q14. Kurtosis measures:
- A.The arithmetic mean
- B.The tailedness of a distribution
- C.The center of the distribution✓ Correct
- D.The measure of skewness
Explanation
Kurtosis measures how heavy or light the tails of a distribution are compared to a normal distribution.
Report an error in this question
Exploratory Data AnalysisMedium
Q15. What is the Interquartile Range (IQR)?
- A.Mean minus Median
- B.Max - Min range
- C.Standard deviation * 2✓ Correct
- D.Q3 - Q1 difference
Explanation
IQR = Q3 - Q1, representing the middle 50% of the data. It is used to identify outliers (values below Q1-1.5*IQR or above Q3+1.5*IQR).
Report an error in this question
Exploratory Data AnalysisMedium
Q16. A violin plot combines which two visualizations?
- A.Box plot and KDE plot✓ Correct
- B.Pie chart and histogram
- C.Scatter plot and bar chart
- D.Line chart and area chart
Explanation
A violin plot combines a box plot with a kernel density estimation (KDE) plot, showing both the summary statistics and distribution shape.
Report an error in this question
Exploratory Data AnalysisMedium
Q17. What does the value_counts() method do in pandas?
- A.Returns the frequency of unique values in a Series
- B.Sums together all values stored in a Series✓ Correct
- C.Counts the total number of rows in a Series
- D.Counts the number of null values in a Series
Explanation
value_counts() returns a Series with the count of each unique value, sorted in descending order of frequency.
Report an error in this question
Exploratory Data AnalysisMedium
Q18. A QQ plot is used to:
- A.Build linear regression prediction models
- B.Plot quarterly financial data over time
- C.Create grouped and stacked bar charts
- D.Check if data follows a particular distribution✓ Correct
Explanation
A QQ (Quantile-Quantile) plot compares the quantiles of the data to the quantiles of a theoretical distribution to assess distributional fit.
Report an error in this question
Exploratory Data AnalysisMedium
Q19. What is a kernel density estimation (KDE) plot?
- A.A circular chart showing proportions
- B.A variant of the standard scatter plot
- C.A smoothed continuous version of a histogram✓ Correct
- D.A type of stacked or grouped bar chart
Explanation
KDE creates a smooth, continuous curve estimating the probability density function of a variable, unlike the discrete bins of a histogram.
Report an error in this question
Exploratory Data AnalysisMedium
Q20. What is the purpose of a log transformation in EDA?
- A.To reduce skewness and make data more normally distributed✓ Correct
- B.To change the data types of columns in the table
- C.To encrypt data columns for compliance and security
- D.To completely delete all outlier values from data
Explanation
Log transformation compresses large values and expands small ones, reducing right skewness and making data more suitable for analysis.
Report an error in this question
Exploratory Data AnalysisHard
Q21. Simpson's paradox in EDA refers to:
- A.A trend that appears in groups but reverses when groups are combined✓ Correct
- B.A visualization error caused by incorrect axis scale calibration
- C.A specific type of missing data pattern observed in large datasets
- D.A random sampling method for selecting representative subsets
Explanation
Simpson's paradox occurs when a trend in separate groups of data reverses when the groups are combined, highlighting the importance of subgroup analysis.
Report an error in this question
Exploratory Data AnalysisHard
Q22. What is the purpose of the Kolmogorov-Smirnov test?
- A.To calculate the arithmetic mean of a data sample
- B.To test if a sample comes from a specific distribution✓ Correct
- C.To encode categorical data into numerical integers
- D.To remove extreme outlier values from the dataset
Explanation
The KS test compares the empirical distribution of a sample to a reference distribution (or two samples) to test distributional assumptions.
Report an error in this question
Exploratory Data AnalysisHard
Q23. What is the curse of dimensionality in EDA?
- A.There are too few features for the model to learn
- B.There is simply too much data volume to process
- C.Data becomes increasingly sparse as dimensions increase✓ Correct
- D.The data is already too clean to analyze further
Explanation
The curse of dimensionality means that as the number of features grows, data becomes sparse and distances between points become less meaningful.
Report an error in this question
Exploratory Data AnalysisHard
Q24. Cramér's V statistic measures:
- A.The skewness of a single distribution
- B.Correlation between two numerical variables✓ Correct
- C.The arithmetic mean of a given dataset
- D.Association between two categorical variables
Explanation
Cramér's V measures the strength of association between two categorical variables, ranging from 0 (no association) to 1 (complete association).
Report an error in this question
Exploratory Data AnalysisHard
Q25. What is the difference between correlation and causation in EDA?
- A.Correlation measures association; causation means one variable directly affects another
- B.Causation is a weaker relationship than correlation✓ Correct
- C.Correlation always directly implies a causal link
- D.They are completely identical concepts in statistics
Explanation
Correlation indicates a statistical association between variables, but does not prove that one causes the other. Causation requires controlled experiments or causal inference methods.
Report an error in this question
Exploratory Data AnalysisHard
Q26. What is the Shapiro-Wilk test used for?
- A.Testing for the presence of outlier values
- B.Testing if a dataset is normally distributed✓ Correct
- C.Testing for the existence of duplicate rows
- D.Testing for the count of missing null values
Explanation
The Shapiro-Wilk test is a statistical test of the null hypothesis that a sample comes from a normally distributed population.
Report an error in this question
Exploratory Data AnalysisHard
Q27. In EDA, what is a bimodal distribution?
- A.A completely uniform distribution
- B.A distribution with two distinct peaks✓ Correct
- C.A distribution with no peaks at all
- D.A distribution with one single peak
Explanation
A bimodal distribution has two distinct peaks (modes), potentially indicating that the data comes from two different underlying populations.
Report an error in this question
Exploratory Data AnalysisHard
Q28. What does the Durbin-Watson statistic test for?
- A.Normality of the data
- B.Homoscedasticity✓ Correct
- C.Multicollinearity issues
- D.Autocorrelation in residuals
Explanation
The Durbin-Watson statistic tests for autocorrelation in the residuals of a regression analysis, with values near 2 indicating no autocorrelation.
Report an error in this question
Exploratory Data AnalysisHard
Q29. What is heteroscedasticity and why is it problematic?
- A.Non-constant variance of residuals, violating regression assumptions
- B.A specific type of variable correlation
- C.A pattern of missing data values✓ Correct
- D.Perfectly constant variance of all residuals
Explanation
Heteroscedasticity means the variance of residuals changes across the range of predictions, violating OLS assumptions and making standard errors unreliable.
Report an error in this question
Exploratory Data AnalysisHard
Q30. What is the purpose of the Andrews curves visualization?
- A.To display proportional breakdowns in a circular pie format
- B.To visualize multivariate data as curves, each representing one observation✓ Correct
- C.To create standard grouped and stacked bar charts for categories
- D.To plot time series trend lines with seasonal decomposition
Explanation
Andrews curves map each observation to a Fourier-like function curve, allowing visual identification of clusters and outliers in high-dimensional data.
Report an error in this question