HomeSubjectsUniversityBlogAbout

Mathematical Foundations

Topic in AI / Machine Learning & Data Analytics

210 total MCQsShowing 30 with explanations10 Easy10 Medium10 Hard

About This Topic

The mathematical foundations of machine learning are linear algebra, calculus, probability and statistics, the tools used to describe and optimize models. Questions define scalars, vectors, matrices and tensors, and test dot products, matrix multiplication, transposes, eigenvalues and eigenvectors. Calculus items cover derivatives, partial derivatives that hold other variables constant, the chain rule and gradients used in gradient descent. Probability questions address distributions such as normal, Bernoulli and binomial, expectation, variance, conditional probability and Bayes' rule. Advanced items ask about entropy, cross-entropy, Kullback-Leibler divergence between distributions, maximum likelihood estimation and the Fisher information matrix.

Below are 30 practice questions from a pool of 210 Mathematical Foundations MCQs, one of 17 topics in AI / Machine Learning & Data Analytics. Each shows the correct answer with an explanation; when you are ready, take a timed quiz to test recall under exam conditions.

Practice Questions

Each question below shows the correct answer with a full explanation. Use these to build conceptual understanding before attempting a timed quiz.

Mathematical FoundationsEasy

Q1. What is a scalar in linear algebra?

  1. A.A higher-order tensor
  2. B.A multi-row data matrix
  3. C.A single numerical value✓ Correct
  4. D.A multi-element vector

Explanation

A scalar is a single numerical value, as opposed to vectors, matrices, or higher-order tensors.

Report an error in this question

Mathematical FoundationsEasy

Q2. What is the mean of the numbers 2, 4, 6, 8, 10?

  1. A.7
  2. B.8
  3. C.5
  4. D.6✓ Correct

Explanation

The mean is (2+4+6+8+10)/5 = 30/5 = 6.

Report an error in this question

Mathematical FoundationsEasy

Q3. A matrix with equal rows and columns is called:

  1. A.Rectangular matrix
  2. B.Diagonal matrix
  3. C.Square matrix✓ Correct
  4. D.Identity matrix

Explanation

A square matrix has the same number of rows and columns.

Report an error in this question

Mathematical FoundationsEasy

Q4. What does probability measure?

  1. A.The computational speed of an algorithm
  2. B.The overall size of a given dataset
  3. C.The likelihood of an event occurring
  4. D.The hardware accuracy of a processor✓ Correct

Explanation

Probability measures the likelihood or chance of an event occurring, ranging from 0 to 1.

Report an error in this question

Mathematical FoundationsEasy

Q5. The derivative of a constant is:

  1. A.1
  2. B.The constant itself
  3. C.0✓ Correct
  4. D.Infinity

Explanation

The derivative of a constant is always 0 because a constant has no rate of change.

Report an error in this question

Mathematical FoundationsEasy

Q6. What is the dot product of vectors [1,2] and [3,4]?

  1. A.7
  2. B.14
  3. C.11✓ Correct
  4. D.10

Explanation

Dot product = (1*3) + (2*4) = 3 + 8 = 11.

Report an error in this question

Mathematical FoundationsEasy

Q7. What is the median of {1, 3, 5, 7, 9}?

  1. A.3
  2. B.5✓ Correct
  3. C.4
  4. D.7

Explanation

The median is the middle value in a sorted set. In {1,3,5,7,9}, the middle value is 5.

Report an error in this question

Mathematical FoundationsEasy

Q8. An identity matrix has:

  1. A.All entries set to the value of zero
  2. B.Entries filled with random real numbers
  3. C.1s on the diagonal and 0s elsewhere✓ Correct
  4. D.All entries set to the value of one

Explanation

An identity matrix has 1s on the main diagonal and 0s in all other positions.

Report an error in this question

Mathematical FoundationsEasy

Q9. The range of a probability value is:

  1. A.-1 to 1
  2. B.Negative infinity to positive infinity
  3. C.0 to 1✓ Correct
  4. D.0 to 100

Explanation

Probability values always range from 0 (impossible) to 1 (certain).

Report an error in this question

Mathematical FoundationsEasy

Q10. What is the transpose of a matrix?

  1. A.The matrix is numerically inverted
  2. B.The matrix entries are all squared
  3. C.The matrix is multiplied by itself
  4. D.Rows and columns are interchanged✓ Correct

Explanation

The transpose of a matrix is obtained by interchanging its rows and columns.

Report an error in this question

Mathematical FoundationsMedium

Q11. The gradient of a function points in the direction of:

  1. A.Steepest descent
  2. B.Random direction
  3. C.Steepest ascent✓ Correct
  4. D.Zero change

Explanation

The gradient vector points in the direction of the steepest increase of the function.

Report an error in this question

Mathematical FoundationsMedium

Q12. Bayes' theorem relates:

  1. A.Prior and posterior probabilities✓ Correct
  2. B.Only matrix decompositions
  3. C.Only the mean and variance
  4. D.Only derivative computations

Explanation

Bayes' theorem describes how to update the probability of a hypothesis given new evidence, relating prior and posterior probabilities.

Report an error in this question

Mathematical FoundationsMedium

Q13. The eigenvalue equation is Av = λv. What is λ?

  1. A.Eigenvector
  2. B.Trace
  3. C.Determinant
  4. D.Eigenvalue✓ Correct

Explanation

In the equation Av = λv, λ is the eigenvalue, a scalar that indicates how the eigenvector v is scaled by matrix A.

Report an error in this question

Mathematical FoundationsMedium

Q14. What is the chain rule used for in calculus?

  1. A.Computing joint probabilities
  2. B.Adding matrices element-wise
  3. C.Differentiating composite functions✓ Correct
  4. D.Sorting data sequentially

Explanation

The chain rule is used to compute the derivative of a composite function: d/dx[f(g(x))] = f'(g(x)) * g'(x).

Report an error in this question

Mathematical FoundationsMedium

Q15. Standard deviation measures:

  1. A.Spread of data around the mean✓ Correct
  2. B.The central tendency of values
  3. C.The minimum observed value
  4. D.The maximum observed value

Explanation

Standard deviation measures how spread out data values are from the mean of the dataset.

Report an error in this question

Mathematical FoundationsMedium

Q16. A positive definite matrix has:

  1. A.All negative eigenvalues
  2. B.Mixed sign eigenvalues
  3. C.All zero eigenvalues
  4. D.All positive eigenvalues✓ Correct

Explanation

A positive definite matrix has all positive eigenvalues, which is important in optimization (guarantees a minimum).

Report an error in this question

Mathematical FoundationsMedium

Q17. In probability, two events are independent if:

  1. A.P(A|B) = P(B)
  2. B.P(A∩B) = P(A) * P(B)✓ Correct
  3. C.P(A∩B) = P(A) + P(B)
  4. D.P(A∩B) = 0

Explanation

Two events are independent if the probability of both occurring equals the product of their individual probabilities.

Report an error in this question

Mathematical FoundationsMedium

Q18. The determinant of a 2x2 matrix [[a,b],[c,d]] is:

  1. A.ab + cd
  2. B.ab - cd
  3. C.ad - bc✓ Correct
  4. D.ad + bc

Explanation

The determinant of a 2x2 matrix is calculated as ad - bc.

Report an error in this question

Mathematical FoundationsMedium

Q19. A convex function has the property that:

  1. A.It is always an increasing function
  2. B.It has no minimum point at all
  3. C.Any local minimum is also a global minimum✓ Correct
  4. D.It has many distinct local minima

Explanation

A key property of convex functions is that any local minimum is also a global minimum, which simplifies optimization.

Report an error in this question

Mathematical FoundationsMedium

Q20. The covariance between two identical variables equals:

  1. A.A value of exactly zero
  2. B.The variance of that variable✓ Correct
  3. C.A value of negative one
  4. D.A value of exactly one

Explanation

Cov(X,X) = Var(X). The covariance of a variable with itself is its variance.

Report an error in this question

Mathematical FoundationsHard

Q21. The Hessian matrix contains:

  1. A.Eigenvalues of the matrix only
  2. B.The determinant values only
  3. C.Second-order partial derivatives✓ Correct
  4. D.First-order derivatives only

Explanation

The Hessian matrix contains all second-order partial derivatives of a scalar-valued function, used to analyze curvature.

Report an error in this question

Mathematical FoundationsHard

Q22. The Kullback-Leibler divergence measures:

  1. A.The spread or variance of a single distribution
  2. B.The central tendency or mean of a distribution
  3. C.How one probability distribution differs from another✓ Correct
  4. D.The linear correlation between two variables

Explanation

KL divergence measures the difference between two probability distributions, commonly used in ML for comparing distributions.

Report an error in this question

Mathematical FoundationsHard

Q23. A Jacobian matrix represents:

  1. A.First-order partial derivatives of a vector-valued function✓ Correct
  2. B.Only the output values of a single scalar function
  3. C.Only the second-order derivatives of a scalar function
  4. D.Only the eigenvalues of a square transformation matrix

Explanation

The Jacobian matrix contains all first-order partial derivatives of a vector-valued function with respect to another vector.

Report an error in this question

Mathematical FoundationsHard

Q24. In the context of ML, why is the log-likelihood used instead of likelihood?

  1. A.It requires significantly less memory during training
  2. B.It converts products to sums, making computation easier✓ Correct
  3. C.It produces more statistically accurate final estimates
  4. D.It runs faster on modern parallel computing hardware

Explanation

Log-likelihood converts the product of probabilities into a sum, making numerical computation more stable and differentiation easier.

Report an error in this question

Mathematical FoundationsHard

Q25. The singular value decomposition (SVD) decomposes a matrix into:

  1. A.U, Sigma, and V-transpose matrices✓ Correct
  2. B.Only the eigenvector components
  3. C.Exactly two factor matrices
  4. D.Only the eigenvalue components

Explanation

SVD decomposes a matrix A into UΣV^T where U and V are orthogonal matrices and Σ is a diagonal matrix of singular values.

Report an error in this question

Mathematical FoundationsHard

Q26. A saddle point in optimization is where:

  1. A.The objective function reaches its overall global minimum
  2. B.The gradient of the function is completely undefined
  3. C.The objective function reaches its overall global maximum
  4. D.The gradient is zero but it is neither a minimum nor maximum✓ Correct

Explanation

At a saddle point, the gradient is zero but the point is a minimum along some directions and a maximum along others.

Report an error in this question

Mathematical FoundationsHard

Q27. The moment generating function uniquely determines:

  1. A.Only the median statistic
  2. B.Only the expected mean value
  3. C.The probability distribution✓ Correct
  4. D.Only the variance estimate

Explanation

If it exists in a neighborhood of zero, the moment generating function uniquely determines the probability distribution.

Report an error in this question

Mathematical FoundationsHard

Q28. In matrix calculus, the gradient of x^T A x with respect to x is:

  1. A.2x
  2. B.Ax
  3. C.A^T
  4. D.(A + A^T)x✓ Correct

Explanation

The gradient of x^T A x with respect to x is (A + A^T)x. If A is symmetric, this simplifies to 2Ax.

Report an error in this question

Mathematical FoundationsHard

Q29. The condition number of a matrix indicates:

  1. A.The trace or diagonal sum of the matrix
  2. B.Sensitivity of the solution to small changes in input✓ Correct
  3. C.The total number of columns in the matrix
  4. D.The total number of rows in the matrix

Explanation

The condition number measures how sensitive a matrix computation is to perturbations in the input, important for numerical stability.

Report an error in this question

Mathematical FoundationsHard

Q30. Jensen's inequality states that for a convex function f:

  1. A.f(E[X]) ≤ E[f(X)]✓ Correct
  2. B.f(E[X]) ≥ E[f(X)]
  3. C.f(E[X]) = 0
  4. D.f(E[X]) = E[f(X)]

Explanation

Jensen's inequality states that for a convex function, the function of the expected value is less than or equal to the expected value of the function.

Report an error in this question

Ready to test yourself on Mathematical Foundations?

Take a timed quiz drawn from 210+ questions on this topic. No signup required — your progress saves in your browser.

Start Mathematical Foundations Quiz