HomeSubjectsUniversityBlogAbout

Deep Learning Fundamentals

Topic in AI / Machine Learning & Data Analytics

210 total MCQsShowing 30 with explanations10 Easy10 Medium10 Hard

About This Topic

Deep learning uses artificial neural networks with many layers of connected neurons to learn increasingly abstract features directly from raw data. Core questions explain perceptrons, weights, biases, forward propagation, and backpropagation, which applies the chain rule to compute gradients for weight updates. Activation functions are heavily tested, including sigmoid, tanh with its output range of -1 to 1, ReLU and Leaky ReLU, which fixes the dying ReLU problem. Expect items on loss functions, optimizers such as SGD, momentum and Adam, vanishing and exploding gradients, Xavier and He initialization, dropout, batch normalization, and mixed-precision training on modern hardware.

Below are 30 practice questions from a pool of 210 Deep Learning Fundamentals MCQs, one of 17 topics in AI / Machine Learning & Data Analytics. Each shows the correct answer with an explanation; when you are ready, take a timed quiz to test recall under exam conditions.

Practice Questions

Each question below shows the correct answer with a full explanation. Use these to build conceptual understanding before attempting a timed quiz.

Deep Learning FundamentalsEasy

Q1. Which framework is commonly used for deep learning?

  1. A.NumPy only
  2. B.Flask
  3. C.TensorFlow✓ Correct
  4. D.Django

Explanation

TensorFlow is one of the most popular deep learning frameworks, providing tools for building and training neural networks.

Report an error in this question

Deep Learning FundamentalsEasy

Q2. What is a neural network?

  1. A.A network routing protocol for data packets
  2. B.A computing system inspired by biological neural networks✓ Correct
  3. C.A comparison-based sorting algorithm for arrays
  4. D.A relational database for storing structured records

Explanation

A neural network is a computational model inspired by the structure of biological neural networks, consisting of interconnected nodes (neurons).

Report an error in this question

Deep Learning FundamentalsEasy

Q3. What is the ReLU activation function?

  1. A.f(x) = x^2
  2. B.f(x) = 1/(1+e^-x)
  3. C.f(x) = max(0, x)✓ Correct
  4. D.f(x) = tanh(x)

Explanation

ReLU (Rectified Linear Unit) outputs the input if positive, otherwise zero: f(x) = max(0, x). It is the most commonly used activation function.

Report an error in this question

Deep Learning FundamentalsEasy

Q4. What is an activation function?

  1. A.A technique for initializing random weight parameters
  2. B.A function that introduces non-linearity into the network✓ Correct
  3. C.A function for loading data from external file sources
  4. D.A method for calculating the total loss of the model

Explanation

Activation functions introduce non-linearity, enabling neural networks to learn complex patterns. Examples include ReLU, sigmoid, and tanh.

Report an error in this question

Deep Learning FundamentalsEasy

Q5. What is a hidden layer?

  1. A.A data preprocessing transformation layer
  2. B.The very last output layer of the network
  3. C.The very first input layer of the network
  4. D.A layer between the input and output layers✓ Correct

Explanation

Hidden layers are intermediate layers between input and output that perform computations and learn feature representations.

Report an error in this question

Deep Learning FundamentalsEasy

Q6. What is backpropagation?

  1. A.A type of non-linear activation function for neurons
  2. B.A specific deep neural network architecture design
  3. C.A data preprocessing technique for cleaning raw inputs
  4. D.An algorithm for computing gradients to update network weights✓ Correct

Explanation

Backpropagation computes the gradient of the loss with respect to each weight by propagating errors backward through the network.

Report an error in this question

Deep Learning FundamentalsEasy

Q7. What is a neuron (node) in a neural network?

  1. A.A basic computational unit that receives inputs and produces an output✓ Correct
  2. B.A physical data storage unit for persisting files on disk
  3. C.A specific type of input feature used in the dataset
  4. D.A gradient-based training algorithm for optimization

Explanation

A neuron receives weighted inputs, applies an activation function, and produces an output that can be passed to other neurons.

Report an error in this question

Deep Learning FundamentalsEasy

Q8. What is gradient descent?

  1. A.A type of multi-layer feedforward neural network architecture with connections
  2. B.An optimization algorithm minimizing loss by updating weights via steepest descent✓ Correct
  3. C.A tree-based hierarchical data structure for nearest-neighbor lookups
  4. D.A non-linear activation function that maps inputs to the range zero to one

Explanation

Gradient descent iteratively adjusts model parameters in the direction that reduces the loss function most rapidly.

Report an error in this question

Deep Learning FundamentalsMedium

Q9. What is the exploding gradient problem?

  1. A.The overall training process runs far too slowly
  2. B.Gradients become extremely small, causing no weight changes
  3. C.The model severely underfits the training data distribution
  4. D.Gradients become extremely large, causing unstable weight updates✓ Correct

Explanation

Exploding gradients occur when gradient values grow exponentially during backpropagation, causing numerical overflow and unstable training.

Report an error in this question

Deep Learning FundamentalsMedium

Q10. What is the vanishing gradient problem?

  1. A.Gradients become very small in early layers, making training slow✓ Correct
  2. B.Gradients become extremely large causing numerical overflow
  3. C.The model has far too many hidden layers to train
  4. D.The training process converges far too quickly to minimum

Explanation

The vanishing gradient problem occurs when gradients diminish exponentially as they propagate back through many layers, hindering learning in early layers.

Report an error in this question

Deep Learning FundamentalsMedium

Q11. What is the Adam optimizer?

  1. A.A non-linear activation function like sigmoid or ReLU
  2. B.An adaptive learning rate optimizer combining momentum and RMSProp✓ Correct
  3. C.A cross-entropy loss function for classification training
  4. D.A type of multi-layer feedforward neural network architecture

Explanation

Adam (Adaptive Moment Estimation) maintains per-parameter learning rates using first and second moment estimates, combining benefits of momentum and RMSProp.

Report an error in this question

Deep Learning FundamentalsHard

Q12. What is the dying ReLU problem?

  1. A.ReLU causes gradients to explode during backpropagation
  2. B.Neurons permanently output zero because they get stuck in the negative region✓ Correct
  3. C.ReLU increases peak memory usage beyond available capacity
  4. D.ReLU activation is computationally too slow for large networks

Explanation

Dying ReLU occurs when neurons consistently receive negative inputs, causing them to always output zero with zero gradients, effectively becoming permanently inactive.

Report an error in this question

Deep Learning FundamentalsMedium

Q13. What is a softmax output layer used for?

  1. A.Feature extraction from the input data
  2. B.Binary classification with only two output classes
  3. C.Regression for continuous value prediction
  4. D.Multi-class classification producing a probability distribution✓ Correct

Explanation

A softmax output layer converts raw scores into probabilities for multi-class classification, where all class probabilities sum to 1.

Report an error in this question

Deep Learning FundamentalsEasy

Q14. What is an epoch in training?

  1. A.A type of non-linear activation function node
  2. B.A single individual data point in the training set
  3. C.A single hidden layer within the neural network
  4. D.One complete pass through the entire training dataset✓ Correct

Explanation

An epoch is one complete iteration through the entire training dataset during the training process.

Report an error in this question

Deep Learning FundamentalsMedium

Q15. What is the difference between a batch and a mini-batch?

  1. A.Batch and mini-batch gradient descent are completely identical✓ Correct
  2. B.A batch is the full dataset; a mini-batch is a subset used per iteration
  3. C.A mini-batch is always larger than a full batch of data
  4. D.A batch always contains exactly one single data sample only

Explanation

A batch uses the entire dataset for one update. A mini-batch uses a smaller subset, balancing computation speed and gradient accuracy.

Report an error in this question

Deep Learning FundamentalsMedium

Q16. What is dropout?

  1. A.Deleting training data samples to reduce dataset size
  2. B.Randomly deactivating neurons during training to prevent overfitting✓ Correct
  3. C.Reducing the learning rate schedule over the epochs
  4. D.Removing network layers permanently from the architecture

Explanation

Dropout randomly sets a fraction of neurons to zero during each training step, preventing co-adaptation and acting as regularization.

Report an error in this question

Deep Learning FundamentalsMedium

Q17. What is the purpose of the learning rate?

  1. A.It determines the total number of hidden layers in the model
  2. B.It explicitly sets the mini-batch size during each epoch
  3. C.It defines the specific neural network architecture layout
  4. D.It controls the step size of weight updates during optimization✓ Correct

Explanation

The learning rate determines how much weights are adjusted with each update. Too high causes instability; too low causes slow convergence.

Report an error in this question

Deep Learning FundamentalsEasy

Q18. What is a loss function?

  1. A.An individual input feature column used for training
  2. B.A non-linear activation function applied at hidden layers
  3. C.A structured collection of labeled data for evaluation
  4. D.A function measuring how well predictions match actual values✓ Correct

Explanation

A loss function quantifies the difference between predicted and actual values, guiding the optimization of model weights.

Report an error in this question

Deep Learning FundamentalsMedium

Q19. What is transfer learning?

  1. A.Moving files between different computer file systems
  2. B.Transferring data records between separate databases
  3. C.Using a pre-trained model on a new but related task✓ Correct
  4. D.Training every model from scratch on random weights

Explanation

Transfer learning leverages knowledge from a model pre-trained on a large dataset, fine-tuning it for a new related task with less data.

Report an error in this question

Deep Learning FundamentalsMedium

Q20. What is batch normalization?

  1. A.A loss function for measuring classification prediction error
  2. B.A preprocessing step applied only to the raw input data
  3. C.Normalizing the inputs of each layer to stabilize and accelerate training✓ Correct
  4. D.A type of non-linear activation function for hidden layers

Explanation

Batch normalization normalizes layer inputs across a mini-batch, reducing internal covariate shift and allowing higher learning rates.

Report an error in this question

Deep Learning FundamentalsMedium

Q21. What is weight initialization and why is it important?

  1. A.Setting initial weights to enable effective training and avoid gradient problems✓ Correct
  2. B.Weights are assigned their final values only after training
  3. C.Weights are always initialized to exactly zero for every neuron
  4. D.The initialization strategy has no effect on the training process

Explanation

Proper weight initialization (e.g., Xavier, He) ensures gradients flow well during training, preventing vanishing or exploding gradients from the start.

Report an error in this question

Deep Learning FundamentalsHard

Q22. What is the purpose of skip connections (residual connections)?

  1. A.Removing entire layers from the network to reduce complexity
  2. B.Skipping the data preprocessing step to speed up training
  3. C.Reducing the size of the training dataset to save storage
  4. D.Allowing gradients to flow directly by adding input to output of a block✓ Correct

Explanation

Skip connections add the input of a block to its output, enabling gradient flow across many layers and allowing networks to learn residual functions, as in ResNets.

Report an error in this question

Deep Learning FundamentalsHard

Q23. What is the difference between SGD with momentum and Adam?

  1. A.SGD momentum uses a fixed learning rate with velocity; Adam adapts per parameter✓ Correct
  2. B.Adam optimizer never converges to a good solution on real data
  3. C.SGD with momentum is always faster than Adam for every problem
  4. D.They are completely identical optimization algorithms with no differences

Explanation

SGD with momentum accumulates velocity but uses a global learning rate. Adam adapts per-parameter learning rates using running averages of gradients and squared gradients.

Report an error in this question

Deep Learning FundamentalsHard

Q24. What is the effect of batch size on training?

  1. A.Smaller batch sizes are always strictly worse for model generalization
  2. B.Larger batch sizes are always strictly better for all training scenarios
  3. C.Larger batches give smoother gradients but may generalize worse; smaller batches add helpful noise✓ Correct
  4. D.Batch size has absolutely no measurable effect on training or generalization

Explanation

Larger batch sizes provide smoother gradient estimates but may converge to sharp minima that generalize poorly. Smaller batches introduce noise that can help escape sharp minima.

Report an error in this question

Deep Learning FundamentalsHard

Q25. What is the Lottery Ticket Hypothesis?

  1. A.All neural network architectures are equally effective regardless of size
  2. B.Sparse network architectures never work well for any real task
  3. C.Dense networks contain sparse subnetworks that can match full network performance✓ Correct
  4. D.Larger networks are always strictly better than smaller alternatives

Explanation

The Lottery Ticket Hypothesis posits that randomly-initialized dense networks contain sparse subnetworks ('winning tickets') that can match the full network's performance.

Report an error in this question

Deep Learning FundamentalsHard

Q26. What is knowledge distillation?

  1. A.Removing irrelevant features from the input feature space
  2. B.Training a smaller student model to mimic a larger teacher model✓ Correct
  3. C.Transferring database records between different storage systems
  4. D.Augmenting the training data with synthetic generated samples

Explanation

Knowledge distillation trains a compact student model to reproduce the behavior of a larger teacher model, compressing knowledge into a smaller, deployable model.

Report an error in this question

Deep Learning FundamentalsHard

Q27. What is the difference between pre-training and fine-tuning?

  1. A.They are completely identical stages in the same training process
  2. B.Fine-tuning always happens before the pre-training phase
  3. C.Pre-training is designed specifically for very small datasets
  4. D.Pre-training learns general representations; fine-tuning adapts to a specific task✓ Correct

Explanation

Pre-training learns general features from large datasets. Fine-tuning adapts those learned representations to a specific downstream task with task-specific data.

Report an error in this question

Deep Learning FundamentalsHard

Q28. What is neural architecture search (NAS)?

  1. A.Automated methods for finding optimal neural network architectures✓ Correct
  2. B.Feature selection and dimensionality reduction algorithms
  3. C.Manual hand-designed network architecture engineering process
  4. D.Data preprocessing and cleaning pipeline automation tools

Explanation

NAS uses optimization techniques to automatically discover optimal network architectures, reducing the need for manual architecture engineering.

Report an error in this question

Deep Learning FundamentalsHard

Q29. What is gradient clipping?

  1. A.Removing all gradients entirely from the backpropagation computation
  2. B.Setting all gradient values to exactly zero after each update step
  3. C.Increasing the learning rate to speed up convergence time
  4. D.Limiting gradient values to a maximum threshold to prevent exploding gradients✓ Correct

Explanation

Gradient clipping caps gradient values when they exceed a threshold, preventing exploding gradients while maintaining training stability.

Report an error in this question

Deep Learning FundamentalsHard

Q30. What is mixed-precision training?

  1. A.Using both 16-bit and 32-bit floating point to speed up training while maintaining accuracy✓ Correct
  2. B.Using only 8-bit integer arithmetic for all network computation steps
  3. C.Using only 64-bit double-precision floats for maximum numerical precision
  4. D.Training neural network models entirely without any GPU acceleration

Explanation

Mixed-precision training uses FP16 for most computations for speed and FP32 for critical accumulations for accuracy, reducing memory usage and increasing throughput.

Report an error in this question

Ready to test yourself on Deep Learning Fundamentals?

Take a timed quiz drawn from 210+ questions on this topic. No signup required — your progress saves in your browser.

Start Deep Learning Fundamentals Quiz