Each question below shows the correct answer with a full explanation. Use these to build conceptual understanding before attempting a timed quiz.
Deep Learning FundamentalsEasy
Q1. Which framework is commonly used for deep learning?
- A.NumPy only
- B.Flask
- C.TensorFlow✓ Correct
- D.Django
Explanation
TensorFlow is one of the most popular deep learning frameworks, providing tools for building and training neural networks.
Report an error in this question
Deep Learning FundamentalsEasy
Q2. What is a neural network?
- A.A network routing protocol for data packets
- B.A computing system inspired by biological neural networks✓ Correct
- C.A comparison-based sorting algorithm for arrays
- D.A relational database for storing structured records
Explanation
A neural network is a computational model inspired by the structure of biological neural networks, consisting of interconnected nodes (neurons).
Report an error in this question
Deep Learning FundamentalsEasy
Q3. What is the ReLU activation function?
- A.f(x) = x^2
- B.f(x) = 1/(1+e^-x)
- C.f(x) = max(0, x)✓ Correct
- D.f(x) = tanh(x)
Explanation
ReLU (Rectified Linear Unit) outputs the input if positive, otherwise zero: f(x) = max(0, x). It is the most commonly used activation function.
Report an error in this question
Deep Learning FundamentalsEasy
Q4. What is an activation function?
- A.A technique for initializing random weight parameters
- B.A function that introduces non-linearity into the network✓ Correct
- C.A function for loading data from external file sources
- D.A method for calculating the total loss of the model
Explanation
Activation functions introduce non-linearity, enabling neural networks to learn complex patterns. Examples include ReLU, sigmoid, and tanh.
Report an error in this question
Deep Learning FundamentalsEasy
Q5. What is a hidden layer?
- A.A data preprocessing transformation layer
- B.The very last output layer of the network
- C.The very first input layer of the network
- D.A layer between the input and output layers✓ Correct
Explanation
Hidden layers are intermediate layers between input and output that perform computations and learn feature representations.
Report an error in this question
Deep Learning FundamentalsEasy
Q6. What is backpropagation?
- A.A type of non-linear activation function for neurons
- B.A specific deep neural network architecture design
- C.A data preprocessing technique for cleaning raw inputs
- D.An algorithm for computing gradients to update network weights✓ Correct
Explanation
Backpropagation computes the gradient of the loss with respect to each weight by propagating errors backward through the network.
Report an error in this question
Deep Learning FundamentalsEasy
Q7. What is a neuron (node) in a neural network?
- A.A basic computational unit that receives inputs and produces an output✓ Correct
- B.A physical data storage unit for persisting files on disk
- C.A specific type of input feature used in the dataset
- D.A gradient-based training algorithm for optimization
Explanation
A neuron receives weighted inputs, applies an activation function, and produces an output that can be passed to other neurons.
Report an error in this question
Deep Learning FundamentalsEasy
Q8. What is gradient descent?
- A.A type of multi-layer feedforward neural network architecture with connections
- B.An optimization algorithm minimizing loss by updating weights via steepest descent✓ Correct
- C.A tree-based hierarchical data structure for nearest-neighbor lookups
- D.A non-linear activation function that maps inputs to the range zero to one
Explanation
Gradient descent iteratively adjusts model parameters in the direction that reduces the loss function most rapidly.
Report an error in this question
Deep Learning FundamentalsMedium
Q9. What is the exploding gradient problem?
- A.The overall training process runs far too slowly
- B.Gradients become extremely small, causing no weight changes
- C.The model severely underfits the training data distribution
- D.Gradients become extremely large, causing unstable weight updates✓ Correct
Explanation
Exploding gradients occur when gradient values grow exponentially during backpropagation, causing numerical overflow and unstable training.
Report an error in this question
Deep Learning FundamentalsMedium
Q10. What is the vanishing gradient problem?
- A.Gradients become very small in early layers, making training slow✓ Correct
- B.Gradients become extremely large causing numerical overflow
- C.The model has far too many hidden layers to train
- D.The training process converges far too quickly to minimum
Explanation
The vanishing gradient problem occurs when gradients diminish exponentially as they propagate back through many layers, hindering learning in early layers.
Report an error in this question
Deep Learning FundamentalsMedium
Q11. What is the Adam optimizer?
- A.A non-linear activation function like sigmoid or ReLU
- B.An adaptive learning rate optimizer combining momentum and RMSProp✓ Correct
- C.A cross-entropy loss function for classification training
- D.A type of multi-layer feedforward neural network architecture
Explanation
Adam (Adaptive Moment Estimation) maintains per-parameter learning rates using first and second moment estimates, combining benefits of momentum and RMSProp.
Report an error in this question
Deep Learning FundamentalsHard
Q12. What is the dying ReLU problem?
- A.ReLU causes gradients to explode during backpropagation
- B.Neurons permanently output zero because they get stuck in the negative region✓ Correct
- C.ReLU increases peak memory usage beyond available capacity
- D.ReLU activation is computationally too slow for large networks
Explanation
Dying ReLU occurs when neurons consistently receive negative inputs, causing them to always output zero with zero gradients, effectively becoming permanently inactive.
Report an error in this question
Deep Learning FundamentalsMedium
Q13. What is a softmax output layer used for?
- A.Feature extraction from the input data
- B.Binary classification with only two output classes
- C.Regression for continuous value prediction
- D.Multi-class classification producing a probability distribution✓ Correct
Explanation
A softmax output layer converts raw scores into probabilities for multi-class classification, where all class probabilities sum to 1.
Report an error in this question
Deep Learning FundamentalsEasy
Q14. What is an epoch in training?
- A.A type of non-linear activation function node
- B.A single individual data point in the training set
- C.A single hidden layer within the neural network
- D.One complete pass through the entire training dataset✓ Correct
Explanation
An epoch is one complete iteration through the entire training dataset during the training process.
Report an error in this question
Deep Learning FundamentalsMedium
Q15. What is the difference between a batch and a mini-batch?
- A.Batch and mini-batch gradient descent are completely identical✓ Correct
- B.A batch is the full dataset; a mini-batch is a subset used per iteration
- C.A mini-batch is always larger than a full batch of data
- D.A batch always contains exactly one single data sample only
Explanation
A batch uses the entire dataset for one update. A mini-batch uses a smaller subset, balancing computation speed and gradient accuracy.
Report an error in this question
Deep Learning FundamentalsMedium
Q16. What is dropout?
- A.Deleting training data samples to reduce dataset size
- B.Randomly deactivating neurons during training to prevent overfitting✓ Correct
- C.Reducing the learning rate schedule over the epochs
- D.Removing network layers permanently from the architecture
Explanation
Dropout randomly sets a fraction of neurons to zero during each training step, preventing co-adaptation and acting as regularization.
Report an error in this question
Deep Learning FundamentalsMedium
Q17. What is the purpose of the learning rate?
- A.It determines the total number of hidden layers in the model
- B.It explicitly sets the mini-batch size during each epoch
- C.It defines the specific neural network architecture layout
- D.It controls the step size of weight updates during optimization✓ Correct
Explanation
The learning rate determines how much weights are adjusted with each update. Too high causes instability; too low causes slow convergence.
Report an error in this question
Deep Learning FundamentalsEasy
Q18. What is a loss function?
- A.An individual input feature column used for training
- B.A non-linear activation function applied at hidden layers
- C.A structured collection of labeled data for evaluation
- D.A function measuring how well predictions match actual values✓ Correct
Explanation
A loss function quantifies the difference between predicted and actual values, guiding the optimization of model weights.
Report an error in this question
Deep Learning FundamentalsMedium
Q19. What is transfer learning?
- A.Moving files between different computer file systems
- B.Transferring data records between separate databases
- C.Using a pre-trained model on a new but related task✓ Correct
- D.Training every model from scratch on random weights
Explanation
Transfer learning leverages knowledge from a model pre-trained on a large dataset, fine-tuning it for a new related task with less data.
Report an error in this question
Deep Learning FundamentalsMedium
Q20. What is batch normalization?
- A.A loss function for measuring classification prediction error
- B.A preprocessing step applied only to the raw input data
- C.Normalizing the inputs of each layer to stabilize and accelerate training✓ Correct
- D.A type of non-linear activation function for hidden layers
Explanation
Batch normalization normalizes layer inputs across a mini-batch, reducing internal covariate shift and allowing higher learning rates.
Report an error in this question
Deep Learning FundamentalsMedium
Q21. What is weight initialization and why is it important?
- A.Setting initial weights to enable effective training and avoid gradient problems✓ Correct
- B.Weights are assigned their final values only after training
- C.Weights are always initialized to exactly zero for every neuron
- D.The initialization strategy has no effect on the training process
Explanation
Proper weight initialization (e.g., Xavier, He) ensures gradients flow well during training, preventing vanishing or exploding gradients from the start.
Report an error in this question
Deep Learning FundamentalsHard
Q22. What is the purpose of skip connections (residual connections)?
- A.Removing entire layers from the network to reduce complexity
- B.Skipping the data preprocessing step to speed up training
- C.Reducing the size of the training dataset to save storage
- D.Allowing gradients to flow directly by adding input to output of a block✓ Correct
Explanation
Skip connections add the input of a block to its output, enabling gradient flow across many layers and allowing networks to learn residual functions, as in ResNets.
Report an error in this question
Deep Learning FundamentalsHard
Q23. What is the difference between SGD with momentum and Adam?
- A.SGD momentum uses a fixed learning rate with velocity; Adam adapts per parameter✓ Correct
- B.Adam optimizer never converges to a good solution on real data
- C.SGD with momentum is always faster than Adam for every problem
- D.They are completely identical optimization algorithms with no differences
Explanation
SGD with momentum accumulates velocity but uses a global learning rate. Adam adapts per-parameter learning rates using running averages of gradients and squared gradients.
Report an error in this question
Deep Learning FundamentalsHard
Q24. What is the effect of batch size on training?
- A.Smaller batch sizes are always strictly worse for model generalization
- B.Larger batch sizes are always strictly better for all training scenarios
- C.Larger batches give smoother gradients but may generalize worse; smaller batches add helpful noise✓ Correct
- D.Batch size has absolutely no measurable effect on training or generalization
Explanation
Larger batch sizes provide smoother gradient estimates but may converge to sharp minima that generalize poorly. Smaller batches introduce noise that can help escape sharp minima.
Report an error in this question
Deep Learning FundamentalsHard
Q25. What is the Lottery Ticket Hypothesis?
- A.All neural network architectures are equally effective regardless of size
- B.Sparse network architectures never work well for any real task
- C.Dense networks contain sparse subnetworks that can match full network performance✓ Correct
- D.Larger networks are always strictly better than smaller alternatives
Explanation
The Lottery Ticket Hypothesis posits that randomly-initialized dense networks contain sparse subnetworks ('winning tickets') that can match the full network's performance.
Report an error in this question
Deep Learning FundamentalsHard
Q26. What is knowledge distillation?
- A.Removing irrelevant features from the input feature space
- B.Training a smaller student model to mimic a larger teacher model✓ Correct
- C.Transferring database records between different storage systems
- D.Augmenting the training data with synthetic generated samples
Explanation
Knowledge distillation trains a compact student model to reproduce the behavior of a larger teacher model, compressing knowledge into a smaller, deployable model.
Report an error in this question
Deep Learning FundamentalsHard
Q27. What is the difference between pre-training and fine-tuning?
- A.They are completely identical stages in the same training process
- B.Fine-tuning always happens before the pre-training phase
- C.Pre-training is designed specifically for very small datasets
- D.Pre-training learns general representations; fine-tuning adapts to a specific task✓ Correct
Explanation
Pre-training learns general features from large datasets. Fine-tuning adapts those learned representations to a specific downstream task with task-specific data.
Report an error in this question
Deep Learning FundamentalsHard
Q28. What is neural architecture search (NAS)?
- A.Automated methods for finding optimal neural network architectures✓ Correct
- B.Feature selection and dimensionality reduction algorithms
- C.Manual hand-designed network architecture engineering process
- D.Data preprocessing and cleaning pipeline automation tools
Explanation
NAS uses optimization techniques to automatically discover optimal network architectures, reducing the need for manual architecture engineering.
Report an error in this question
Deep Learning FundamentalsHard
Q29. What is gradient clipping?
- A.Removing all gradients entirely from the backpropagation computation
- B.Setting all gradient values to exactly zero after each update step
- C.Increasing the learning rate to speed up convergence time
- D.Limiting gradient values to a maximum threshold to prevent exploding gradients✓ Correct
Explanation
Gradient clipping caps gradient values when they exceed a threshold, preventing exploding gradients while maintaining training stability.
Report an error in this question
Deep Learning FundamentalsHard
Q30. What is mixed-precision training?
- A.Using both 16-bit and 32-bit floating point to speed up training while maintaining accuracy✓ Correct
- B.Using only 8-bit integer arithmetic for all network computation steps
- C.Using only 64-bit double-precision floats for maximum numerical precision
- D.Training neural network models entirely without any GPU acceleration
Explanation
Mixed-precision training uses FP16 for most computations for speed and FP32 for critical accumulations for accuracy, reducing memory usage and increasing throughput.
Report an error in this question