Each question below shows the correct answer with a full explanation. Use these to build conceptual understanding before attempting a timed quiz.
Advanced Deep LearningEasy
Q1. What is the purpose of a convolutional filter (kernel)?
- A.To add random noise for data augmentation purposes
- B.To store training data on disk for later retrieval access
- C.To increase the spatial resolution of the input image
- D.To detect specific features like edges or textures in input data✓ Correct
Explanation
Convolutional filters slide over input data to detect local patterns like edges, textures, and shapes through learned weights.
Report an error in this question
Advanced Deep LearningEasy
Q2. What is a Convolutional Neural Network (CNN)?
- A.A neural network designed for processing grid-like data such as images✓ Correct
- B.A relational database system for storing structured records
- C.A recurrent network designed exclusively for text sequence data
- D.A comparison-based sorting algorithm for numerical arrays
Explanation
CNNs use convolutional layers to automatically learn spatial hierarchies of features, making them ideal for image processing.
Report an error in this question
Advanced Deep LearningEasy
Q3. What is a GAN?
- A.An unsupervised clustering method based on density estimation
- B.A specialized variant of convolutional neural networks for images
- C.Generative Adversarial Network - two networks competing to generate realistic data✓ Correct
- D.A supervised regression algorithm for predicting numeric values
Explanation
A GAN consists of a generator that creates fake data and a discriminator that tries to distinguish real from fake, training adversarially.
Report an error in this question
Advanced Deep LearningEasy
Q4. What is an autoencoder?
- A.A neural network that learns to compress and reconstruct its input✓ Correct
- B.A comparison-based sorting algorithm for numerical arrays
- C.A data loading utility for reading files into memory
- D.A supervised classifier that predicts categorical labels from data
Explanation
An autoencoder learns to encode input into a lower-dimensional representation and then decode it back, useful for dimensionality reduction and denoising.
Report an error in this question
Advanced Deep LearningEasy
Q5. What is a Recurrent Neural Network (RNN)?
- A.A neural network for sequential data with memory of previous inputs✓ Correct
- B.A convolutional network designed exclusively for image pixel data
- C.An unsupervised density-based clustering algorithm for grouping
- D.A static feedforward network with no temporal connections
Explanation
RNNs process sequential data by maintaining hidden states that capture information from previous time steps.
Report an error in this question
Advanced Deep LearningEasy
Q6. What is a pooling layer in a CNN?
- A.A data loading layer for reading external files
- B.The final output layer for class predictions
- C.A layer that increases the total number of neurons
- D.A layer that reduces spatial dimensions by down-sampling✓ Correct
Explanation
Pooling layers reduce the spatial dimensions of feature maps (e.g., max pooling, average pooling), decreasing computation and providing translation invariance.
Report an error in this question
Advanced Deep LearningEasy
Q7. What is LSTM?
- A.Long Short-Term Memory - an RNN variant handling long-term dependencies✓ Correct
- B.A cross-entropy loss function for multi-class classification
- C.A specialized type of convolutional neural network for images
- D.A structured data format for storing tabular information
Explanation
LSTM is a special RNN architecture with gates (forget, input, output) that can learn long-term dependencies, solving the vanishing gradient problem in standard RNNs.
Report an error in this question
Advanced Deep LearningMedium
Q8. What is a Variational Autoencoder (VAE)?
- A.A linear regression model for numeric outputs
- B.A generative model that learns a probabilistic latent space✓ Correct
- C.A standard deterministic autoencoder without sampling
- D.A supervised classifier for categorical prediction
Explanation
A VAE learns a probability distribution in the latent space, enabling generation of new data by sampling from this distribution, unlike standard autoencoders.
Report an error in this question
Advanced Deep LearningEasy
Q9. What framework was developed by Facebook (Meta) for deep learning?
- A.Caffe
- B.Keras
- C.PyTorch✓ Correct
- D.TensorFlow
Explanation
PyTorch was developed by Meta (Facebook) and is widely used for deep learning research and production.
Report an error in this question
Advanced Deep LearningMedium
Q10. What is the Transformer architecture?
- A.A model based on self-attention mechanisms without recurrence✓ Correct
- B.A generative adversarial network for data synthesis
- C.A density-based clustering algorithm for grouping
- D.A standard convolutional neural network for image tasks
Explanation
The Transformer uses self-attention to process all positions simultaneously, enabling parallelization and capturing long-range dependencies more effectively than RNNs.
Report an error in this question
Advanced Deep LearningEasy
Q11. What is a generative model?
- A.A model that only performs supervised categorical classification
- B.A model that only performs numerical regression prediction
- C.A model that learns to generate new data similar to training data✓ Correct
- D.A model that only performs unsupervised data point clustering
Explanation
Generative models learn the underlying distribution of training data and can generate new, similar data samples.
Report an error in this question
Advanced Deep LearningEasy
Q12. What is a pre-trained model?
- A.A randomly initialized model with no learned representations
- B.A model with absolutely no prior training on any data
- C.A model that has been permanently deleted from storage
- D.A model already trained on a large dataset for use on new tasks✓ Correct
Explanation
Pre-trained models have been trained on large datasets (e.g., ImageNet) and can be fine-tuned for specific tasks, saving time and data.
Report an error in this question
Advanced Deep LearningMedium
Q13. What is 1x1 convolution used for?
- A.Increasing the spatial dimensions of the feature maps in the network✓ Correct
- B.Changing the number of channels and adding non-linearity without changing spatial dimensions
- C.Performing spatial pooling to reduce the feature map dimensions
- D.Applying batch normalization across all the feature map channels
Explanation
A 1x1 convolution changes the depth (number of channels) of feature maps, adds non-linearity, and reduces computation, as used in Inception networks.
Report an error in this question
Advanced Deep LearningMedium
Q14. What is self-attention?
- A.A cross-entropy loss function that measures classification prediction error
- B.A type of spatial pooling operation that reduces feature map dimensions
- C.A mechanism relating different positions within a sequence to compute representations✓ Correct
- D.A data augmentation method that generates synthetic training examples
Explanation
Self-attention computes attention weights between all pairs of positions in a sequence, allowing each position to attend to all others for contextual understanding.
Report an error in this question
Advanced Deep LearningMedium
Q15. What is the difference between GRU and LSTM?
- A.GRU is always more accurate than LSTM on all tasks
- B.GRU and LSTM are completely identical architectures
- C.GRU has fewer gates (2 vs 3) and is computationally simpler
- D.LSTM has fewer learnable parameters than the GRU✓ Correct
Explanation
GRU (Gated Recurrent Unit) uses 2 gates (reset and update) compared to LSTM's 3 gates, making it computationally simpler while achieving comparable performance.
Report an error in this question
Advanced Deep LearningMedium
Q16. What is data parallelism in deep learning?
- A.A type of batch normalization applied across distributed nodes
- B.Distributing training data across multiple GPUs, each running a model copy✓ Correct
- C.Reducing the overall size of the training data by sampling
- D.Using a single GPU for all training and inference computation
Explanation
Data parallelism replicates the model across multiple GPUs, each processing different data batches, and synchronizes gradients for faster training.
Report an error in this question
Advanced Deep LearningMedium
Q17. What is depthwise separable convolution?
- A.Performing a pooling operation to reduce the spatial feature dimensions✓ Correct
- B.Splitting convolution into depthwise and pointwise operations to reduce computation
- C.Applying a normalization technique to stabilize the training process
- D.Performing a standard full convolution operation on the input feature maps
Explanation
Depthwise separable convolution performs convolution on each channel independently, then combines with 1x1 convolutions, drastically reducing parameters and computation.
Report an error in this question
Advanced Deep LearningMedium
Q18. What is the encoder-decoder architecture?
- A.A cross-entropy loss function for measuring prediction quality
- B.A type of generative adversarial network with two competing agents
- C.A structured data format for storing model configuration files
- D.A structure where an encoder compresses input and a decoder generates output✓ Correct
Explanation
Encoder-decoder architectures compress input into a representation (encoding) and then generate the desired output (decoding), used in translation, segmentation, and more.
Report an error in this question
Advanced Deep LearningHard
Q19. What is the difference between causal and bidirectional self-attention?
- A.Causal attention only looks at previous positions; bidirectional looks at all positions
- B.Bidirectional attention only looks backward at previous tokens in sequence
- C.Causal and bidirectional attention are completely identical mechanisms✓ Correct
- D.Causal attention actually looks at all positions in the input sequence
Explanation
Causal (masked) attention prevents attending to future positions, used in autoregressive models like GPT. Bidirectional attention attends to all positions, used in models like BERT.
Report an error in this question
Advanced Deep LearningMedium
Q20. What is a residual network (ResNet)?
- A.A very shallow network with only a single hidden layer
- B.A standard recurrent neural network for sequences
- C.A deep network using skip connections that add input to output of layers✓ Correct
- D.A network architecture with no inter-layer connections
Explanation
ResNet introduces skip connections that bypass layers, enabling training of very deep networks (100+ layers) by mitigating the vanishing gradient problem.
Report an error in this question
Advanced Deep LearningMedium
Q21. What is the purpose of attention mechanisms?
- A.To allow models to focus on relevant parts of the input✓ Correct
- B.To increase the mini-batch size during training
- C.To initialize the network weights before training
- D.To reduce the learning rate over training epochs
Explanation
Attention mechanisms enable models to selectively focus on the most relevant parts of the input when producing each part of the output.
Report an error in this question
Advanced Deep LearningHard
Q22. What is the mode collapse problem in GANs?
- A.The generator produces limited variety instead of the full data distribution✓ Correct
- B.The model architecture has grown far too large for memory
- C.The discriminator fails completely and cannot distinguish anything
- D.The training process runs far too slowly to converge in time
Explanation
Mode collapse occurs when the generator learns to produce only a few types of outputs that fool the discriminator, failing to capture the full diversity of the data distribution.
Report an error in this question
Advanced Deep LearningHard
Q23. What is the Flash Attention algorithm?
- A.An IO-aware exact attention algorithm reducing memory from quadratic to linear✓ Correct
- B.A data loading technique for prefetching training batches from disk
- C.An approximate attention method that trades accuracy for modest speed
- D.A new non-linear activation function for transformer hidden layers
Explanation
Flash Attention computes exact attention by tiling and recomputation, reducing GPU memory reads/writes and achieving linear memory complexity while being faster than standard attention.
Report an error in this question
Advanced Deep LearningHard
Q24. What is Neural ODE?
- A.Modeling continuous-depth neural networks using ordinary differential equations✓ Correct
- B.A generative adversarial network variant for image synthesis
- C.A standard discrete-layer feedforward neural network architecture
- D.An unsupervised density-based technique for data clustering
Explanation
Neural ODEs treat hidden state dynamics as continuous-time ODEs solved with ODE solvers, enabling adaptive computation and continuous-depth networks.
Report an error in this question
Advanced Deep LearningHard
Q25. What is the concept of equivariance in CNNs?
- A.Output is always identical regardless of any transformation of input✓ Correct
- B.All layers in the network have exactly equal learned weight values
- C.Output transforms in the same way as the input (e.g., shift in input causes shift in output)
- D.The neural network architecture is perfectly symmetric across all layers
Explanation
Equivariance means if the input is transformed (e.g., translated), the output transforms correspondingly. CNNs are translation equivariant in convolutional layers.
Report an error in this question
Advanced Deep LearningHard
Q26. What is quantization in deep learning?
- A.Increasing the overall model size by adding more trainable parameters
- B.Applying data augmentation transforms to expand the training set
- C.Reducing weight and activation precision to lower bit representations for efficiency✓ Correct
- D.Adding more hidden layers to increase the depth of the network
Explanation
Quantization converts model weights and activations from high-precision (FP32) to lower-precision formats (INT8, INT4), reducing model size and inference latency with minimal accuracy loss.
Report an error in this question
Advanced Deep LearningHard
Q27. What is the Mixture of Experts (MoE) architecture?
- A.A standard model ensemble using simple majority voting
- B.A data preprocessing method for feature normalization
- C.A single monolithic large network without any routing
- D.A model using a gating network to route inputs to specialized sub-networks✓ Correct
Explanation
MoE uses a gating mechanism to dynamically route each input to a subset of specialized expert networks, enabling larger model capacity with manageable computation.
Report an error in this question
Advanced Deep LearningHard
Q28. What is a diffusion model?
- A.A recurrent neural network architecture for sequence modeling
- B.An unsupervised density-based clustering algorithm for data
- C.A generative model that learns to reverse a gradual noising process✓ Correct
- D.A generative adversarial network variant using two competing agents
Explanation
Diffusion models generate data by learning to reverse a gradual process of adding noise to data, achieving state-of-the-art results in image generation.
Report an error in this question
Advanced Deep LearningHard
Q29. What is the Wasserstein GAN (WGAN) and how does it improve training?
- A.It uses only convolutional layers without any dense layers
- B.It uses the Wasserstein distance for more stable training gradients✓ Correct
- C.It completely removes the discriminator from the architecture
- D.It uses standard cross-entropy loss for the discriminator
Explanation
WGAN uses the Wasserstein (Earth Mover's) distance instead of JS divergence, providing smoother gradients even when distributions don't overlap, leading to more stable training.
Report an error in this question
Advanced Deep LearningHard
Q30. What is the difference between model parallelism and data parallelism?
- A.Data parallelism splits the model layers across devices for training
- B.Model parallelism splits the model across devices; data parallelism replicates model with split data✓ Correct
- C.Model parallelism splits the data while keeping the full model on one device
- D.They are completely identical distributed training strategies with no differences
Explanation
Data parallelism replicates the model on each device with different data. Model parallelism distributes different parts of the model across devices, needed when the model doesn't fit on one device.
Report an error in this question