HomeSubjectsUniversityBlogAbout

Advanced Deep Learning

Topic in AI / Machine Learning & Data Analytics

210 total MCQsShowing 30 with explanations10 Easy10 Medium10 Hard

About This Topic

Advanced deep learning covers specialized neural architectures, namely convolutional, recurrent, attention-based and generative models. Questions ask how convolution, pooling, stride and padding work in CNNs and why residual connections in ResNet ease training. Sequence items cover RNNs, the vanishing gradient problem, and LSTM and GRU gates, such as how the forget gate discards old memory. Transformer questions address self-attention, multi-head attention, positional encoding and attention's quadratic cost in sequence length. Generative topics include GANs with generator and discriminator, variational autoencoders, diffusion models and normalizing flows, plus transfer learning and fine-tuning.

Below are 30 practice questions from a pool of 210 Advanced Deep Learning MCQs, one of 17 topics in AI / Machine Learning & Data Analytics. Each shows the correct answer with an explanation; when you are ready, take a timed quiz to test recall under exam conditions.

Practice Questions

Each question below shows the correct answer with a full explanation. Use these to build conceptual understanding before attempting a timed quiz.

Advanced Deep LearningEasy

Q1. What is the purpose of a convolutional filter (kernel)?

  1. A.To add random noise for data augmentation purposes
  2. B.To store training data on disk for later retrieval access
  3. C.To increase the spatial resolution of the input image
  4. D.To detect specific features like edges or textures in input data✓ Correct

Explanation

Convolutional filters slide over input data to detect local patterns like edges, textures, and shapes through learned weights.

Report an error in this question

Advanced Deep LearningEasy

Q2. What is a Convolutional Neural Network (CNN)?

  1. A.A neural network designed for processing grid-like data such as images✓ Correct
  2. B.A relational database system for storing structured records
  3. C.A recurrent network designed exclusively for text sequence data
  4. D.A comparison-based sorting algorithm for numerical arrays

Explanation

CNNs use convolutional layers to automatically learn spatial hierarchies of features, making them ideal for image processing.

Report an error in this question

Advanced Deep LearningEasy

Q3. What is a GAN?

  1. A.An unsupervised clustering method based on density estimation
  2. B.A specialized variant of convolutional neural networks for images
  3. C.Generative Adversarial Network - two networks competing to generate realistic data✓ Correct
  4. D.A supervised regression algorithm for predicting numeric values

Explanation

A GAN consists of a generator that creates fake data and a discriminator that tries to distinguish real from fake, training adversarially.

Report an error in this question

Advanced Deep LearningEasy

Q4. What is an autoencoder?

  1. A.A neural network that learns to compress and reconstruct its input✓ Correct
  2. B.A comparison-based sorting algorithm for numerical arrays
  3. C.A data loading utility for reading files into memory
  4. D.A supervised classifier that predicts categorical labels from data

Explanation

An autoencoder learns to encode input into a lower-dimensional representation and then decode it back, useful for dimensionality reduction and denoising.

Report an error in this question

Advanced Deep LearningEasy

Q5. What is a Recurrent Neural Network (RNN)?

  1. A.A neural network for sequential data with memory of previous inputs✓ Correct
  2. B.A convolutional network designed exclusively for image pixel data
  3. C.An unsupervised density-based clustering algorithm for grouping
  4. D.A static feedforward network with no temporal connections

Explanation

RNNs process sequential data by maintaining hidden states that capture information from previous time steps.

Report an error in this question

Advanced Deep LearningEasy

Q6. What is a pooling layer in a CNN?

  1. A.A data loading layer for reading external files
  2. B.The final output layer for class predictions
  3. C.A layer that increases the total number of neurons
  4. D.A layer that reduces spatial dimensions by down-sampling✓ Correct

Explanation

Pooling layers reduce the spatial dimensions of feature maps (e.g., max pooling, average pooling), decreasing computation and providing translation invariance.

Report an error in this question

Advanced Deep LearningEasy

Q7. What is LSTM?

  1. A.Long Short-Term Memory - an RNN variant handling long-term dependencies✓ Correct
  2. B.A cross-entropy loss function for multi-class classification
  3. C.A specialized type of convolutional neural network for images
  4. D.A structured data format for storing tabular information

Explanation

LSTM is a special RNN architecture with gates (forget, input, output) that can learn long-term dependencies, solving the vanishing gradient problem in standard RNNs.

Report an error in this question

Advanced Deep LearningMedium

Q8. What is a Variational Autoencoder (VAE)?

  1. A.A linear regression model for numeric outputs
  2. B.A generative model that learns a probabilistic latent space✓ Correct
  3. C.A standard deterministic autoencoder without sampling
  4. D.A supervised classifier for categorical prediction

Explanation

A VAE learns a probability distribution in the latent space, enabling generation of new data by sampling from this distribution, unlike standard autoencoders.

Report an error in this question

Advanced Deep LearningEasy

Q9. What framework was developed by Facebook (Meta) for deep learning?

  1. A.Caffe
  2. B.Keras
  3. C.PyTorch✓ Correct
  4. D.TensorFlow

Explanation

PyTorch was developed by Meta (Facebook) and is widely used for deep learning research and production.

Report an error in this question

Advanced Deep LearningMedium

Q10. What is the Transformer architecture?

  1. A.A model based on self-attention mechanisms without recurrence✓ Correct
  2. B.A generative adversarial network for data synthesis
  3. C.A density-based clustering algorithm for grouping
  4. D.A standard convolutional neural network for image tasks

Explanation

The Transformer uses self-attention to process all positions simultaneously, enabling parallelization and capturing long-range dependencies more effectively than RNNs.

Report an error in this question

Advanced Deep LearningEasy

Q11. What is a generative model?

  1. A.A model that only performs supervised categorical classification
  2. B.A model that only performs numerical regression prediction
  3. C.A model that learns to generate new data similar to training data✓ Correct
  4. D.A model that only performs unsupervised data point clustering

Explanation

Generative models learn the underlying distribution of training data and can generate new, similar data samples.

Report an error in this question

Advanced Deep LearningEasy

Q12. What is a pre-trained model?

  1. A.A randomly initialized model with no learned representations
  2. B.A model with absolutely no prior training on any data
  3. C.A model that has been permanently deleted from storage
  4. D.A model already trained on a large dataset for use on new tasks✓ Correct

Explanation

Pre-trained models have been trained on large datasets (e.g., ImageNet) and can be fine-tuned for specific tasks, saving time and data.

Report an error in this question

Advanced Deep LearningMedium

Q13. What is 1x1 convolution used for?

  1. A.Increasing the spatial dimensions of the feature maps in the network✓ Correct
  2. B.Changing the number of channels and adding non-linearity without changing spatial dimensions
  3. C.Performing spatial pooling to reduce the feature map dimensions
  4. D.Applying batch normalization across all the feature map channels

Explanation

A 1x1 convolution changes the depth (number of channels) of feature maps, adds non-linearity, and reduces computation, as used in Inception networks.

Report an error in this question

Advanced Deep LearningMedium

Q14. What is self-attention?

  1. A.A cross-entropy loss function that measures classification prediction error
  2. B.A type of spatial pooling operation that reduces feature map dimensions
  3. C.A mechanism relating different positions within a sequence to compute representations✓ Correct
  4. D.A data augmentation method that generates synthetic training examples

Explanation

Self-attention computes attention weights between all pairs of positions in a sequence, allowing each position to attend to all others for contextual understanding.

Report an error in this question

Advanced Deep LearningMedium

Q15. What is the difference between GRU and LSTM?

  1. A.GRU is always more accurate than LSTM on all tasks
  2. B.GRU and LSTM are completely identical architectures
  3. C.GRU has fewer gates (2 vs 3) and is computationally simpler
  4. D.LSTM has fewer learnable parameters than the GRU✓ Correct

Explanation

GRU (Gated Recurrent Unit) uses 2 gates (reset and update) compared to LSTM's 3 gates, making it computationally simpler while achieving comparable performance.

Report an error in this question

Advanced Deep LearningMedium

Q16. What is data parallelism in deep learning?

  1. A.A type of batch normalization applied across distributed nodes
  2. B.Distributing training data across multiple GPUs, each running a model copy✓ Correct
  3. C.Reducing the overall size of the training data by sampling
  4. D.Using a single GPU for all training and inference computation

Explanation

Data parallelism replicates the model across multiple GPUs, each processing different data batches, and synchronizes gradients for faster training.

Report an error in this question

Advanced Deep LearningMedium

Q17. What is depthwise separable convolution?

  1. A.Performing a pooling operation to reduce the spatial feature dimensions✓ Correct
  2. B.Splitting convolution into depthwise and pointwise operations to reduce computation
  3. C.Applying a normalization technique to stabilize the training process
  4. D.Performing a standard full convolution operation on the input feature maps

Explanation

Depthwise separable convolution performs convolution on each channel independently, then combines with 1x1 convolutions, drastically reducing parameters and computation.

Report an error in this question

Advanced Deep LearningMedium

Q18. What is the encoder-decoder architecture?

  1. A.A cross-entropy loss function for measuring prediction quality
  2. B.A type of generative adversarial network with two competing agents
  3. C.A structured data format for storing model configuration files
  4. D.A structure where an encoder compresses input and a decoder generates output✓ Correct

Explanation

Encoder-decoder architectures compress input into a representation (encoding) and then generate the desired output (decoding), used in translation, segmentation, and more.

Report an error in this question

Advanced Deep LearningHard

Q19. What is the difference between causal and bidirectional self-attention?

  1. A.Causal attention only looks at previous positions; bidirectional looks at all positions
  2. B.Bidirectional attention only looks backward at previous tokens in sequence
  3. C.Causal and bidirectional attention are completely identical mechanisms✓ Correct
  4. D.Causal attention actually looks at all positions in the input sequence

Explanation

Causal (masked) attention prevents attending to future positions, used in autoregressive models like GPT. Bidirectional attention attends to all positions, used in models like BERT.

Report an error in this question

Advanced Deep LearningMedium

Q20. What is a residual network (ResNet)?

  1. A.A very shallow network with only a single hidden layer
  2. B.A standard recurrent neural network for sequences
  3. C.A deep network using skip connections that add input to output of layers✓ Correct
  4. D.A network architecture with no inter-layer connections

Explanation

ResNet introduces skip connections that bypass layers, enabling training of very deep networks (100+ layers) by mitigating the vanishing gradient problem.

Report an error in this question

Advanced Deep LearningMedium

Q21. What is the purpose of attention mechanisms?

  1. A.To allow models to focus on relevant parts of the input✓ Correct
  2. B.To increase the mini-batch size during training
  3. C.To initialize the network weights before training
  4. D.To reduce the learning rate over training epochs

Explanation

Attention mechanisms enable models to selectively focus on the most relevant parts of the input when producing each part of the output.

Report an error in this question

Advanced Deep LearningHard

Q22. What is the mode collapse problem in GANs?

  1. A.The generator produces limited variety instead of the full data distribution✓ Correct
  2. B.The model architecture has grown far too large for memory
  3. C.The discriminator fails completely and cannot distinguish anything
  4. D.The training process runs far too slowly to converge in time

Explanation

Mode collapse occurs when the generator learns to produce only a few types of outputs that fool the discriminator, failing to capture the full diversity of the data distribution.

Report an error in this question

Advanced Deep LearningHard

Q23. What is the Flash Attention algorithm?

  1. A.An IO-aware exact attention algorithm reducing memory from quadratic to linear✓ Correct
  2. B.A data loading technique for prefetching training batches from disk
  3. C.An approximate attention method that trades accuracy for modest speed
  4. D.A new non-linear activation function for transformer hidden layers

Explanation

Flash Attention computes exact attention by tiling and recomputation, reducing GPU memory reads/writes and achieving linear memory complexity while being faster than standard attention.

Report an error in this question

Advanced Deep LearningHard

Q24. What is Neural ODE?

  1. A.Modeling continuous-depth neural networks using ordinary differential equations✓ Correct
  2. B.A generative adversarial network variant for image synthesis
  3. C.A standard discrete-layer feedforward neural network architecture
  4. D.An unsupervised density-based technique for data clustering

Explanation

Neural ODEs treat hidden state dynamics as continuous-time ODEs solved with ODE solvers, enabling adaptive computation and continuous-depth networks.

Report an error in this question

Advanced Deep LearningHard

Q25. What is the concept of equivariance in CNNs?

  1. A.Output is always identical regardless of any transformation of input✓ Correct
  2. B.All layers in the network have exactly equal learned weight values
  3. C.Output transforms in the same way as the input (e.g., shift in input causes shift in output)
  4. D.The neural network architecture is perfectly symmetric across all layers

Explanation

Equivariance means if the input is transformed (e.g., translated), the output transforms correspondingly. CNNs are translation equivariant in convolutional layers.

Report an error in this question

Advanced Deep LearningHard

Q26. What is quantization in deep learning?

  1. A.Increasing the overall model size by adding more trainable parameters
  2. B.Applying data augmentation transforms to expand the training set
  3. C.Reducing weight and activation precision to lower bit representations for efficiency✓ Correct
  4. D.Adding more hidden layers to increase the depth of the network

Explanation

Quantization converts model weights and activations from high-precision (FP32) to lower-precision formats (INT8, INT4), reducing model size and inference latency with minimal accuracy loss.

Report an error in this question

Advanced Deep LearningHard

Q27. What is the Mixture of Experts (MoE) architecture?

  1. A.A standard model ensemble using simple majority voting
  2. B.A data preprocessing method for feature normalization
  3. C.A single monolithic large network without any routing
  4. D.A model using a gating network to route inputs to specialized sub-networks✓ Correct

Explanation

MoE uses a gating mechanism to dynamically route each input to a subset of specialized expert networks, enabling larger model capacity with manageable computation.

Report an error in this question

Advanced Deep LearningHard

Q28. What is a diffusion model?

  1. A.A recurrent neural network architecture for sequence modeling
  2. B.An unsupervised density-based clustering algorithm for data
  3. C.A generative model that learns to reverse a gradual noising process✓ Correct
  4. D.A generative adversarial network variant using two competing agents

Explanation

Diffusion models generate data by learning to reverse a gradual process of adding noise to data, achieving state-of-the-art results in image generation.

Report an error in this question

Advanced Deep LearningHard

Q29. What is the Wasserstein GAN (WGAN) and how does it improve training?

  1. A.It uses only convolutional layers without any dense layers
  2. B.It uses the Wasserstein distance for more stable training gradients✓ Correct
  3. C.It completely removes the discriminator from the architecture
  4. D.It uses standard cross-entropy loss for the discriminator

Explanation

WGAN uses the Wasserstein (Earth Mover's) distance instead of JS divergence, providing smoother gradients even when distributions don't overlap, leading to more stable training.

Report an error in this question

Advanced Deep LearningHard

Q30. What is the difference between model parallelism and data parallelism?

  1. A.Data parallelism splits the model layers across devices for training
  2. B.Model parallelism splits the model across devices; data parallelism replicates model with split data✓ Correct
  3. C.Model parallelism splits the data while keeping the full model on one device
  4. D.They are completely identical distributed training strategies with no differences

Explanation

Data parallelism replicates the model on each device with different data. Model parallelism distributes different parts of the model across devices, needed when the model doesn't fit on one device.

Report an error in this question

Ready to test yourself on Advanced Deep Learning?

Take a timed quiz drawn from 210+ questions on this topic. No signup required — your progress saves in your browser.

Start Advanced Deep Learning Quiz