HomeSubjectsUniversityBlogAbout
210 total MCQsShowing 30 with explanations10 Easy10 Medium10 Hard

About This Topic

Computer vision is the field that lets computers extract meaningful information from images and video, such as recognizing objects, faces and motion. Questions cover image basics like pixels, colour channels and filtering, classic edge detection with Sobel and Canny, and CNN-based image classification. Data augmentation items ask about flipping, rotation, cropping and colour jitter. Detection questions compare two-stage detectors such as Faster R-CNN, which uses a region proposal network, with one-stage detectors like YOLO and SSD, and define IoU and non-maximum suppression. Segmentation with U-Net and Mask R-CNN, and 3D representations like neural radiance fields, also appear.

Below are 30 practice questions from a pool of 210 Computer Vision MCQs, one of 17 topics in AI / Machine Learning & Data Analytics. Each shows the correct answer with an explanation; when you are ready, take a timed quiz to test recall under exam conditions.

Practice Questions

Each question below shows the correct answer with a full explanation. Use these to build conceptual understanding before attempting a timed quiz.

Computer VisionEasy

Q1. What is object detection?

  1. A.Removing detected objects from an existing image frame
  2. B.Classifying the entire image into a single category label
  3. C.Identifying and locating objects within an image with bounding boxes✓ Correct
  4. D.Generating entirely new photorealistic synthetic images

Explanation

Object detection identifies what objects are present in an image and where they are located, typically using bounding boxes.

Report an error in this question

Computer VisionEasy

Q2. What is a pixel?

  1. A.An input feature column name
  2. B.A gradient training algorithm
  3. C.A type of deep neural network
  4. D.The smallest unit of a digital image✓ Correct

Explanation

A pixel (picture element) is the smallest addressable element of a digital image, with color values representing visual information.

Report an error in this question

Computer VisionEasy

Q3. What is image classification?

  1. A.Segmenting an image into parts
  2. B.Assigning a label to an entire image
  3. C.Detecting objects in a given image✓ Correct
  4. D.Generating an image from scratch

Explanation

Image classification assigns a single category label to an entire image based on its visual content.

Report an error in this question

Computer VisionEasy

Q4. Computer vision enables computers to:

  1. A.Interpret and understand visual information from images and videos✓ Correct
  2. B.Design and build responsive web page user interfaces
  3. C.Manage and query relational database table records
  4. D.Process only natural language text data and documents

Explanation

Computer vision is a field of AI that enables computers to extract meaningful information from images, videos, and other visual inputs.

Report an error in this question

Computer VisionEasy

Q5. What does a convolution operation do to an image?

  1. A.Applies a filter to extract features like edges and textures✓ Correct
  2. B.Increases the image resolution to a higher pixel count
  3. C.Changes the file format of the image container file
  4. D.Deletes the image entirely from disk storage space

Explanation

Convolution slides a filter over the image, computing element-wise multiplications and sums to detect local features like edges, corners, and textures.

Report an error in this question

Computer VisionHard

Q6. What is contrastive learning in computer vision (e.g., SimCLR)?

  1. A.A two-stage region-proposal based object detection algorithm pipeline
  2. B.Learning visual representations by contrasting augmented views of the same image against others✓ Correct
  3. C.A fully supervised classification method that requires complete labeled training data
  4. D.A generative adversarial network training method for synthesizing realistic images

Explanation

SimCLR and similar methods learn representations by maximizing agreement between differently augmented views of the same image while minimizing agreement with other images.

Report an error in this question

Computer VisionMedium

Q7. What is image feature extraction?

  1. A.Adding textual annotation overlays onto the image pixels
  2. B.Compressing images into smaller file archive formats
  3. C.Identifying and representing distinctive visual patterns in images✓ Correct
  4. D.Deleting specific features or layers from the image data

Explanation

Feature extraction identifies meaningful visual patterns (edges, textures, shapes) that can be represented as numerical features for downstream tasks.

Report an error in this question

Computer VisionMedium

Q8. What is the difference between object detection and image classification?

  1. A.Image classification actually locates and draws bounding boxes around objects
  2. B.Detection locates multiple objects with bounding boxes; classification assigns one label to the whole image✓ Correct
  3. C.Object detection assigns a single label to the entire image without any localization
  4. D.Object detection and image classification are completely identical tasks

Explanation

Image classification assigns a single label to the entire image, while object detection identifies and locates multiple objects with bounding boxes and class labels.

Report an error in this question

Computer VisionEasy

Q9. What is image segmentation?

  1. A.Resizing an image to different specified dimensions
  2. B.Partitioning an image into meaningful regions at the pixel level
  3. C.Rotating an image by a specified number of degrees✓ Correct
  4. D.Cropping an image to a smaller rectangular area

Explanation

Image segmentation assigns a label to every pixel in an image, dividing it into meaningful regions or objects.

Report an error in this question

Computer VisionHard

Q10. What is neural style transfer?

  1. A.Classifying images into predefined categorical labels with CNNs
  2. B.Applying the artistic style of one image to the content of another using CNNs✓ Correct
  3. C.Training a brand new neural network from scratch on random data
  4. D.Applying data augmentation transforms to expand the training set

Explanation

Neural style transfer separates and recombines the content of one image with the style of another by optimizing an image to match content and style feature representations from different layers of a CNN.

Report an error in this question

Computer VisionMedium

Q11. What is optical flow?

  1. A.A type of spatial convolutional image filter kernel operation
  2. B.A standard color model representation for digital images
  3. C.The pattern of apparent motion of objects between consecutive frames✓ Correct
  4. D.A container file format for storing compressed video data

Explanation

Optical flow estimates the motion of pixels between consecutive video frames, used in video analysis, action recognition, and motion tracking.

Report an error in this question

Computer VisionHard

Q12. What is the Feature Pyramid Network (FPN)?

  1. A.A single-scale object detector that operates at only one fixed resolution level
  2. B.A standard image classification network without multi-scale feature fusion
  3. C.A generative adversarial network for producing synthetic realistic images
  4. D.A multi-scale feature extractor combining low-res semantic and high-res spatial features✓ Correct

Explanation

FPN creates a top-down architecture with lateral connections to build a feature pyramid, enabling detection of objects at multiple scales with rich semantics at all levels.

Report an error in this question

Computer VisionEasy

Q13. RGB stands for:

  1. A.Random, Generated, Binary
  2. B.Recursive, Generative, Base
  3. C.Real, Gradient, Bias
  4. D.Red, Green, Blue✓ Correct

Explanation

RGB (Red, Green, Blue) is the standard color model for digital images, where each pixel has three channel values.

Report an error in this question

Computer VisionEasy

Q14. Face detection is an example of:

  1. A.Reinforcement learning
  2. B.Computer vision✓ Correct
  3. C.Data analytics
  4. D.Natural language processing

Explanation

Face detection uses computer vision techniques to identify and locate human faces in images or video.

Report an error in this question

Computer VisionHard

Q15. What is the concept of deformable convolutions?

  1. A.Standard fixed-grid convolutions with no adaptive sampling offsets
  2. B.Convolutions with learnable offsets that adapt the sampling grid to object shapes✓ Correct
  3. C.One-by-one pointwise convolutions for changing channel depth
  4. D.Dilated convolutions that increase the receptive field without offsets

Explanation

Deformable convolutions add learnable offsets to the regular grid sampling positions, enabling the convolution to adapt to geometric transformations and object deformations.

Report an error in this question

Computer VisionEasy

Q16. What is a grayscale image?

  1. A.An image with only shades of gray (single channel)✓ Correct
  2. B.A volumetric three-dimensional image stack
  3. C.A binary black and white image with two values
  4. D.A full-color image with three RGB channels

Explanation

A grayscale image has a single channel where each pixel represents intensity from black (0) to white (255).

Report an error in this question

Computer VisionMedium

Q17. What are anchor boxes in object detection?

  1. A.The pixel coordinate system used for image representation
  2. B.The individual color channels within the input image data
  3. C.The border regions around the edges of an input image✓ Correct
  4. D.Predefined bounding box shapes that help detect objects of various sizes and ratios

Explanation

Anchor boxes are predefined bounding boxes of various sizes and aspect ratios used as reference shapes for predicting object locations in detection models.

Report an error in this question

Computer VisionMedium

Q18. What is the YOLO algorithm?

  1. A.A pixel-level semantic image segmentation algorithm
  2. B.A supervised multi-class image classification algorithm
  3. C.You Only Look Once - a real-time object detection algorithm✓ Correct
  4. D.An unsupervised K-means image clustering algorithm

Explanation

YOLO processes the entire image in a single pass through the network, predicting bounding boxes and class probabilities simultaneously for real-time object detection.

Report an error in this question

Computer VisionMedium

Q19. What is the IoU (Intersection over Union) metric?

  1. A.A metric measuring overlap between predicted and ground truth bounding boxes
  2. B.A metric that measures overall image quality and clarity
  3. C.A metric that assesses the overall color accuracy of images✓ Correct
  4. D.A metric for evaluating the spatial resolution of images

Explanation

IoU calculates the ratio of the intersection area to the union area of predicted and ground truth bounding boxes, commonly used to evaluate object detection.

Report an error in this question

Computer VisionHard

Q20. What is the difference between one-stage and two-stage object detectors?

  1. A.One-stage (YOLO) detects directly; two-stage (R-CNN) proposes regions then classifies✓ Correct
  2. B.One-stage detectors are always significantly slower than two-stage methods
  3. C.They are completely identical detection approaches with no performance differences
  4. D.Two-stage detectors do not use region proposals at any processing stage

Explanation

Two-stage detectors like Faster R-CNN first generate region proposals, then classify each. One-stage detectors like YOLO and SSD directly predict boxes and classes in one pass, trading accuracy for speed.

Report an error in this question

Computer VisionMedium

Q21. What is data augmentation in computer vision?

  1. A.Deleting images from the dataset to make it smaller and more manageable
  2. B.Reducing the overall image quality to speed up the training process
  3. C.Applying transformations like rotation, flipping, and cropping to increase training data
  4. D.Converting all images to grayscale only without any other transformation✓ Correct

Explanation

Data augmentation applies random transformations (rotation, flipping, scaling, color jittering) to training images, increasing dataset diversity and reducing overfitting.

Report an error in this question

Computer VisionMedium

Q22. What is the purpose of max pooling?

  1. A.Adding random noise to the feature map pixel values
  2. B.Changing the color representation of the image data
  3. C.Reducing spatial dimensions while retaining the most prominent features✓ Correct
  4. D.Increasing the spatial size of the feature map dimensions

Explanation

Max pooling selects the maximum value in each pooling window, reducing spatial dimensions while preserving the most activated features and providing some translation invariance.

Report an error in this question

Computer VisionMedium

Q23. What is the difference between semantic and instance segmentation?

  1. A.Instance segmentation completely ignores the object class labels
  2. B.Semantic segmentation is always the better approach overall✓ Correct
  3. C.Semantic labels every pixel by class; instance distinguishes individual objects of the same class
  4. D.Semantic and instance segmentation are completely identical approaches

Explanation

Semantic segmentation labels each pixel with a class but doesn't distinguish between instances. Instance segmentation identifies each individual object separately.

Report an error in this question

Computer VisionHard

Q24. What is panoptic segmentation?

  1. A.A task that performs only semantic segmentation without instances
  2. B.A task that performs only instance segmentation without semantics
  3. C.A task that is limited to standard object detection only
  4. D.A unified task combining semantic segmentation (stuff) and instance segmentation (things)✓ Correct

Explanation

Panoptic segmentation unifies semantic segmentation for amorphous regions (stuff like sky, road) and instance segmentation for countable objects (things like cars, people).

Report an error in this question

Computer VisionHard

Q25. What is the Vision Transformer (ViT)?

  1. A.A recurrent neural network architecture designed for image sequences
  2. B.A generative adversarial network for producing synthetic images
  3. C.A specialized convolutional neural network variant for image processing
  4. D.Applying the Transformer architecture directly to image patches for classification✓ Correct

Explanation

ViT splits images into patches, treats them as tokens, and applies a standard Transformer encoder for image classification, challenging the dominance of CNNs.

Report an error in this question

Computer VisionEasy

Q26. What is image resizing?

  1. A.Rotating the image by a specified angle value
  2. B.Changing the dimensions (width and height) of an image✓ Correct
  3. C.Deleting the image from the disk file system
  4. D.Changing the color space of the image channels

Explanation

Image resizing changes the width and height of an image, often required to match the input dimensions expected by neural networks.

Report an error in this question

Computer VisionMedium

Q27. What is non-maximum suppression (NMS)?

  1. A.A batch normalization method for stabilizing training
  2. B.A post-processing step that removes overlapping duplicate detections✓ Correct
  3. C.A gradient-based model training optimization technique
  4. D.A non-linear activation function for hidden layers

Explanation

NMS eliminates redundant overlapping bounding boxes by keeping only the detection with the highest confidence for each object.

Report an error in this question

Computer VisionHard

Q28. What is depth estimation in computer vision?

  1. A.Predicting the distance of each pixel from the camera✓ Correct
  2. B.Measuring the resolution of the image in pixels
  3. C.Classifying the entire image into a category
  4. D.Counting the total number of objects in a scene

Explanation

Monocular depth estimation predicts a depth map from a single image, while stereo depth estimation uses two cameras. It is crucial for 3D understanding and autonomous driving.

Report an error in this question

Computer VisionHard

Q29. What is the CLIP model?

  1. A.A model learning visual concepts from language supervision, connecting images and text✓ Correct
  2. B.A single-modality object detector that only processes image pixel data
  3. C.A standard supervised image classifier trained only on labeled image categories
  4. D.A text-only language model without any visual understanding capabilities

Explanation

CLIP (Contrastive Language-Image Pre-training) learns to connect images and text by training on image-text pairs, enabling zero-shot image classification using natural language descriptions.

Report an error in this question

Computer VisionHard

Q30. What is 3D point cloud processing?

  1. A.Processing natural language text data from document sources
  2. B.Processing unstructured 3D spatial data points from sensors like LiDAR✓ Correct
  3. C.Processing continuous audio waveform signals from microphones
  4. D.Processing standard two-dimensional image pixel data from cameras

Explanation

Point cloud processing handles sets of 3D points from sensors like LiDAR, using architectures like PointNet that directly process unstructured 3D data for tasks like segmentation and detection.

Report an error in this question

Ready to test yourself on Computer Vision?

Take a timed quiz drawn from 210+ questions on this topic. No signup required — your progress saves in your browser.

Start Computer Vision Quiz