MauryaHub PYQ Practice

da5006_2026T1_ET_FN.pdf

Deep Learning for Computer Vision · End Term · Jan 2026 FN

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 2 MCQ · 2.0 marks

A standard RGB image of size 256×256 is stored as 8-bit unsigned integers per channel. What is the total memory (in bytes) to store the raw image (no compression)?
  1. 256×256 = 65,536 bytes
  2. 256×256×2 = 131,072 bytes
  3. 256×256×3 = 196,608 bytes
  4. 256×256×3×8 = 1,572,864 bytes

A published solution is not available for this question yet.

Question 3 MCQ · 2.0 marks

Which statement is correct?
  1. Convolution equals correlation for all filters.
  2. Convolution flips the kernel spatially; correlation does not.
  3. Correlation flips the kernel; convolution does not.
  4. Both operations always preserve energy.

A published solution is not available for this question yet.

Question 4 MCQ · 2.0 marks

You apply a 5×5 filter with stride 1 and **no padding** on a 32×32 image. Output size is:
  1. 32×32
  2. 28×28
  3. 27×27
  4. 26×26

A published solution is not available for this question yet.

Question 5 MCQ · 2.0 marks

The Sobel operator primarily estimates:
  1. [[IMAGE:c3314bd6a89b0d7f_3_2]]
    Source diagram or notation
  2. [[IMAGE:c3314bd6a89b0d7f_3_3]]
    Source diagram or notation
  3. [[IMAGE:c3314bd6a89b0d7f_3_4]]
    Source diagram or notation
  4. [[IMAGE:c3314bd6a89b0d7f_3_5]]
    Source diagram or notation

A published solution is not available for this question yet.

Question 6 MCQ · 2.0 marks

SIFT is designed to be robust mainly to:
  1. Scale and rotation changes
  2. Only translation (shifts)
  3. Only illumination changes
  4. Only affine viewpoint changes with no residual error

A published solution is not available for this question yet.

Question 7 MCQ · 2.0 marks

In classical scale-space theory, increasing σ in Gaussian smoothing generally:
  1. Enhances high-frequency details
  2. Suppresses high-frequency details and removes fine structures
  3. Has no effect on edges
  4. Increases noise

A published solution is not available for this question yet.

Question 8 MCQ · 2.0 marks

Backpropagation is best described as:
  1. A method to avoid overfitting
  2. Efficient computation of gradients using the chain rule through computation graphs
  3. A method that works only for linear models
  4. An optimizer like Adam

A published solution is not available for this question yet.

Question 9 MCQ · 2.0 marks

Which is a key effect of momentum?
  1. Makes gradients unbiased
  2. Accumulates a velocity term
  3. Gurantees to converge unlike vanilla Gradient Descent
  4. Eliminates need for setting learning rate by adpatly calculaitng it on the fly.

A published solution is not available for this question yet.

Question 10 MCQ · 2.0 marks

L2 regularization on weights most directly encourages:
  1. Sparse weights with many exact zeros
  2. Smaller-magnitude weights
  3. Larger-magnitude weights
  4. Larger learning rates

A published solution is not available for this question yet.

Question 11 MCQ · 2.0 marks

A key reason CNNs are parameter-efficient compared to fully connected nets for images is:
  1. Weight sharing and local receptive fields
  2. They never use nonlinearities thus easier to train
  3. They always use pooling
  4. Compared to vanilla neural network and RNN, they are robust to gradient explosion and vanishing gradient.

A published solution is not available for this question yet.

Question 12 MCQ · 2.0 marks

Stacking multiple 3×3 convolutions (stride 1) increases receptive field because:
  1. Each layer composes local neighborhoods into larger effective context
  2. Each layer reduces spatial resolution
  3. It removes nonlinearities thereby increasing the effective context
  4. It forces global averaging thereby increasing the contribution of input pixels

A published solution is not available for this question yet.

Question 13 MCQ · 2.0 marks

In Faster R-CNN, the Region Proposal Network (RPN) outputs:
  1. Only class probabilities + pixel-wise masks
  2. Objectness scores + bounding box proposals
  3. Pixel-wise masks
  4. Optical flow

A published solution is not available for this question yet.

Question 14 MCQ · 2.0 marks

Which statement is correct ?
  1. Semantic segmentation separates individual object instances.
  2. Instance segmentation distinguishes different object instances of the same class.
  3. Both are identical tasks.
  4. Neither predicts labels per pixel.

A published solution is not available for this question yet.

Question 15 MCQ · 2.0 marks

Vanishing gradients in vanilla RNNs occur mainly due to:
  1. High deapth of the RNN layers
  2. Repeated multiplication by Jacobians across timesteps
  3. Too many convolution layers before RNN
  4. Use of softmax

A published solution is not available for this question yet.

Question 16 MSQ · 3.0 marks

Select all correct statements:
  1. Harris corner detector responds strongly where gradient changes in two orthogonal directions.
  2. Difference-of-Gaussians approximates Laplacian-of-Gaussian for blob detection.
  3. SIFT descriptors are computed from raw pixel intensities without gradients.
  4. Canny edge detector is a data driven method to calculate edges in an image.

A published solution is not available for this question yet.

Question 17 MSQ · 3.0 marks

Select all correct statements :
  1. Adam adapts per-parameter learning rates using estimates of first and second moments.
  2. Early stopping can act as a form of regularization.
  3. Dropout increases training accuracy by removing bad training data samples .
  4. Data augmentation typically improves generalization in vision tasks.

A published solution is not available for this question yet.

Question 18 MSQ · 3.0 marks

Select all correct statements:
  1. Self-attention computes pairwise interactions between tokens within a sequence.
  2. Positional encodings are needed because self-attention alone is permutation- invariant.
  3. Transformers must use recurrence to handle sequences.
  4. Multi-head attention allows the model to attend to different subspaces/relations.

A published solution is not available for this question yet.

Question 19 NAT · 4.0 marks

Box A: top-left (0,0), bottom-right (8,8). Box B: top-left (4,4), bottom-right (12,12). Compute IoU as a decimal.

    A published solution is not available for this question yet.

    Question 20 NAT · 4.0 marks

    For a single query, suppose similarity with the positive is s\(^{+}\)=2 and with two negatives are s1=1, s2=0. InfoNCE loss: L = −log( exp(s\(^{+}\)) / (exp(s\(^{+}\))+exp(s1)+exp(s2)) ). Using exp(2)=7.39, exp(1)=2.72, exp(0)=1, compute L.

      A published solution is not available for this question yet.

      Question 21 NAT · 4.0 marks

      Input image: 64×64. Convolution: kernel 7×7, stride 2, padding 3. Compute output spatial size.

        A published solution is not available for this question yet.

        Question 22 NAT · 4.0 marks

        A fully connected network maps input 100 → hidden 50 → output 10. Ignore bias. What is the total number of weights?

          A published solution is not available for this question yet.

          Question 23 NAT · 4.0 marks

          Input feature map has M=64 channels. You apply depthwise 3×3 followed by pointwise 1×1 to produce N=128 channels. Ignore bias. What is the Total parameters?

            A published solution is not available for this question yet.