MauryaHub PYQ Practice

da5006_2026T2_Q2_NA.pdf

Deep Learning for Computer Vision · Quiz 2 · May 2026

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 2 MCQ · 1.0 marks

Why do very deep plain CNNs (without skip connections) sometimes show **higher training error** than shallower CNNs?
  1. They always overfit because parameter count increases.
  2. Optimization becomes harder; skip connections make learning identity mappings easier and improve gradient flow.
  3. BatchNorm guarantees this never happens.
  4. Using sigmoid activations instead of ReLU always fixes it.

A published solution is not available for this question yet.

Question 3 MCQ · 1.0 marks

MobileNetV1 reduces computation primarily by using:
  1. Dilated convolutions everywhere
  2. Depthwise separable convolutions (depthwise + pointwise)
  3. ROI pooling
  4. Fully-connected layers instead of convolutions

A published solution is not available for this question yet.

Question 4 MCQ · 1.0 marks

EfficientNet’s key scaling idea is:
  1. Only scale depth while keeping width and resolution fixed
  2. Only scale resolution while keeping depth and width fixed
  3. Compound scaling of depth, width, and resolution using a single coefficient
  4. Scale width but reduce depth to keep parameters constant

A published solution is not available for this question yet.

Question 5 MCQ · 1.0 marks

Which technique most directly helps reduce catastrophic forgetting early in fine-tuning?
  1. Use very high weight decay on the backbone
  2. Use differential learning rates (small LR for backbone, larger LR for head)
  3. Replace ReLU with sigmoid
  4. Remove skip connections

A published solution is not available for this question yet.

Question 6 MCQ · 1.0 marks

Which method produces a heatmap by weighting convolutional feature maps using gradients of a target class score?
  1. PCA
  2. Grad-CAM
  3. Histogram equalization
  4. t-SNE

A published solution is not available for this question yet.

Question 7 MCQ · 1.0 marks

Which detector is a classic **two-stage** detector with an RPN + ROI feature extraction?
  1. SSD
  2. YOLO
  3. Faster R-CNN
  4. RetinaNet

A published solution is not available for this question yet.

Question 8 MCQ · 1.0 marks

RetinaNet is notable primarily because it introduced:
  1. ROI Align
  2. Focal loss to address class imbalance
  3. Depthwise separable convolutions
  4. Self-attention

A published solution is not available for this question yet.

Question 9 MCQ · 1.0 marks

A model that predicts a class label for **every pixel** without separating object instances is:
  1. Object detection
  2. Instance segmentation
  3. Semantic segmentation
  4. Captioning

A published solution is not available for this question yet.

Question 10 MCQ · 1.0 marks

Which statement is true?
  1. Transformers rely on recurrence to model sequences.
  2. Transformers cannot handle variable-length sequences.
  3. Self-attention enables direct interactions between any pair of tokens within a layer.
  4. Transformers require optical flow for vision tasks.

A published solution is not available for this question yet.

Question 11 MCQ · 2.0 marks

In a standard ResNet-50 bottleneck block, the three convolutions are typically:
  1. 3×3, 3×3, 3×3
  2. 1×1, 3×3, 1×1
  3. 5×5, 3×3, 1×1
  4. 1×1, 1×1, 3×3

A published solution is not available for this question yet.

Question 12 MCQ · 2.0 marks

In Inception/GoogLeNet modules, the main purpose of 1×1 convolutions is to:
  1. Increase spatial resolution
  2. Reduce channel dimensionality (bottleneck) and add non-linearity
  3. Replace pooling
  4. Implement residual learning

A published solution is not available for this question yet.

Question 13 MCQ · 2.0 marks

Which statement is most accurate?
  1. Two-stage detectors do not use CNN backbones.
  2. Single-stage detectors cannot regress bounding boxes.
  3. Two-stage detectors typically generate proposals then classify/refine
  4. Single-stage detectors require ROI pooling.

A published solution is not available for this question yet.

Question 14 MCQ · 2.0 marks

Vanishing gradients in vanilla RNNs are largely caused by:
  1. Too much data augmentation
  2. Repeated multiplication by Jacobians whose spectral norm is often < 1
  3. Using convolutions before the RNN
  4. Using attention layers

A published solution is not available for this question yet.

Question 15 MSQ · 3.0 marks

Which statements can be true in practice when fine-tuning with small batch sizes?
  1. Freezing BatchNorm running statistics can improve stability.
  2. If the new dataset distribution differs, re-estimating BN stats may help.
  3. BN always improves performance even with batch size 1.
  4. Setting BN to eval disables learned affine parameters (γ, β).

A published solution is not available for this question yet.

Question 16 MSQ · 3.0 marks

Which statements are true ?
  1. Smooth L1 (Huber) is often used for box regression due to robustness to outliers.
  2. Focal loss primarily addresses foreground/background (class) imbalance.
  3. IoU loss cannot be used for box regression.
  4. In two-stage detectors, the RPN has its own objectness + box regression losses.

A published solution is not available for this question yet.

Question 17 MSQ · 3.0 marks

Which statements are true?
  1. Soft attention is differentiable and can be trained with backpropagation.
  2. Hard attention often needs sampling + REINFORCE (or similar) due to non- differentiability.
  3. Hard attention always has higher test-time compute than soft attention.
  4. Cross-attention can align decoder queries to encoder keys/values (e.g., words attending to image regions).

A published solution is not available for this question yet.

Question 18 NAT · 4.0 marks

Input feature map: 14×14×128, output: 14×14×256, kernel: 3×3, stride 1, same padding, ignore bias. Compute **number of parameters** for a standard convolution layer.

    A published solution is not available for this question yet.

    Question 19 NAT · 4.0 marks

    Input feature map: 14×14×128, output: 14×14×256, kernel: 3×3, stride 1, same padding, ignore bias, compute total **MACs** (multiply-accumulates).

      A published solution is not available for this question yet.

      Question 20 NAT · 4.0 marks

      A ResNet bottleneck block takes 256 channels in and uses: 1×1 conv to 256 channels → 3×3 conv at 256 → 1×1 conv to 1024 channels. Ignore bias. Compute total parameters in these **three convolutions** (exclude projection shortcut).

        A published solution is not available for this question yet.

        Question 21 NAT · 4.0 marks

        A detector uses three pyramid levels: P3 = 80×80, P4 = 40×40, P5 = 20×20. At each location, it uses 9 anchors (3 scales × 3 aspect ratios). Compute total anchors per image.

          A published solution is not available for this question yet.

          Question 22 NAT · 4.0 marks

          Box A: top-left (2,2), bottom-right (10,12) Box B: top-left (6,5), bottom-right (14,15) Compute IoU rounded to 3 decimals.

            A published solution is not available for this question yet.

            Question 23 NAT · 4.0 marks

            Let query q = [1,2]^T and keys:k1 = [1,0]^T, k2 = [0,1]^T, k3 = [1,1]^T. Scores si = q^T ki. Use softmax over scores. Use exp(1)=2.72, exp(2)=7.39, exp(3)=20.09. What is α3 (the attention weight for k3)? Provide a value in a reasonable range.

              A published solution is not available for this question yet.