da5006_2026T2_Q2_NA.pdf
Deep Learning for Computer Vision · Quiz 2 · May 2026
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 MCQ · 1.0 marks
Why do very deep plain CNNs (without skip connections) sometimes show **higher training error**
than shallower CNNs?
They always overfit because parameter count increases.
Optimization becomes harder; skip connections make learning identity
mappings easier and improve gradient flow.
BatchNorm guarantees this never happens.
Using sigmoid activations instead of ReLU always fixes it.
A published solution is not available for this question yet.
Question 3 MCQ · 1.0 marks
MobileNetV1 reduces computation primarily by using:
Dilated convolutions everywhere
Depthwise separable convolutions (depthwise + pointwise)
ROI pooling
Fully-connected layers instead of convolutions
A published solution is not available for this question yet.
Question 4 MCQ · 1.0 marks
EfficientNet’s key scaling idea is:
Only scale depth while keeping width and resolution fixed
Only scale resolution while keeping depth and width fixed
Compound scaling of depth, width, and resolution using a single coefficient
Scale width but reduce depth to keep parameters constant
A published solution is not available for this question yet.
Question 5 MCQ · 1.0 marks
Which technique most directly helps reduce catastrophic forgetting early in fine-tuning?
Use very high weight decay on the backbone
Use differential learning rates (small LR for backbone, larger LR for head)
Replace ReLU with sigmoid
Remove skip connections
A published solution is not available for this question yet.
Question 6 MCQ · 1.0 marks
Which method produces a heatmap by weighting convolutional feature maps using gradients of a
target class score?
PCA
Grad-CAM
Histogram equalization
t-SNE
A published solution is not available for this question yet.
Question 7 MCQ · 1.0 marks
Which detector is a classic **two-stage** detector with an RPN + ROI feature extraction?
SSD
YOLO
Faster R-CNN
RetinaNet
A published solution is not available for this question yet.
Question 8 MCQ · 1.0 marks
RetinaNet is notable primarily because it introduced:
ROI Align
Focal loss to address class imbalance
Depthwise separable convolutions
Self-attention
A published solution is not available for this question yet.
Question 9 MCQ · 1.0 marks
A model that predicts a class label for **every pixel** without separating object instances is:
Object detection
Instance segmentation
Semantic segmentation
Captioning
A published solution is not available for this question yet.
Question 10 MCQ · 1.0 marks
Which statement is true?
Transformers rely on recurrence to model sequences.
Transformers cannot handle variable-length sequences.
Self-attention enables direct interactions between any pair of tokens within a
layer.
Transformers require optical flow for vision tasks.
A published solution is not available for this question yet.
Question 11 MCQ · 2.0 marks
In a standard ResNet-50 bottleneck block, the three convolutions are typically:
3×3, 3×3, 3×3
1×1, 3×3, 1×1
5×5, 3×3, 1×1
1×1, 1×1, 3×3
A published solution is not available for this question yet.
Question 12 MCQ · 2.0 marks
In Inception/GoogLeNet modules, the main purpose of 1×1 convolutions is to:
Increase spatial resolution
Reduce channel dimensionality (bottleneck) and add non-linearity
Replace pooling
Implement residual learning
A published solution is not available for this question yet.
Question 13 MCQ · 2.0 marks
Which statement is most accurate?
Two-stage detectors do not use CNN backbones.
Single-stage detectors cannot regress bounding boxes.
Two-stage detectors typically generate proposals then classify/refine
Single-stage detectors require ROI pooling.
A published solution is not available for this question yet.
Question 14 MCQ · 2.0 marks
Vanishing gradients in vanilla RNNs are largely caused by:
Too much data augmentation
Repeated multiplication by Jacobians whose spectral norm is often < 1
Using convolutions before the RNN
Using attention layers
A published solution is not available for this question yet.
Question 15 MSQ · 3.0 marks
Which statements can be true in practice when fine-tuning with small batch sizes?
Freezing BatchNorm running statistics can improve stability.
If the new dataset distribution differs, re-estimating BN stats may help.
BN always improves performance even with batch size 1.
Setting BN to eval disables learned affine parameters (γ, β).
A published solution is not available for this question yet.
Question 16 MSQ · 3.0 marks
Which statements are true ?
Smooth L1 (Huber) is often used for box regression due to robustness to
outliers.
Focal loss primarily addresses foreground/background (class) imbalance.
IoU loss cannot be used for box regression.
In two-stage detectors, the RPN has its own objectness + box regression
losses.
A published solution is not available for this question yet.
Question 17 MSQ · 3.0 marks
Which statements are true?
Soft attention is differentiable and can be trained with backpropagation.
Hard attention often needs sampling + REINFORCE (or similar) due to non-
differentiability.
Hard attention always has higher test-time compute than soft attention.
Cross-attention can align decoder queries to encoder keys/values (e.g., words
attending to image regions).
A published solution is not available for this question yet.
Question 18 NAT · 4.0 marks
Input feature map: 14×14×128, output: 14×14×256, kernel: 3×3, stride 1, same padding, ignore
bias.
Compute **number of parameters** for a standard convolution layer.
A published solution is not available for this question yet.
Question 19 NAT · 4.0 marks
Input feature map: 14×14×128, output: 14×14×256, kernel: 3×3, stride 1, same padding, ignore
bias, compute total **MACs** (multiply-accumulates).
A published solution is not available for this question yet.
Question 20 NAT · 4.0 marks
A ResNet bottleneck block takes 256 channels in and uses:
1×1 conv to 256 channels → 3×3 conv at 256 → 1×1 conv to 1024 channels. Ignore bias.
Compute total parameters in these **three convolutions** (exclude projection shortcut).
A published solution is not available for this question yet.
Question 21 NAT · 4.0 marks
A detector uses three pyramid levels: P3 = 80×80, P4 = 40×40, P5 = 20×20.
At each location, it uses 9 anchors (3 scales × 3 aspect ratios).
Compute total anchors per image.
A published solution is not available for this question yet.
Question 22 NAT · 4.0 marks
Box A: top-left (2,2), bottom-right (10,12) Box B: top-left (6,5), bottom-right (14,15) Compute IoU
rounded to 3 decimals.
A published solution is not available for this question yet.
Question 23 NAT · 4.0 marks
Let query q = [1,2]^T and keys:k1 = [1,0]^T, k2 = [0,1]^T, k3 = [1,1]^T. Scores si = q^T ki. Use
softmax over scores. Use exp(1)=2.72, exp(2)=7.39, exp(3)=20.09.
What is α3 (the attention weight for k3)? Provide a value in a reasonable range.
A published solution is not available for this question yet.