da5006_2026T1_Q2_NA.pdf
Deep Learning for Computer Vision · Quiz 2 · Jan 2026
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 MCQ · 1.0 marks
Why do very deep plain CNNs (without skip connections) sometimes show **higher training error**
than shallower CNNs?
The increased capacity leads to severe overfitting on the training set before
convergence can be reached.
Optimization becomes harder; skip connections make learning identity
mappings easier and improve gradient flow.
The receptive field becomes excessively large, causing the network to lose
fine-grained spatial information.
Saturated activations cause exploding gradients that cannot be mitigated by
standard initialization techniques.
A published solution is not available for this question yet.
Question 3 MCQ · 1.0 marks
In a standard ResNet-50 bottleneck block, the three convolutions are typically:
3×3, 1×1, 3×3
1×1, 3×3, 1×1
1×1, 5×5, 1×1
3×3, 3×3, 3×3
A published solution is not available for this question yet.
Question 4 MCQ · 1.0 marks
In Inception/GoogLeNet modules, the main purpose of 1×1 convolutions is to:
Increase the receptive field of subsequent convolutional layers without
downsampling.
Reduce channel dimensionality (bottleneck) and add non-linearity
Act as a differentiable substitute for max-pooling operations to preserve exact
spatial hierarchies.
Project the feature maps into a higher-dimensional space to separate
entangled features.
A published solution is not available for this question yet.
Question 5 MCQ · 1.0 marks
MobileNetV1 reduces computation primarily by using:
Grouped convolutions followed by spatial pyramid pooling
Depthwise separable convolutions (depthwise + pointwise)
Asymmetric convolutions (e.g., factorizing a 3×3 into 3×1 followed by 1×3)
Low-rank matrix factorization of dense classification layers
A published solution is not available for this question yet.
Question 6 MCQ · 1.0 marks
EfficientNet’s key scaling idea is:
Scaling the network depth logarithmically while linearly increasing the input
resolution
Utilizing neural architecture search (NAS) to independently optimize depth,
width, and resolution for each block
Compound scaling of depth, width, and resolution using a single coefficient
Adjusting the channel multiplier dynamically based on the target device's
hardware constraints
A published solution is not available for this question yet.
Question 7 MCQ · 1.0 marks
You have a pretrained CNN backbone and only **500** labeled images for a new task. A strong first
baseline is:
Fine-tune the entire network immediately using a very small learning rate to
preserve pre-trained features
Freeze backbone, train a new classifier head; optionally unfreeze later with
smaller LR
Train a completely new architecture from scratch using heavy data
augmentation to compensate for the lack of data
Unfreeze the backbone from the bottom layers upwards, gradually increasing
the learning rate
A published solution is not available for this question yet.
Question 8 MCQ · 1.0 marks
When fine-tuning a model, what is a common way to prevent the new training from "overwriting"
or destroying the useful features already learned during pre-training?
Applying a high dropout rate exclusively to the newly initialized classification
head
Use differential learning rates (small LR for the pre-trained backbone, larger
LR for the new head)
Re-initializing the weights of the last convolutional block before training
Freezing the classification head and only fine-tuning the backbone features
A published solution is not available for this question yet.
Question 9 MCQ · 1.0 marks
Which method produces a heatmap by weighting convolutional feature maps using gradients of a
target class score?
Guided Backpropagation
Grad-CAM
SmoothGrad
Integrated Gradients
A published solution is not available for this question yet.
Question 10 MCQ · 1.0 marks
Which detector is a classic **two-stage** detector with an RPN + ROI feature extraction?
SSD
YOLO
Faster R-CNN
DETR
A published solution is not available for this question yet.
Question 11 MCQ · 1.0 marks
Which statement is most accurate?
Two-stage detectors optimize a single joint loss function for classification and
localization, whereas single-stage uses decoupled losses.
Single-stage detectors rely on selective search for region generation, whereas
two-stage detectors use a dedicated sub-network.
Two-stage detectors typically generate proposals then classify/refine; single-
stage predicts boxes/classes densely.
Two-stage detectors process the entire image in one pass, while single-stage
detectors evaluate crops sequentially.
A published solution is not available for this question yet.
Question 12 MCQ · 1.0 marks
RetinaNet is notable primarily because it introduced:
A novel feature pyramid network (FPN) architecture that eliminates the need
for anchor boxes
Focal loss to address class imbalance
The concept of Region of Interest (RoI) pooling to align features extracted
from different scales
A dynamic routing mechanism that directly assigns bounding box predictions
to ground truth objects
A published solution is not available for this question yet.
Question 13 MCQ · 1.0 marks
A model that predicts a class label for **every pixel** without separating object instances is:
Panoptic segmentation
Instance segmentation
Semantic segmentation
Image matting
A published solution is not available for this question yet.
Question 14 MCQ · 1.0 marks
Vanishing gradients in vanilla RNNs are largely caused by:
The continuous addition of bias terms at each timestep which suppresses the
gradient signal
Repeated multiplication by Jacobians whose spectral norm is often < 1
The use of unbounded activation functions like ReLU that cause activations to
decay over time
The inability of the standard cross-entropy loss function to distinguish
between short-term and long-term dependencies
A published solution is not available for this question yet.
Question 15 MCQ · 1.0 marks
Which statement is true?
Transformers process sequences sequentially but utilize parallelized loss
computation to speed up training.
The attention mechanism strictly limits the context window size, whereas
RNNs have a theoretically infinite context window in practice.
Self-attention enables direct interactions between any pair of tokens within a
layer.
Transformers inherently encode positional information in their feed-forward
weights, eliminating the need for explicit sequence tracking.
A published solution is not available for this question yet.
Question 16 MSQ · 1.0 marks
Which statements can be true in practice when fine-tuning with small batch sizes?
Freezing BatchNorm running statistics can improve stability.
If the new dataset distribution differs, re-estimating BN stats may help.
Setting BatchNorm to training mode with a batch size of 1 provides an
unbiased estimate of the population variance.
Calling .eval() on a BatchNorm layer disables the application of the learned
affine parameters (γ, β).
A published solution is not available for this question yet.
Question 17 MSQ · 1.0 marks
Which statements are true?
Smooth L1 (Huber) is often used for box regression due to robustness to
outliers.
Focal loss primarily addresses foreground/background (class) imbalance.
Standard Cross-Entropy loss scales the loss of easy examples to zero more
aggressively than Focal Loss.
In two-stage detectors, the RPN has its own objectness + box regression
losses.
A published solution is not available for this question yet.
Question 18 MSQ · 1.0 marks
Which statements are true ?
Soft attention is differentiable and can be trained with backpropagation.
Hard attention often needs sampling + REINFORCE (or similar) due to non-
differentiability.
Soft attention significantly reduces the computational complexity of the
forward pass compared to Hard attention.
Cross-attention can align decoder queries to encoder keys/values (e.g., words
attending to image regions).
A published solution is not available for this question yet.
Question 19 NAT · 1.0 marks
Input feature map: 14×14×128, output: 14×14×256, kernel: 3×3, stride 1, same padding, ignore
bias.
Compute **number of parameters** for a standard convolution layer.
A published solution is not available for this question yet.
Question 20 NAT · 1.0 marks
Using the same setup (Input feature map: 14×14×128, output: 14×14×256, kernel: 3×3, stride 1,
same padding, ignore bias), compute total **MACs operation** (multiply-accumulates). A single MAC
operation multiplies two numbers and adds the result to an accumulator.
A published solution is not available for this question yet.
Question 21 NAT · 1.0 marks
A ResNet bottleneck block takes 256 channels in and uses:
1×1 conv to 256 channels → 3×3 conv at 256 → 1×1 conv to 1024 channels. Ignore bias.
Compute total parameters in these **three convolutions** (exclude projection shortcut).
A published solution is not available for this question yet.
Question 22 NAT · 1.0 marks
A detector uses three pyramid levels: P3 = 80×80, P4 = 40×40, P5 = 20×20.
At each location, it uses 9 anchors (3 scales × 3 aspect ratios).
Compute total anchors per image.
A published solution is not available for this question yet.
Question 23 NAT · 1.0 marks
Box A: top-left (2,2), bottom-right (10,12)
Box B: top-left (6,5), bottom-right (14,15)
Compute IoU rounded to 3 decimals.
A published solution is not available for this question yet.
Question 24 NAT · 1.0 marks
Let query q = [1,2]\(^{T}\) and keys:
k1 = [1,0]\(^{T}\), k2 = [0,1]\(^{T}\), k3 = [1,1]\(^{T}\).
Scores si = q\(^{T}\) ki. Use softmax over scores.
Use exp(1)=2.72, exp(2)=7.39, exp(3)=20.09.
What is α3 (the attention weight for k3)? Provide a value in a reasonable range.
A published solution is not available for this question yet.