da5006_2026T1_ET_FN.pdf
Deep Learning for Computer Vision · End Term · Jan 2026 FN
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 MCQ · 2.0 marks
A standard RGB image of size 256×256 is stored as 8-bit unsigned integers per channel. What is
the total memory (in bytes) to store the raw image (no compression)?
256×256 = 65,536 bytes
256×256×2 = 131,072 bytes
256×256×3 = 196,608 bytes
256×256×3×8 = 1,572,864 bytes
A published solution is not available for this question yet.
Question 3 MCQ · 2.0 marks
Which statement is correct?
Convolution equals correlation for all filters.
Convolution flips the kernel spatially; correlation does not.
Correlation flips the kernel; convolution does not.
Both operations always preserve energy.
A published solution is not available for this question yet.
Question 4 MCQ · 2.0 marks
You apply a 5×5 filter with stride 1 and **no padding** on a 32×32 image. Output size is:
32×32
28×28
27×27
26×26
A published solution is not available for this question yet.
Question 5 MCQ · 2.0 marks
The Sobel operator primarily estimates:
[[IMAGE:c3314bd6a89b0d7f_3_2]]

[[IMAGE:c3314bd6a89b0d7f_3_3]]

[[IMAGE:c3314bd6a89b0d7f_3_4]]

[[IMAGE:c3314bd6a89b0d7f_3_5]]

A published solution is not available for this question yet.
Question 6 MCQ · 2.0 marks
SIFT is designed to be robust mainly to:
Scale and rotation changes
Only translation (shifts)
Only illumination changes
Only affine viewpoint changes with no residual error
A published solution is not available for this question yet.
Question 7 MCQ · 2.0 marks
In classical scale-space theory, increasing σ in Gaussian smoothing generally:
Enhances high-frequency details
Suppresses high-frequency details and removes fine structures
Has no effect on edges
Increases noise
A published solution is not available for this question yet.
Question 8 MCQ · 2.0 marks
Backpropagation is best described as:
A method to avoid overfitting
Efficient computation of gradients using the chain rule through computation
graphs
A method that works only for linear models
An optimizer like Adam
A published solution is not available for this question yet.
Question 9 MCQ · 2.0 marks
Which is a key effect of momentum?
Makes gradients unbiased
Accumulates a velocity term
Gurantees to converge unlike vanilla Gradient Descent
Eliminates need for setting learning rate by adpatly calculaitng it on the fly.
A published solution is not available for this question yet.
Question 10 MCQ · 2.0 marks
L2 regularization on weights most directly encourages:
Sparse weights with many exact zeros
Smaller-magnitude weights
Larger-magnitude weights
Larger learning rates
A published solution is not available for this question yet.
Question 11 MCQ · 2.0 marks
A key reason CNNs are parameter-efficient compared to fully connected nets for images is:
Weight sharing and local receptive fields
They never use nonlinearities thus easier to train
They always use pooling
Compared to vanilla neural network and RNN, they are robust to gradient
explosion and vanishing gradient.
A published solution is not available for this question yet.
Question 12 MCQ · 2.0 marks
Stacking multiple 3×3 convolutions (stride 1) increases receptive field because:
Each layer composes local neighborhoods into larger effective context
Each layer reduces spatial resolution
It removes nonlinearities thereby increasing the effective context
It forces global averaging thereby increasing the contribution of input pixels
A published solution is not available for this question yet.
Question 13 MCQ · 2.0 marks
In Faster R-CNN, the Region Proposal Network (RPN) outputs:
Only class probabilities + pixel-wise masks
Objectness scores + bounding box proposals
Pixel-wise masks
Optical flow
A published solution is not available for this question yet.
Question 14 MCQ · 2.0 marks
Which statement is correct ?
Semantic segmentation separates individual object instances.
Instance segmentation distinguishes different object instances of the same
class.
Both are identical tasks.
Neither predicts labels per pixel.
A published solution is not available for this question yet.
Question 15 MCQ · 2.0 marks
Vanishing gradients in vanilla RNNs occur mainly due to:
High deapth of the RNN layers
Repeated multiplication by Jacobians across timesteps
Too many convolution layers before RNN
Use of softmax
A published solution is not available for this question yet.
Question 16 MSQ · 3.0 marks
Select all correct statements:
Harris corner detector responds strongly where gradient changes in two
orthogonal directions.
Difference-of-Gaussians approximates Laplacian-of-Gaussian for blob
detection.
SIFT descriptors are computed from raw pixel intensities without gradients.
Canny edge detector is a data driven method to calculate edges in an image.
A published solution is not available for this question yet.
Question 17 MSQ · 3.0 marks
Select all correct statements :
Adam adapts per-parameter learning rates using estimates of first and second
moments.
Early stopping can act as a form of regularization.
Dropout increases training accuracy by removing bad training data samples .
Data augmentation typically improves generalization in vision tasks.
A published solution is not available for this question yet.
Question 18 MSQ · 3.0 marks
Select all correct statements:
Self-attention computes pairwise interactions between tokens within a
sequence.
Positional encodings are needed because self-attention alone is permutation-
invariant.
Transformers must use recurrence to handle sequences.
Multi-head attention allows the model to attend to different
subspaces/relations.
A published solution is not available for this question yet.
Question 19 NAT · 4.0 marks
Box A: top-left (0,0), bottom-right (8,8).
Box B: top-left (4,4), bottom-right (12,12).
Compute IoU as a decimal.
A published solution is not available for this question yet.
Question 20 NAT · 4.0 marks
For a single query, suppose similarity with the positive is s\(^{+}\)=2 and with two negatives are s1=1,
s2=0.
InfoNCE loss: L = −log( exp(s\(^{+}\)) / (exp(s\(^{+}\))+exp(s1)+exp(s2)) ).
Using exp(2)=7.39, exp(1)=2.72, exp(0)=1, compute L.
A published solution is not available for this question yet.
Question 21 NAT · 4.0 marks
Input image: 64×64. Convolution: kernel 7×7, stride 2, padding 3. Compute output spatial size.
A published solution is not available for this question yet.
Question 22 NAT · 4.0 marks
A fully connected network maps input 100 → hidden 50 → output 10. Ignore bias. What is the total
number of weights?
A published solution is not available for this question yet.
Question 23 NAT · 4.0 marks
Input feature map has M=64 channels. You apply depthwise 3×3 followed by pointwise 1×1 to
produce N=128 channels. Ignore bias. What is the Total parameters?
A published solution is not available for this question yet.