da5013_2026T2_ET_FN.pdf
Deep Learning Practice · End Term · May 2026 FN
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 NAT · 1.0 marks
[[IMAGE:cf41183a839418fe_3_2]]
Based on the above data, answer the given subquestions.
Based on the actual output of [[IMAGE:cf41183a839418fe_4_3]] for a 128×128 input, how many elements are there
per sample when it is flattened (i.e., what should [[IMAGE:cf41183a839418fe_4_4]] evaluate to)?



A published solution is not available for this question yet.
Question 3 MCQ · 1.0 marks
[[IMAGE:cf41183a839418fe_3_2]]
Based on the above data, answer the given subquestions.
Which statement best explains why this code raises a [[IMAGE:cf41183a839418fe_4_5]] ?


The classifier's declared input size (12544) does not match the actual flattened
feature size produced by [[IMAGE:cf41183a839418fe_4_6]] .

[[IMAGE:cf41183a839418fe_4_7]] uses an invalid stride value for a 3×3 kernel.

[[IMAGE:cf41183a839418fe_4_8]] cannot be applied immediately after [[IMAGE:cf41183a839418fe_4_9]] .


The batch size of 8 is incompatible with [[IMAGE:cf41183a839418fe_4_10]] .

A published solution is not available for this question yet.
Question 4 NAT · 1.0 marks
[[IMAGE:cf41183a839418fe_3_2]]
Based on the above data, answer the given subquestions.
Using the actual flattened feature size from previous question number 2, how many learnable
parameters (weights + bias) would a fully connected layer with 10 outputs have?

A published solution is not available for this question yet.
Question 5 MSQ · 2.0 marks
[[IMAGE:cf41183a839418fe_3_2]]
Based on the above data, answer the given subquestions.
Which of the following statements about the (corrected) model are true?

The output of [[IMAGE:cf41183a839418fe_5_11]] (before [[IMAGE:cf41183a839418fe_5_12]] ) has spatial size 62×62.


[[IMAGE:cf41183a839418fe_5_13]] applied to a 62×62 input produces exactly 31×31, with no
dimension lost to rounding.

[[IMAGE:cf41183a839418fe_5_14]] 's stride of 2 also contributes to spatial downsampling in this network.

The number of channels in the final output of [[IMAGE:cf41183a839418fe_5_15]] is 32.

A published solution is not available for this question yet.
Question 6 NAT · 1.0 marks
[[IMAGE:cf41183a839418fe_5_16]]
Recall: [[IMAGE:cf41183a839418fe_5_17]] for transpose Convolution ).
The corresponding encoder to the above decoder downsampled an original 120×120 input via 3
stride-2 stages down to 15×15.
Based on the above data, answer the given subquestions.
What is the output height after [[IMAGE:cf41183a839418fe_6_18]] ? The width is the same; enter one integer only.



A published solution is not available for this question yet.
Question 7 NAT · 1.0 marks
[[IMAGE:cf41183a839418fe_5_16]]
Recall: [[IMAGE:cf41183a839418fe_5_17]] for transpose Convolution ).
The corresponding encoder to the above decoder downsampled an original 120×120 input via 3
stride-2 stages down to 15×15.
Based on the above data, answer the given subquestions.
What is the output height after [[IMAGE:cf41183a839418fe_6_19]] ? The width is the same; enter one integer only.



A published solution is not available for this question yet.
Question 8 MCQ · 2.0 marks
[[IMAGE:cf41183a839418fe_5_16]]
Recall: [[IMAGE:cf41183a839418fe_5_17]] for transpose Convolution ).
The corresponding encoder to the above decoder downsampled an original 120×120 input via 3
stride-2 stages down to 15×15.
Based on the above data, answer the given subquestions.
Which layer is responsible for the decoder failing to reach the target 120×120 output, and why?


[[IMAGE:cf41183a839418fe_6_20]] , because its kernel size is too small.

[[IMAGE:cf41183a839418fe_6_21]] , because it uses the wrong number of input channels.

[[IMAGE:cf41183a839418fe_7_22]] , because it uses [[IMAGE:cf41183a839418fe_7_23]] instead of [[IMAGE:cf41183a839418fe_7_24]] , so it fails to double the
spatial resolution.



None of the layers — the decoder already produces a 120×120 output.
A published solution is not available for this question yet.
Question 9 NAT · 1.0 marks
[[IMAGE:cf41183a839418fe_5_16]]
Recall: [[IMAGE:cf41183a839418fe_5_17]] for transpose Convolution ).
The corresponding encoder to the above decoder downsampled an original 120×120 input via 3
stride-2 stages down to 15×15.
Based on the above data, answer the given subquestions.
Suppose instead the decoder's input [[IMAGE:cf41183a839418fe_7_25]] had spatial size 30×30 (instead of 15×15).
After applying the correction identified in previous question, what would the final output height
be? The width is the same; enter one integer only.



A published solution is not available for this question yet.
Question 10 NAT · 1.0 marks
[[IMAGE:cf41183a839418fe_7_26]]
Consider an input of size **32×32×32** (32 channels), a standard
[[IMAGE:cf41183a839418fe_7_27]] , and the
[[IMAGE:cf41183a839418fe_8_28]] module above. Ignore bias
terms throughout, and assume spatial size is preserved (32×32 output) in both cases.
Based on the above data, answer the given subquestions.
How many learnable parameters does the standard
[[IMAGE:cf41183a839418fe_8_29]] have?




A published solution is not available for this question yet.
Question 11 NAT · 1.0 marks
[[IMAGE:cf41183a839418fe_7_26]]
Consider an input of size **32×32×32** (32 channels), a standard
[[IMAGE:cf41183a839418fe_7_27]] , and the
[[IMAGE:cf41183a839418fe_8_28]] module above. Ignore bias
terms throughout, and assume spatial size is preserved (32×32 output) in both cases.
Based on the above data, answer the given subquestions.
How many total learnable parameters does
[[IMAGE:cf41183a839418fe_8_30]] have (depthwise +
pointwise combined)?




A published solution is not available for this question yet.
Question 12 NAT · 1.0 marks
[[IMAGE:cf41183a839418fe_7_26]]
Consider an input of size **32×32×32** (32 channels), a standard
[[IMAGE:cf41183a839418fe_7_27]] , and the
[[IMAGE:cf41183a839418fe_8_28]] module above. Ignore bias
terms throughout, and assume spatial size is preserved (32×32 output) in both cases.
Based on the above data, answer the given subquestions.
What is the parameter compression ratio (standard ÷ depthwise-separable), rounded to 2 decimal
places?



A published solution is not available for this question yet.
Question 13 MSQ · 2.0 marks
[[IMAGE:cf41183a839418fe_7_26]]
Consider an input of size **32×32×32** (32 channels), a standard
[[IMAGE:cf41183a839418fe_7_27]] , and the
[[IMAGE:cf41183a839418fe_8_28]] module above. Ignore bias
terms throughout, and assume spatial size is preserved (32×32 output) in both cases.
Based on the above data, answer the given subquestions.
Which of the following statements about [[IMAGE:cf41183a839418fe_9_31]] are correct?




The depthwise layer must use [[IMAGE:cf41183a839418fe_9_32]] so that each input channel is
convolved with its own separate filter.

The pointwise layer uses a 1×1 kernel to mix information across channels after
the depthwise step.
Replacing a standard convolution with a depthwise-separable convolution
always improves model accuracy.
The compression ratio achieved depends on both the kernel size and the
number of output channels.
A published solution is not available for this question yet.
Question 14 MSQ · 1.0 marks
Assume the [[IMAGE:cf41183a839418fe_9_33]] architecture itself has already been corrected so that its classifier input size
matches the feature output. Focus only on the training/checkpoint/inference pipeline below.
[[IMAGE:cf41183a839418fe_10_34]]
Based on the above data, answer the given subquestions.
Which of the following statements correctly identify a real bug in the given code ?


[[IMAGE:cf41183a839418fe_11_35]] is never called, so gradients accumulate across
batches instead of being reset before each step.

[[IMAGE:cf41183a839418fe_11_36]] saves the full model object, but the reload code
calls [[IMAGE:cf41183a839418fe_11_37]] , which expects a state dict rather than a
model instance.


[[IMAGE:cf41183a839418fe_11_38]] is never moved to [[IMAGE:cf41183a839418fe_11_39]] , so [[IMAGE:cf41183a839418fe_11_40]] will raise a device-
mismatch error whenever [[IMAGE:cf41183a839418fe_11_41]] is a GPU.




[[IMAGE:cf41183a839418fe_11_42]] is the wrong loss function for a 10-class
classification problem.

A published solution is not available for this question yet.
Question 15 MCQ · 1.0 marks
Assume the [[IMAGE:cf41183a839418fe_9_33]] architecture itself has already been corrected so that its classifier input size
matches the feature output. Focus only on the training/checkpoint/inference pipeline below.
[[IMAGE:cf41183a839418fe_10_34]]
Based on the above data, answer the given subquestions.
Suppose the device-mismatch bug is fixed, but [[IMAGE:cf41183a839418fe_11_43]] is still never called before
inference. Given that [[IMAGE:cf41183a839418fe_11_44]] uses BatchNorm, what is the most likely consequence?




No effect — BatchNorm behaves identically in train and eval mode.
BatchNorm continues to use statistics from the current inference batch
instead of the stored running statistics, so predictions can differ from proper evaluation-mode
inference and can depend on batch composition.
The model automatically switches to eval mode whenever [[IMAGE:cf41183a839418fe_11_45]]
is used.

The prediction will be correct but computed twice as slowly.
A published solution is not available for this question yet.
Question 16 NAT · 1.0 marks
Assume the [[IMAGE:cf41183a839418fe_9_33]] architecture itself has already been corrected so that its classifier input size
matches the feature output. Focus only on the training/checkpoint/inference pipeline below.
[[IMAGE:cf41183a839418fe_10_34]]
Based on the above data, answer the given subquestions.
In the fully corrected pipeline, if [[IMAGE:cf41183a839418fe_11_46]] yields 20 batches per epoch over 5 epochs, how
many total times should [[IMAGE:cf41183a839418fe_11_47]] be called?




A published solution is not available for this question yet.
Question 17 NAT · 2.0 marks
A convolutional layer receives a feature map with **16 input channels** and applies **32 filters**, each
of spatial size **5 × 5**. Following the "generic level of CNN" view (a filter spans all input channels),
calculate the total number of **weights (excluding bias)** in this layer.
A published solution is not available for this question yet.
Question 18 NAT · 2.0 marks
A convolutional layer is applied to a **32 × 32** input using a **5 × 5** filter with **zero-padding of 2** and
**stride 2**. Using the standard convolution output-size formula, calculate the **output width (one**
**spatial dimension)**. The height is the same; enter one integer only.
A published solution is not available for this question yet.
Question 19 NAT · 2.0 marks
The final classification head applies a **softmax** over the pre-activation scores (logits). For a 3-class
problem, the logits produced for a particular image are **[3, 1, 0]**. Compute the softmax **probability**
**assigned to the first class** (round to 2 decimal places).
A published solution is not available for this question yet.
Question 20 NAT · 2.0 marks
Following the VGG design philosophy of replacing a single large-kernel convolution with a cascade
of small **3 × 3, stride-1** convolutions, how many such 3 × 3 layers must be **stacked** so that the
resulting stack has the same effective (one-dimensional) receptive field as a **single 11 × 11**
**convolution** (the kernel size used in AlexNet's first layer)?
A published solution is not available for this question yet.
Question 21 NAT · 2.0 marks
A YOLOv1-style detector divides the input image into an **S = 14** grid. Each grid cell predicts **B = 3**
bounding boxes (each with [[IMAGE:cf41183a839418fe_13_48]] ) and a single set of class scores over **C =**
**25** classes.
Compute the **total number of scalar values** in the final output tensor of the network.

A published solution is not available for this question yet.
Question 22 NAT · 2.0 marks
The decoder of a depth-estimation network upsamples feature maps using a **transposed**
**convolution** ("up-convolution") with **kernel size 4, stride 2, and padding 1**. If the input feature
map has spatial size **30 × 30**, what is the output height? The width is the same; enter one integer
only.
A published solution is not available for this question yet.
Question 23 NAT · 2.0 marks
In the original U-Net used for dense prediction (and adaptable to depth estimation), each encoder
stage applies **two successive 3×3 convolutions with no padding (stride 1)**, followed by a **2×2**
**max-pooling with stride 2**. If the input feature map to one such encoder stage has spatial size
**140 × 140**, what is the output height of the feature map **after both convolutions and the max-**
**pooling** of that stage? The width is the same; enter one integer only.
A published solution is not available for this question yet.
Question 24 NAT · 2.0 marks
The SRGAN discriminator is a stack of **8 convolutional layers**, all using **3 × 3 kernels with**
**padding 1**, and strides alternating as **s = 1, 2, 1, 2, 1, 2, 1, 2** across the eight layers (the stride-1
layers preserve spatial size; the stride-2 layers downsample).
If the input image is **96 × 96**, what is the output height of the feature map **after the last (8th)**
**convolutional layer**, just before the dense layers? The width is the same; enter one integer only.
A published solution is not available for this question yet.
Question 25 NAT · 2.0 marks
Efficient restoration networks (e.g., LaKDNet) replace standard convolutions with **depth-wise**
**separable convolutions** (a depth-wise convolution followed by a point-wise 1×1 convolution).
An input feature map of size **32 × 32 × 16** (16 channels) is processed by a depth-wise separable
convolution that uses **3 × 3** depth-wise kernels and produces **32 output channels**, with stride 1
and padding 1 (so the spatial size is preserved).
Calculate the **total number of multiplications** for the full depth-wise separable operation (depth-
wise + point-wise). Give the answer in millions (e.g., if 1,234,567 write 1.23).
A published solution is not available for this question yet.
Question 26 MSQ · 2.0 marks
Compared with the sigmoid activation, which of the following are **correct** reasons for preferring
the **ReLU** activation in deep CNNs?
For positive inputs, ReLU does not saturate, so it avoids the gradient-killing
that occurs in the flat tails of the sigmoid.
ReLU is preferred because it bounds every activation to the interval [0, 1],
preventing activations from growing large.
ReLU is cheaper to evaluate — a simple thresholding at zero — whereas
sigmoid/tanh require expensive exponentials.
The sigmoid's derivative peaks at only **0.25**, so multiplying many such small
factors across layers pushes gradients toward zero (vanishing gradient).
A published solution is not available for this question yet.
Question 27 MSQ · 2.0 marks
VGGNet replaces large convolution kernels with stacks of small 3 × 3 convolutions. Which of the
following statements correctly describe the **advantages** of this design choice?
A stack of three 3 × 3 conv layers (stride 1) has the same effective receptive
field as one 7 × 7 conv layer.
For C channels per layer, three stacked 3 × 3 layers use fewer parameters than
a single 7 × 7 layer.
A single 3 × 3 convolution by itself already spans the same receptive field as a
7 × 7 convolution, so stacking is unnecessary.
Inserting a ReLU non-linearity between successive 3 × 3 layers makes the
composite function more discriminative than a single linear projection over the same region.
A published solution is not available for this question yet.
Question 28 MSQ · 2.0 marks
Which of the following statements correctly describe the **Region Proposal Network (RPN)** in
Faster R-CNN?
The RPN is fully convolutional and **shares** its convolutional feature maps with
the downstream detection network.
The RPN relies on an **external Selective Search** module to generate its region
proposals.
At each sliding-window location, the RPN predicts objectness scores and box
refinements **relative to k predefined anchor boxes**.
Anchors of multiple **scales and aspect ratios** allow the RPN to propose
regions for objects of very different sizes and shapes.
A published solution is not available for this question yet.
Question 29 MSQ · 2.0 marks
Consider the multi-task loss used to train the original **YOLOv1** model. Which of the following
statements about it are correct?
The loss predicts the **square roots** of width and height so that a fixed absolute
error is penalized **more** for small boxes than for large boxes.
The hyperparameters are set to **λ_coord = 5** (upweighting localization error)
and **λ_noobj = 0.5** (downweighting confidence error for cells with no object).
The **classification** term for a cell is only penalized when an object is actually
present in that grid cell.
For the predictor "responsible" for an object, the target **confidence** value is
fixed at 1, independent of the box's IoU with the ground truth.
A published solution is not available for this question yet.
Question 30 MSQ · 2.0 marks
Regarding the **channel-attention module in CBAM**, which of the following are correct?
It exploits the **inter-channel** relationship of features, learning "what" is
meaningful by re-weighting feature channels.
It produces a **spatial** attention map of size H × W × 1 that highlights where to
focus, discarding all channel information.
The spatial dimensions of the input feature map are squeezed using **both**
**average-pooling and max-pooling**, producing two channel descriptors.
The pooled descriptors are passed through a **shared MLP**, combined, and a
sigmoid produces the channel attention weights that rescale the input feature map.
A published solution is not available for this question yet.
Question 31 MSQ · 2.0 marks
SRGAN performs photo-realistic super-resolution using a GAN. Which of the following statements
are correct?
Its **perceptual loss** combines a **content loss** and an **adversarial loss**, rather
than relying on pixel-wise MSE alone.
The **content loss** is computed on **feature maps of a pre-trained VGG**
**network** (perceptual similarity), not purely in pixel space.
Purely MSE-optimized solutions tend to be **overly smooth** because they
approximate a pixel-wise average of many plausible HR solutions, losing high-frequency texture.
The adversarial term makes SRGAN maximize PSNR, which is why it always
achieves a **higher PSNR** than the MSE-based SRResNet.
A published solution is not available for this question yet.
Question 32 MCQ · 2.0 marks
In the **multi-scale deep network** for single-image depth estimation (Eigen et al.), the architecture
is split into a **coarse network** and a **fine network**. Which statement best explains their respective
roles?
The coarse network uses a large receptive field (via deeper layers and fully-
connected layers) to predict the global scene depth, while the fine network refines this prediction
locally using input-image details.
The fine network predicts the global scene structure, and the coarse network
only sharpens edges.
Both networks are architecturally identical and their outputs are simply
averaged (ensembled).
The coarse network processes each pixel independently and therefore ignores
global scene context.
A published solution is not available for this question yet.