da5013_2025T3_ET_FN.pdf
Deep Learning Practice · End Term · Sep 2025 FN
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 NAT · 1.0 marks
Read the data below and answer the given subquestions regarding Object Detection Evaluation.
An object detector is evaluated on a single class ”Car”. The dataset has 10 ground truth ”Car”
objects. The model makes 10 detections. Based on an IoU threshold of 0.5:
• 6 detections are True Positives (TP).
• 4 detections are False Positives (FP).
• (Consequently, there are 4 False Negatives (FN), as 10 ground truths - 6 found = 4 missed).
Calculate the Precision of the model. (Answer as a decimal between 0 and 1).
A published solution is not available for this question yet.
Question 3 NAT · 2.0 marks
Read the data below and answer the given subquestions regarding Object Detection Evaluation.
An object detector is evaluated on a single class ”Car”. The dataset has 10 ground truth ”Car”
objects. The model makes 10 detections. Based on an IoU threshold of 0.5:
• 6 detections are True Positives (TP).
• 4 detections are False Positives (FP).
• (Consequently, there are 4 False Negatives (FN), as 10 ground truths - 6 found = 4 missed).
Calculate the Recall of the model. (Answer as a decimal between 0 and 1).
A published solution is not available for this question yet.
Question 4 NAT · 2.0 marks
Read the data below and answer the given subquestions regarding Object Detection Evaluation.
An object detector is evaluated on a single class ”Car”. The dataset has 10 ground truth ”Car”
objects. The model makes 10 detections. Based on an IoU threshold of 0.5:
• 6 detections are True Positives (TP).
• 4 detections are False Positives (FP).
• (Consequently, there are 4 False Negatives (FN), as 10 ground truths - 6 found = 4 missed).
If the IoU threshold is increased to 0.9, one of the previous True Positives becomes a False
Positive. What is the new Precision? (Round to 2 decimal places).
A published solution is not available for this question yet.
Question 5 NAT · 5.0 marks
An input feature map has dimensions 28 × 28 × 192. We apply a 1 × 1 convolution layer with 64
filters. What is the total number of operations (multiplications) required? (Give answer in millions,
e.g., if 1,000,000, write 1).
A published solution is not available for this question yet.
Question 6 NAT · 5.0 marks
A stereo camera system has a baseline B = 0.5 meters and a focal length f = 100 pixels.
If the disparity d for a specific point is calculated to be 10 pixels, what is the depth Z of that point
in meters?
A published solution is not available for this question yet.
Question 7 NAT · 5.0 marks
In the original YOLO paper, the image is divided into a 7 × 7 grid (S = 7). The model predicts 2
bounding boxes per cell (B = 2) and detects 20 classes (C = 20). What is the depth (number of
channels) of the final output tensor?
A published solution is not available for this question yet.
Question 8 NAT · 5.0 marks
In AlexNet, the first convolutional layer uses filters of size 11 × 11 × 3. If there are 96 such filters,
calculate the total number of weights (excluding bias) in this layer.
A published solution is not available for this question yet.
Question 9 MSQ · 5.0 marks
In Unsupervised Monocular Depth Estimation using Left-Right Consistency :
The model is trained using stereo image pairs but requires only a single image
at test time.
It uses the predicted disparity to warp the left image to match the right image
(or vice versa).
The loss function is based on the photometric reconstruction error between
the warped image and the target image.
It requires ground truth depth maps captured by LiDAR for training.
A published solution is not available for this question yet.
Question 10 MSQ · 5.0 marks
Why is Fast R-CNN faster than the original R-CNN?
It processes the whole image with a CNN once, rather than passing each
region proposal separately.
It uses RoI Pooling to extract features for each region from the shared feature
map.
It eliminates the need for region proposals entirely.
It uses a single-stage regression like YOLO.
A published solution is not available for this question yet.
Question 11 MSQ · 5.0 marks
Which of the following statements are true regarding the evaluation of Object Detection models?
IoU (Intersection over Union) measures the overlap between the predicted box
and the ground truth box.
A detection is considered a False Positive if the IoU is below a certain threshold
(e.g., 0.5).
mAP (Mean Average Precision) is calculated by averaging the Area Under the
Precision-Recall Curve across all classes.
High Precision implies that the model detects all ground truth objects, even if
it generates many false alarms.
A published solution is not available for this question yet.
Question 12 MSQ · 5.0 marks
Which of the following loss components are typically used in SRGAN (Super-Resolution GAN) to
achieve photo-realistic results?
Adversarial Loss: To encourage the generator to produce solutions on the
natural image manifold.
Content (Perceptual) Loss: Based on feature maps from a pre-trained VGG
network.
Pixel-wise MSE only: To strictly maximize PSNR.
Classification Loss: To classify the image into 1000 classes.
A published solution is not available for this question yet.
Question 13 MSQ · 5.0 marks
Which of the following statements correctly describe deep learning architectures used for
Monocular Depth Estimation (such as Eigen et al. or U-Net)? (Select all that apply)
The ”Coarse” network (or encoder) captures global scene structure but may
lack high-frequency details.
The ”Fine” network (or decoder) refines local details.
Max pooling is used to increase the spatial resolution of the depth map.
Skip connections are often used to preserve spatial information lost during
down sampling.
A published solution is not available for this question yet.
Question 14 MSQ · 5.0 marks
Select the correct characteristics of the YOLO (You Only Look Once) object detection framework.
It formulates object detection as a single regression problem, predicting
bounding boxes and class probabilities simultaneously.
It sees the entire image during training and testing, encoding contextual
information about classes.
It uses a Region Proposal Network (RPN) to generate proposals before
classification.
It splits the input image into an S × S grid.
A published solution is not available for this question yet.
Question 15 MSQ · 5.0 marks
Which of the following statements about ResNet (Residual Networks) are correct?
It solves the degradation problem where deeper networks have higher
training error than shallower ones.
It uses ”skip connections” (or identity shortcuts) to allow gradients to flow
more easily during backpropagation.
It requires significantly more parameters than VGGNet to achieve similar
accuracy.
It requires significantly less parameters than VGGNet to achieve similar
accuracy.
A published solution is not available for this question yet.
Question 16 MSQ · 5.0 marks
Which of the following are advantages of Convolutional Neural Networks (CNNs) over standard
Multi-Layer Perceptrons (MLPs) for image data?
Parameter Sharing: Weights are shared across the image, reducing the total
parameter count.
Local Connectivity: Neurons are connected only to a local region of the input.
Translation Invariance: Features can be detected regardless of their position in
the image.
Global Connectivity: Every pixel connects to every hidden neuron.
A published solution is not available for this question yet.
Question 17 MCQ · 5.0 marks
Which type of noise is caused by the statistical quantum fluctuations in the number of photons
sensed by the camera sensor?
Quantization noise
Photon shot noise
Read noise
Salt and pepper noise
A published solution is not available for this question yet.
Question 18 MCQ · 5.0 marks
What is the main motivation for using 1×1 convolutions (bottleneck layers) in the Inception module
of GoogLeNet?
To reduce the dimensionality (depth) of feature maps before expensive 3 × 3
or 5 × 5 convolutions.
To increase the spatial dimensions of the image.
To introduce non-linearity without changing channel dimensions.
To perform global average pooling.
A published solution is not available for this question yet.
Question 19 MCQ · 5.0 marks
In the YOLO (v1) algorithm, detection is modeled as a regression problem. If an object’s center
falls into a specific grid cell, which entity is responsible for detecting that object?
The neighboring grid cells.
The anchor box with the lowest IoU.
That specific grid cell.
The fully connected layer at the end only.
A published solution is not available for this question yet.
Question 20 MCQ · 5.0 marks
Which of the following cues is NOT available for Monocular Depth Estimation (depth from a single
image), making the task ”ill-posed”?
Texture gradient
Object size
Stereo Parallax
Linear perspective (vanishing points)
A published solution is not available for this question yet.
Question 21 MCQ · 5.0 marks
In the Faster R-CNN architecture, what is the specific role of the Region Proposal Network (RPN)?
To classify the object into one of the 1000 ImageNet classes.
To map sliding windows to object/not-object scores and bounding box
coordinates.
To perform Non-Maximum Suppression (NMS) on the final output.
To compute the mAP score during training.
A published solution is not available for this question yet.
Question 22 MCQ · 5.0 marks
In the VGGNet architecture, the design philosophy relies on stacking small convolutional filters.
What is the effective receptive field of a stack of three 3 × 3 convolutional layers (stride 1)?
3 x 3
5 x 5
7 x 7
9 x 9
A published solution is not available for this question yet.
Question 23 MCQ · 5.0 marks
In the MPRNet (Multi-stage Progressive Image Restoration) architecture, what is the primary
purpose of the Cross-stage Feature Fusion (CSFF) module?
To downsample the image to reduce memory usage.
To pass features from one stage to the next to minimize information loss.
To calculate the Perceptual Loss using a pretrained VGG network.
To generate random noise for the Generator.
A published solution is not available for this question yet.