da5013_2026T1_ET_FN.pdf
Deep Learning Practice · End Term · Jan 2026 FN
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 NAT · 1.0 marks
Read the following configuration file and code snippet related to training and annotations in a
YOLO model, and answer the given subquestions.
A dataset configuration file (YAML) is defined as:
path: /data/helmet_dataset
train: images/train
val: images/val
nc: 5
names: [’helmet’, ’no_helmet’, ’person’,’car’, ’truck’]
A sample annotation in YOLO format is given below:
2 0.25 0.60 0.40 0.20
This represents: classid, xcenter, ycenter,width, height.
Assume the corresponding image has dimensions: 640 (width) × 480 (height).
Based on the YAML configuration file, what is the total number of object classes the model is
trained to detect?
A published solution is not available for this question yet.
Question 3 NAT · 2.0 marks
Read the following configuration file and code snippet related to training and annotations in a
YOLO model, and answer the given subquestions.
A dataset configuration file (YAML) is defined as:
path: /data/helmet_dataset
train: images/train
val: images/val
nc: 5
names: [’helmet’, ’no_helmet’, ’person’,’car’, ’truck’]
A sample annotation in YOLO format is given below:
2 0.25 0.60 0.40 0.20
This represents: classid, xcenter, ycenter,width, height.
Assume the corresponding image has dimensions: 640 (width) × 480 (height).
What is the x-coordinate (in pixels) of the center of the bounding box?
A published solution is not available for this question yet.
Question 4 NAT · 2.0 marks
Read the following configuration file and code snippet related to training and annotations in a
YOLO model, and answer the given subquestions.
A dataset configuration file (YAML) is defined as:
path: /data/helmet_dataset
train: images/train
val: images/val
nc: 5
names: [’helmet’, ’no_helmet’, ’person’,’car’, ’truck’]
A sample annotation in YOLO format is given below:
2 0.25 0.60 0.40 0.20
This represents: classid, xcenter, ycenter,width, height.
Assume the corresponding image has dimensions: 640 (width) × 480 (height).
What is the x-coordinate (in pixels) of the top-left corner of the bounding box?
A published solution is not available for this question yet.
Question 5 NAT · 5.0 marks
Consider a convolutional neural network with an input image of size 84 × 84 × 3. The architecture
consists of the following layers:
1.A convolutional layer with 16 filters of size 4 × 4, a stride of 2, and no padding.
2. A ReLU activation layer.
3. A second convolutional layer with 32 filters of size 2×2, a stride of 2, and no padding.
4. A flattening layer.
What is the dimension of the resulting flattened layer?
A published solution is not available for this question yet.
Question 6 NAT · 5.0 marks
A bottleneck layer in a ResNet receives an input feature map of dimensions 14×14×1024. We apply
a 1×1 convolution layer with 256 filters to reduce dimensionality. Calculate the total number of
multiplications required for this operation. (Give answer in millions, e.g., if 65,452,123, write
65.45).
A published solution is not available for this question yet.
Question 7 NAT · 5.0 marks
A modified YOLO-style architecture divides an image into a 13 × 13 grid (S = 13). If the model is
designed to predict B = 3 bounding boxes per grid cell and is trained to recognize C = 80 distinct
classes (similar to COCO dataset), what is the depth (number of channels) of the final output
tensor?
A published solution is not available for this question yet.
Question 8 NAT · 5.0 marks
In the VGG16 architecture, consider the second convolutional layer of the first block. It takes an
input with 64 channels and applies 64 filters of size 3 × 3. Calculate the total number of weights
(excluding bias) for this specific layer.
A published solution is not available for this question yet.
Question 9 MSQ · 5.0 marks
When fine-tuning a large vision model (e.g., a pre-trained ResNet or Vision Transformer) for a
specific competition dataset, which of the following techniques are most effective for improving
training efficiency (speed and memory usage)?
Mixed Precision Training (FP16): Using 16-bit floating-point numbers instead of
32-bit for certain operations to speed up computation and reduce GPU memory consumption.
Freeze the Backbone: Keeping the early layers of the pre-trained model non-
trainable (requires grad=False) and only training the final classification head.
Increase Batch Size: Utilizing larger batches improves training efficiency by
lowering memory usage.
Data Prefetching: Using multiple CPU workers (e.g., num workers > 0 in
PyTorch) to load and augment data in the background while the GPU processes the current batch.
A published solution is not available for this question yet.
Question 10 MSQ · 5.0 marks
When training a deep learning model for a vision competition (such as a ”Cat vs Dog” classifier or a
YOLO detector), which of the following statements correctly describe the technical practices for
managing hardware and model state?
To utilize a GPU in PyTorch, you must explicitly move both the model
parameters and the input data tensors to the same device (e.g.,.to(’cuda’)).
Saving a model using torch.save(model.state dict(), PATH) is generally
preferred over saving the entire model object because it only stores the learnable parameters
(weights and biases).
When resuming training or performing inference, you must first instantiate
the model architecture and then load the weights using model.load state dict(torch.load(PATH,
map location=device)).
The state dict of a model includes the architecture's source code, allowing the
weights to be loaded onto a completely different model class without errors.
A published solution is not available for this question yet.
Question 11 MSQ · 5.0 marks
Consider the following Python code snippet using the Ultralytics YOLOv8 API:
from ultralytics import YOLO
model = YOLO("yolov8n.pt")
for name, module in model.model.named_modules():
print(name, type(module))
Based on the architecture of YOLOv8, which of the following statements correctly describe the
model?
The model predicts object locations directly without using predefined anchor
boxes.
The model uses separate components within its detection head for
classification and bounding box regression.
The model performs detection in a single forward pass without a separate
proposal stage.
The model internally generates region proposals using a Region Proposal
Network (RPN) before classification.
A published solution is not available for this question yet.
Question 12 MSQ · 5.0 marks
Which of the following loss components are typically used in SRGAN (Super-Resolution GAN) to
achieve photo-realistic results?
Adversarial Loss: To encourage the generator to produce solutions on the
natural image manifold.
Content (Perceptual) Loss: Based on feature maps from a pre-trained VGG
network.
Pixel-wise MSE only: To strictly maximize PSNR.
Classification Loss: To classify the image into 1000 classes.
A published solution is not available for this question yet.
Question 13 MSQ · 5.0 marks
Regarding the U-Net architecture, which is widely used for tasks like medical image segmentation
and can be adapted for depth estimation:
It features a symmetric architecture consisting of a contracting path (encoder)
and an expansive path (decoder).
Skip connections concatenate high-resolution features from the contracting
path directly to the up sampled features in the expansive path.
The expansive path uses transposed convolutions (or up-convolutions) to
increase the spatial resolution of the feature maps.
It relies on Global Average Pooling at every layer to ensure that spatial
information is discarded in favor of global context.
A published solution is not available for this question yet.
Question 14 MSQ · 5.0 marks
Consider the following code:
model = torchvision.models.detection.fasterrcnn_resnet50_fpn(pretrained=True)
images = [torch.randn(3, 600, 800)]
outputs = model(images)
Unlike YOLO-based models, the number of bounding boxes in outputs[0]['boxes'] varies for each
image.
What architectural characteristic of the model explains this behavior?
The model generates a variable number of region proposals based on the
image content.
The model filters region proposals using confidence scores and non maximum
suppression.
The model uses a fixed grid structure to predict bounding boxes.
The model predicts a constant number of bounding boxes for every image.
A published solution is not available for this question yet.
Question 15 MSQ · 5.0 marks
Which of the following statements correctly describe the characteristics and motivations of the
Inception (GoogLeNet) architecture?
It uses Inception modules that apply multiple filter sizes (1×1, 3×3, 5×5) in
parallel to capture multi-scale features.
It utilizes 1×1 convolutions as bottleneck layers to reduce dimensionality
before computationally expensive operations.
It incorporates Auxiliary Classifiers in intermediate layers to inject additional
gradient signal and combat the vanishing gradient problem.
It relies exclusively on stacking 3 × 3 filters to achieve its receptive field, similar
to the VGGNet design philosophy.
A published solution is not available for this question yet.
Question 16 MSQ · 5.0 marks
In the context of Convolutional Neural Networks (CNNs), what is the primary mechanism that
allows the network to handle inputs of varying spatial sizes while maintaining a fixed number of
parameters, a characteristic that distinguishes them from standard Multi-Layer Perceptrons
(MLPs)?
Neurons are connected only to a local region of the input.
Max Pooling: Reducing the spatial resolution to a single pixel before the first
hidden layer.
Parameter Sharing: Using the same set of weights (filters) across different
spatial locations of the input.
Fully Connected Layers: Replacing all convolutional layers with dense layers to
increase the receptive field.
A published solution is not available for this question yet.
Question 17 MCQ · 5.0 marks
The first layer of the VGG16 model is a Conv2d with kernel = 3, stride = 1, padding = 1.
Consider the following Python code snippet using PyTorch and a pre-trained VGG16 model to
process an input image:
import torch
import torch.nn as nn
from torchvision import models
model = models.vgg16(pretrained=True)
feature_extractor = model.features[0]
input_image = torch.randn(1, 3, 224, 224)
output = feature_extractor(input_image)
print(output.shape)
What will be the shape of the output tensor printed by this code?
torch.Size([1, 3, 224, 224])
torch.Size([1, 64, 112, 112])
torch.Size([1, 64, 224, 224])
torch.Size([1, 64, 222, 222])
A published solution is not available for this question yet.
Question 18 MCQ · 5.0 marks
In the ResNet (Residual Network) architecture, 1 × 1 convolutions are frequently used in the
”bottleneck” building block. Beyond dimensionality reduction, what is an additional benefit of
using these layers compared to a standard building block?
They allow the network to increase the spatial resolution of feature maps to
recover lost details.
They allow for the addition of non-linearity (via activation functions) without
increasing the receptive field or computational cost excessively.
They are used to perform Max Pooling operations within the residual
mapping.
They eliminate the need for skip connections by allowing gradients to flow
through the 1 × 1 filters instead.
A published solution is not available for this question yet.
Question 19 MCQ · 5.0 marks
In the YOLO (You Only Look Once) framework, if multiple bounding boxes are predicted by a single
grid cell for the same object, how does the algorithm determine which specific
bounding box is responsible for that prediction during the training process?
It is observed that the number of bounding boxes printed varies for different input images.
Which of the following statements best explains this behavior?
All predicted boxes for that cell are responsible and updated simultaneously.
The bounding box that has the highest class probability score.
The bounding box that has the highest Intersection over Union (IoU) with the
ground truth.
The bounding box that is physically closest to the center of the image.
A published solution is not available for this question yet.
Question 20 MCQ · 5.0 marks
Consider the following code snippet using a YOLO-based object detection model:
from ultralytics import YOLO
model = YOLO("yolov8n.pt")
results = model("image.jpg")
boxes = results[0].boxes.xyxy
print(boxes.shape)
It is observed that the number of bounding boxes printed varies for different input images.
Which of the following statements best explains this behavior?
The model divides the image into a fixed grid and always predicts a fixed
number of boxes.
The model filters predictions based on confidence scores and non maximum
suppression, resulting in a variable number of final detections.
The model uses a Region Proposal Network (RPN) to generate candidate
regions before prediction.
The number of bounding boxes is fixed by the number of classes in the
dataset.
A published solution is not available for this question yet.
Question 21 MCQ · 5.0 marks
In the SRResNet and SRGAN architectures for image super-resolution, which specialized layer is
used to upsample the feature maps by rearranging elements from the channel dimension into the
spatial dimension?
The Max Pooling layer with a stride of 2.
The Bilinear Interpolation layer for smooth scaling.
The Pixel Shuffle (Sub-pixel Convolution) layer.
The Global Average Pooling layer to reduce spatial dimensions.
A published solution is not available for this question yet.
Question 22 MCQ · 5.0 marks
In the Fast R-CNN architecture, which component is responsible for extracting a fixed-size feature
vector from a shared convolutional feature map for each region proposal?
The Region Proposal Network (RPN) which generates anchors.
The Region of Interest (RoI) Pooling layer.
A stack of three 3 × 3 convolutional layers.
The Global Average Pooling layer used in GoogLeNet.
A published solution is not available for this question yet.
Question 23 MCQ · 5.0 marks
In the GoogLeNet (Inception v1) architecture, 1 × 1 convolutions (bottleneck layers) are applied
before larger 3 × 3 and 5 × 5 convolutions. What is the primary motivation for including these 1 × 1
filters within an Inception module?
Transposed Convolution.
To reduce the dimensionality (depth) of feature maps to manage
computational complexity.
To replace the need for skip connections and solve the degradation problem.
To strictly enforce a mean of 0 and variance of 1 across the channel
dimension.
A published solution is not available for this question yet.