MauryaHub PYQ Practice

da5013_2026T1_ET_FN.pdf

Deep Learning Practice · End Term · Jan 2026 FN

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 2 NAT · 1.0 marks

Read the following configuration file and code snippet related to training and annotations in a YOLO model, and answer the given subquestions. A dataset configuration file (YAML) is defined as: path: /data/helmet_dataset train: images/train val: images/val nc: 5 names: [’helmet’, ’no_helmet’, ’person’,’car’, ’truck’] A sample annotation in YOLO format is given below: 2 0.25 0.60 0.40 0.20 This represents: classid, xcenter, ycenter,width, height. Assume the corresponding image has dimensions: 640 (width) × 480 (height).
Based on the YAML configuration file, what is the total number of object classes the model is trained to detect?

    A published solution is not available for this question yet.

    Question 3 NAT · 2.0 marks

    Read the following configuration file and code snippet related to training and annotations in a YOLO model, and answer the given subquestions. A dataset configuration file (YAML) is defined as: path: /data/helmet_dataset train: images/train val: images/val nc: 5 names: [’helmet’, ’no_helmet’, ’person’,’car’, ’truck’] A sample annotation in YOLO format is given below: 2 0.25 0.60 0.40 0.20 This represents: classid, xcenter, ycenter,width, height. Assume the corresponding image has dimensions: 640 (width) × 480 (height).
    What is the x-coordinate (in pixels) of the center of the bounding box?

      A published solution is not available for this question yet.

      Question 4 NAT · 2.0 marks

      Read the following configuration file and code snippet related to training and annotations in a YOLO model, and answer the given subquestions. A dataset configuration file (YAML) is defined as: path: /data/helmet_dataset train: images/train val: images/val nc: 5 names: [’helmet’, ’no_helmet’, ’person’,’car’, ’truck’] A sample annotation in YOLO format is given below: 2 0.25 0.60 0.40 0.20 This represents: classid, xcenter, ycenter,width, height. Assume the corresponding image has dimensions: 640 (width) × 480 (height).
      What is the x-coordinate (in pixels) of the top-left corner of the bounding box?

        A published solution is not available for this question yet.

        Question 5 NAT · 5.0 marks

        Consider a convolutional neural network with an input image of size 84 × 84 × 3. The architecture consists of the following layers: 1.A convolutional layer with 16 filters of size 4 × 4, a stride of 2, and no padding. 2. A ReLU activation layer. 3. A second convolutional layer with 32 filters of size 2×2, a stride of 2, and no padding. 4. A flattening layer. What is the dimension of the resulting flattened layer?

          A published solution is not available for this question yet.

          Question 6 NAT · 5.0 marks

          A bottleneck layer in a ResNet receives an input feature map of dimensions 14×14×1024. We apply a 1×1 convolution layer with 256 filters to reduce dimensionality. Calculate the total number of multiplications required for this operation. (Give answer in millions, e.g., if 65,452,123, write 65.45).

            A published solution is not available for this question yet.

            Question 7 NAT · 5.0 marks

            A modified YOLO-style architecture divides an image into a 13 × 13 grid (S = 13). If the model is designed to predict B = 3 bounding boxes per grid cell and is trained to recognize C = 80 distinct classes (similar to COCO dataset), what is the depth (number of channels) of the final output tensor?

              A published solution is not available for this question yet.

              Question 8 NAT · 5.0 marks

              In the VGG16 architecture, consider the second convolutional layer of the first block. It takes an input with 64 channels and applies 64 filters of size 3 × 3. Calculate the total number of weights (excluding bias) for this specific layer.

                A published solution is not available for this question yet.

                Question 9 MSQ · 5.0 marks

                When fine-tuning a large vision model (e.g., a pre-trained ResNet or Vision Transformer) for a specific competition dataset, which of the following techniques are most effective for improving training efficiency (speed and memory usage)?
                1. Mixed Precision Training (FP16): Using 16-bit floating-point numbers instead of 32-bit for certain operations to speed up computation and reduce GPU memory consumption.
                2. Freeze the Backbone: Keeping the early layers of the pre-trained model non- trainable (requires grad=False) and only training the final classification head.
                3. Increase Batch Size: Utilizing larger batches improves training efficiency by lowering memory usage.
                4. Data Prefetching: Using multiple CPU workers (e.g., num workers > 0 in PyTorch) to load and augment data in the background while the GPU processes the current batch.

                A published solution is not available for this question yet.

                Question 10 MSQ · 5.0 marks

                When training a deep learning model for a vision competition (such as a ”Cat vs Dog” classifier or a YOLO detector), which of the following statements correctly describe the technical practices for managing hardware and model state?
                1. To utilize a GPU in PyTorch, you must explicitly move both the model parameters and the input data tensors to the same device (e.g.,.to(’cuda’)).
                2. Saving a model using torch.save(model.state dict(), PATH) is generally preferred over saving the entire model object because it only stores the learnable parameters (weights and biases).
                3. When resuming training or performing inference, you must first instantiate the model architecture and then load the weights using model.load state dict(torch.load(PATH, map location=device)).
                4. The state dict of a model includes the architecture's source code, allowing the weights to be loaded onto a completely different model class without errors.

                A published solution is not available for this question yet.

                Question 11 MSQ · 5.0 marks

                Consider the following Python code snippet using the Ultralytics YOLOv8 API: from ultralytics import YOLO model = YOLO("yolov8n.pt") for name, module in model.model.named_modules(): print(name, type(module)) Based on the architecture of YOLOv8, which of the following statements correctly describe the model?
                1. The model predicts object locations directly without using predefined anchor boxes.
                2. The model uses separate components within its detection head for classification and bounding box regression.
                3. The model performs detection in a single forward pass without a separate proposal stage.
                4. The model internally generates region proposals using a Region Proposal Network (RPN) before classification.

                A published solution is not available for this question yet.

                Question 12 MSQ · 5.0 marks

                Which of the following loss components are typically used in SRGAN (Super-Resolution GAN) to achieve photo-realistic results?
                1. Adversarial Loss: To encourage the generator to produce solutions on the natural image manifold.
                2. Content (Perceptual) Loss: Based on feature maps from a pre-trained VGG network.
                3. Pixel-wise MSE only: To strictly maximize PSNR.
                4. Classification Loss: To classify the image into 1000 classes.

                A published solution is not available for this question yet.

                Question 13 MSQ · 5.0 marks

                Regarding the U-Net architecture, which is widely used for tasks like medical image segmentation and can be adapted for depth estimation:
                1. It features a symmetric architecture consisting of a contracting path (encoder) and an expansive path (decoder).
                2. Skip connections concatenate high-resolution features from the contracting path directly to the up sampled features in the expansive path.
                3. The expansive path uses transposed convolutions (or up-convolutions) to increase the spatial resolution of the feature maps.
                4. It relies on Global Average Pooling at every layer to ensure that spatial information is discarded in favor of global context.

                A published solution is not available for this question yet.

                Question 14 MSQ · 5.0 marks

                Consider the following code: model = torchvision.models.detection.fasterrcnn_resnet50_fpn(pretrained=True) images = [torch.randn(3, 600, 800)] outputs = model(images) Unlike YOLO-based models, the number of bounding boxes in outputs[0]['boxes'] varies for each image. What architectural characteristic of the model explains this behavior?
                1. The model generates a variable number of region proposals based on the image content.
                2. The model filters region proposals using confidence scores and non maximum suppression.
                3. The model uses a fixed grid structure to predict bounding boxes.
                4. The model predicts a constant number of bounding boxes for every image.

                A published solution is not available for this question yet.

                Question 15 MSQ · 5.0 marks

                Which of the following statements correctly describe the characteristics and motivations of the Inception (GoogLeNet) architecture?
                1. It uses Inception modules that apply multiple filter sizes (1×1, 3×3, 5×5) in parallel to capture multi-scale features.
                2. It utilizes 1×1 convolutions as bottleneck layers to reduce dimensionality before computationally expensive operations.
                3. It incorporates Auxiliary Classifiers in intermediate layers to inject additional gradient signal and combat the vanishing gradient problem.
                4. It relies exclusively on stacking 3 × 3 filters to achieve its receptive field, similar to the VGGNet design philosophy.

                A published solution is not available for this question yet.

                Question 16 MSQ · 5.0 marks

                In the context of Convolutional Neural Networks (CNNs), what is the primary mechanism that allows the network to handle inputs of varying spatial sizes while maintaining a fixed number of parameters, a characteristic that distinguishes them from standard Multi-Layer Perceptrons (MLPs)?
                1. Neurons are connected only to a local region of the input.
                2. Max Pooling: Reducing the spatial resolution to a single pixel before the first hidden layer.
                3. Parameter Sharing: Using the same set of weights (filters) across different spatial locations of the input.
                4. Fully Connected Layers: Replacing all convolutional layers with dense layers to increase the receptive field.

                A published solution is not available for this question yet.

                Question 17 MCQ · 5.0 marks

                The first layer of the VGG16 model is a Conv2d with kernel = 3, stride = 1, padding = 1. Consider the following Python code snippet using PyTorch and a pre-trained VGG16 model to process an input image: import torch import torch.nn as nn from torchvision import models model = models.vgg16(pretrained=True) feature_extractor = model.features[0] input_image = torch.randn(1, 3, 224, 224) output = feature_extractor(input_image) print(output.shape) What will be the shape of the output tensor printed by this code?
                1. torch.Size([1, 3, 224, 224])
                2. torch.Size([1, 64, 112, 112])
                3. torch.Size([1, 64, 224, 224])
                4. torch.Size([1, 64, 222, 222])

                A published solution is not available for this question yet.

                Question 18 MCQ · 5.0 marks

                In the ResNet (Residual Network) architecture, 1 × 1 convolutions are frequently used in the ”bottleneck” building block. Beyond dimensionality reduction, what is an additional benefit of using these layers compared to a standard building block?
                1. They allow the network to increase the spatial resolution of feature maps to recover lost details.
                2. They allow for the addition of non-linearity (via activation functions) without increasing the receptive field or computational cost excessively.
                3. They are used to perform Max Pooling operations within the residual mapping.
                4. They eliminate the need for skip connections by allowing gradients to flow through the 1 × 1 filters instead.

                A published solution is not available for this question yet.

                Question 19 MCQ · 5.0 marks

                In the YOLO (You Only Look Once) framework, if multiple bounding boxes are predicted by a single grid cell for the same object, how does the algorithm determine which specific bounding box is responsible for that prediction during the training process? It is observed that the number of bounding boxes printed varies for different input images. Which of the following statements best explains this behavior?
                1. All predicted boxes for that cell are responsible and updated simultaneously.
                2. The bounding box that has the highest class probability score.
                3. The bounding box that has the highest Intersection over Union (IoU) with the ground truth.
                4. The bounding box that is physically closest to the center of the image.

                A published solution is not available for this question yet.

                Question 20 MCQ · 5.0 marks

                Consider the following code snippet using a YOLO-based object detection model: from ultralytics import YOLO model = YOLO("yolov8n.pt") results = model("image.jpg") boxes = results[0].boxes.xyxy print(boxes.shape) It is observed that the number of bounding boxes printed varies for different input images. Which of the following statements best explains this behavior?
                1. The model divides the image into a fixed grid and always predicts a fixed number of boxes.
                2. The model filters predictions based on confidence scores and non maximum suppression, resulting in a variable number of final detections.
                3. The model uses a Region Proposal Network (RPN) to generate candidate regions before prediction.
                4. The number of bounding boxes is fixed by the number of classes in the dataset.

                A published solution is not available for this question yet.

                Question 21 MCQ · 5.0 marks

                In the SRResNet and SRGAN architectures for image super-resolution, which specialized layer is used to upsample the feature maps by rearranging elements from the channel dimension into the spatial dimension?
                1. The Max Pooling layer with a stride of 2.
                2. The Bilinear Interpolation layer for smooth scaling.
                3. The Pixel Shuffle (Sub-pixel Convolution) layer.
                4. The Global Average Pooling layer to reduce spatial dimensions.

                A published solution is not available for this question yet.

                Question 22 MCQ · 5.0 marks

                In the Fast R-CNN architecture, which component is responsible for extracting a fixed-size feature vector from a shared convolutional feature map for each region proposal?
                1. The Region Proposal Network (RPN) which generates anchors.
                2. The Region of Interest (RoI) Pooling layer.
                3. A stack of three 3 × 3 convolutional layers.
                4. The Global Average Pooling layer used in GoogLeNet.

                A published solution is not available for this question yet.

                Question 23 MCQ · 5.0 marks

                In the GoogLeNet (Inception v1) architecture, 1 × 1 convolutions (bottleneck layers) are applied before larger 3 × 3 and 5 × 5 convolutions. What is the primary motivation for including these 1 × 1 filters within an Inception module?
                1. Transposed Convolution.
                2. To reduce the dimensionality (depth) of feature maps to manage computational complexity.
                3. To replace the need for skip connections and solve the degradation problem.
                4. To strictly enforce a mean of 0 and variance of 1 across the channel dimension.

                A published solution is not available for this question yet.