MauryaHub PYQ Practice

da5013_2025T3_ET_FN.pdf

Deep Learning Practice · End Term · Sep 2025 FN

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 2 NAT · 1.0 marks

Read the data below and answer the given subquestions regarding Object Detection Evaluation. An object detector is evaluated on a single class ”Car”. The dataset has 10 ground truth ”Car” objects. The model makes 10 detections. Based on an IoU threshold of 0.5: • 6 detections are True Positives (TP). • 4 detections are False Positives (FP). • (Consequently, there are 4 False Negatives (FN), as 10 ground truths - 6 found = 4 missed).
Calculate the Precision of the model. (Answer as a decimal between 0 and 1).

    A published solution is not available for this question yet.

    Question 3 NAT · 2.0 marks

    Read the data below and answer the given subquestions regarding Object Detection Evaluation. An object detector is evaluated on a single class ”Car”. The dataset has 10 ground truth ”Car” objects. The model makes 10 detections. Based on an IoU threshold of 0.5: • 6 detections are True Positives (TP). • 4 detections are False Positives (FP). • (Consequently, there are 4 False Negatives (FN), as 10 ground truths - 6 found = 4 missed).
    Calculate the Recall of the model. (Answer as a decimal between 0 and 1).

      A published solution is not available for this question yet.

      Question 4 NAT · 2.0 marks

      Read the data below and answer the given subquestions regarding Object Detection Evaluation. An object detector is evaluated on a single class ”Car”. The dataset has 10 ground truth ”Car” objects. The model makes 10 detections. Based on an IoU threshold of 0.5: • 6 detections are True Positives (TP). • 4 detections are False Positives (FP). • (Consequently, there are 4 False Negatives (FN), as 10 ground truths - 6 found = 4 missed).
      If the IoU threshold is increased to 0.9, one of the previous True Positives becomes a False Positive. What is the new Precision? (Round to 2 decimal places).

        A published solution is not available for this question yet.

        Question 5 NAT · 5.0 marks

        An input feature map has dimensions 28 × 28 × 192. We apply a 1 × 1 convolution layer with 64 filters. What is the total number of operations (multiplications) required? (Give answer in millions, e.g., if 1,000,000, write 1).

          A published solution is not available for this question yet.

          Question 6 NAT · 5.0 marks

          A stereo camera system has a baseline B = 0.5 meters and a focal length f = 100 pixels. If the disparity d for a specific point is calculated to be 10 pixels, what is the depth Z of that point in meters?

            A published solution is not available for this question yet.

            Question 7 NAT · 5.0 marks

            In the original YOLO paper, the image is divided into a 7 × 7 grid (S = 7). The model predicts 2 bounding boxes per cell (B = 2) and detects 20 classes (C = 20). What is the depth (number of channels) of the final output tensor?

              A published solution is not available for this question yet.

              Question 8 NAT · 5.0 marks

              In AlexNet, the first convolutional layer uses filters of size 11 × 11 × 3. If there are 96 such filters, calculate the total number of weights (excluding bias) in this layer.

                A published solution is not available for this question yet.

                Question 9 MSQ · 5.0 marks

                In Unsupervised Monocular Depth Estimation using Left-Right Consistency :
                1. The model is trained using stereo image pairs but requires only a single image at test time.
                2. It uses the predicted disparity to warp the left image to match the right image (or vice versa).
                3. The loss function is based on the photometric reconstruction error between the warped image and the target image.
                4. It requires ground truth depth maps captured by LiDAR for training.

                A published solution is not available for this question yet.

                Question 10 MSQ · 5.0 marks

                Why is Fast R-CNN faster than the original R-CNN?
                1. It processes the whole image with a CNN once, rather than passing each region proposal separately.
                2. It uses RoI Pooling to extract features for each region from the shared feature map.
                3. It eliminates the need for region proposals entirely.
                4. It uses a single-stage regression like YOLO.

                A published solution is not available for this question yet.

                Question 11 MSQ · 5.0 marks

                Which of the following statements are true regarding the evaluation of Object Detection models?
                1. IoU (Intersection over Union) measures the overlap between the predicted box and the ground truth box.
                2. A detection is considered a False Positive if the IoU is below a certain threshold (e.g., 0.5).
                3. mAP (Mean Average Precision) is calculated by averaging the Area Under the Precision-Recall Curve across all classes.
                4. High Precision implies that the model detects all ground truth objects, even if it generates many false alarms.

                A published solution is not available for this question yet.

                Question 12 MSQ · 5.0 marks

                Which of the following loss components are typically used in SRGAN (Super-Resolution GAN) to achieve photo-realistic results?
                1. Adversarial Loss: To encourage the generator to produce solutions on the natural image manifold.
                2. Content (Perceptual) Loss: Based on feature maps from a pre-trained VGG network.
                3. Pixel-wise MSE only: To strictly maximize PSNR.
                4. Classification Loss: To classify the image into 1000 classes.

                A published solution is not available for this question yet.

                Question 13 MSQ · 5.0 marks

                Which of the following statements correctly describe deep learning architectures used for Monocular Depth Estimation (such as Eigen et al. or U-Net)? (Select all that apply)
                1. The ”Coarse” network (or encoder) captures global scene structure but may lack high-frequency details.
                2. The ”Fine” network (or decoder) refines local details.
                3. Max pooling is used to increase the spatial resolution of the depth map.
                4. Skip connections are often used to preserve spatial information lost during down sampling.

                A published solution is not available for this question yet.

                Question 14 MSQ · 5.0 marks

                Select the correct characteristics of the YOLO (You Only Look Once) object detection framework.
                1. It formulates object detection as a single regression problem, predicting bounding boxes and class probabilities simultaneously.
                2. It sees the entire image during training and testing, encoding contextual information about classes.
                3. It uses a Region Proposal Network (RPN) to generate proposals before classification.
                4. It splits the input image into an S × S grid.

                A published solution is not available for this question yet.

                Question 15 MSQ · 5.0 marks

                Which of the following statements about ResNet (Residual Networks) are correct?
                1. It solves the degradation problem where deeper networks have higher training error than shallower ones.
                2. It uses ”skip connections” (or identity shortcuts) to allow gradients to flow more easily during backpropagation.
                3. It requires significantly more parameters than VGGNet to achieve similar accuracy.
                4. It requires significantly less parameters than VGGNet to achieve similar accuracy.

                A published solution is not available for this question yet.

                Question 16 MSQ · 5.0 marks

                Which of the following are advantages of Convolutional Neural Networks (CNNs) over standard Multi-Layer Perceptrons (MLPs) for image data?
                1. Parameter Sharing: Weights are shared across the image, reducing the total parameter count.
                2. Local Connectivity: Neurons are connected only to a local region of the input.
                3. Translation Invariance: Features can be detected regardless of their position in the image.
                4. Global Connectivity: Every pixel connects to every hidden neuron.

                A published solution is not available for this question yet.

                Question 17 MCQ · 5.0 marks

                Which type of noise is caused by the statistical quantum fluctuations in the number of photons sensed by the camera sensor?
                1. Quantization noise
                2. Photon shot noise
                3. Read noise
                4. Salt and pepper noise

                A published solution is not available for this question yet.

                Question 18 MCQ · 5.0 marks

                What is the main motivation for using 1×1 convolutions (bottleneck layers) in the Inception module of GoogLeNet?
                1. To reduce the dimensionality (depth) of feature maps before expensive 3 × 3 or 5 × 5 convolutions.
                2. To increase the spatial dimensions of the image.
                3. To introduce non-linearity without changing channel dimensions.
                4. To perform global average pooling.

                A published solution is not available for this question yet.

                Question 19 MCQ · 5.0 marks

                In the YOLO (v1) algorithm, detection is modeled as a regression problem. If an object’s center falls into a specific grid cell, which entity is responsible for detecting that object?
                1. The neighboring grid cells.
                2. The anchor box with the lowest IoU.
                3. That specific grid cell.
                4. The fully connected layer at the end only.

                A published solution is not available for this question yet.

                Question 20 MCQ · 5.0 marks

                Which of the following cues is NOT available for Monocular Depth Estimation (depth from a single image), making the task ”ill-posed”?
                1. Texture gradient
                2. Object size
                3. Stereo Parallax
                4. Linear perspective (vanishing points)

                A published solution is not available for this question yet.

                Question 21 MCQ · 5.0 marks

                In the Faster R-CNN architecture, what is the specific role of the Region Proposal Network (RPN)?
                1. To classify the object into one of the 1000 ImageNet classes.
                2. To map sliding windows to object/not-object scores and bounding box coordinates.
                3. To perform Non-Maximum Suppression (NMS) on the final output.
                4. To compute the mAP score during training.

                A published solution is not available for this question yet.

                Question 22 MCQ · 5.0 marks

                In the VGGNet architecture, the design philosophy relies on stacking small convolutional filters. What is the effective receptive field of a stack of three 3 × 3 convolutional layers (stride 1)?
                1. 3 x 3
                2. 5 x 5
                3. 7 x 7
                4. 9 x 9

                A published solution is not available for this question yet.

                Question 23 MCQ · 5.0 marks

                In the MPRNet (Multi-stage Progressive Image Restoration) architecture, what is the primary purpose of the Cross-stage Feature Fusion (CSFF) module?
                1. To downsample the image to reduce memory usage.
                2. To pass features from one stage to the next to minimize information loss.
                3. To calculate the Perceptual Loss using a pretrained VGG network.
                4. To generate random noise for the Generator.

                A published solution is not available for this question yet.