MauryaHub PYQ Practice

da5013_2026T2_ET_FN.pdf

Deep Learning Practice · End Term · May 2026 FN

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 2 NAT · 1.0 marks

[[IMAGE:cf41183a839418fe_3_2]] Based on the above data, answer the given subquestions.
Based on the actual output of [[IMAGE:cf41183a839418fe_4_3]] for a 128×128 input, how many elements are there per sample when it is flattened (i.e., what should [[IMAGE:cf41183a839418fe_4_4]] evaluate to)?
Source diagram or notationSource diagram or notationSource diagram or notation

    A published solution is not available for this question yet.

    Question 3 MCQ · 1.0 marks

    [[IMAGE:cf41183a839418fe_3_2]] Based on the above data, answer the given subquestions.
    Which statement best explains why this code raises a [[IMAGE:cf41183a839418fe_4_5]] ?
    Source diagram or notationSource diagram or notation
    1. The classifier's declared input size (12544) does not match the actual flattened feature size produced by [[IMAGE:cf41183a839418fe_4_6]] .
      Source diagram or notation
    2. [[IMAGE:cf41183a839418fe_4_7]] uses an invalid stride value for a 3×3 kernel.
      Source diagram or notation
    3. [[IMAGE:cf41183a839418fe_4_8]] cannot be applied immediately after [[IMAGE:cf41183a839418fe_4_9]] .
      Source diagram or notationSource diagram or notation
    4. The batch size of 8 is incompatible with [[IMAGE:cf41183a839418fe_4_10]] .
      Source diagram or notation

    A published solution is not available for this question yet.

    Question 4 NAT · 1.0 marks

    [[IMAGE:cf41183a839418fe_3_2]] Based on the above data, answer the given subquestions.
    Using the actual flattened feature size from previous question number 2, how many learnable parameters (weights + bias) would a fully connected layer with 10 outputs have?
    Source diagram or notation

      A published solution is not available for this question yet.

      Question 5 MSQ · 2.0 marks

      [[IMAGE:cf41183a839418fe_3_2]] Based on the above data, answer the given subquestions.
      Which of the following statements about the (corrected) model are true?
      Source diagram or notation
      1. The output of [[IMAGE:cf41183a839418fe_5_11]] (before [[IMAGE:cf41183a839418fe_5_12]] ) has spatial size 62×62.
        Source diagram or notationSource diagram or notation
      2. [[IMAGE:cf41183a839418fe_5_13]] applied to a 62×62 input produces exactly 31×31, with no dimension lost to rounding.
        Source diagram or notation
      3. [[IMAGE:cf41183a839418fe_5_14]] 's stride of 2 also contributes to spatial downsampling in this network.
        Source diagram or notation
      4. The number of channels in the final output of [[IMAGE:cf41183a839418fe_5_15]] is 32.
        Source diagram or notation

      A published solution is not available for this question yet.

      Question 6 NAT · 1.0 marks

      [[IMAGE:cf41183a839418fe_5_16]] Recall: [[IMAGE:cf41183a839418fe_5_17]] for transpose Convolution ). The corresponding encoder to the above decoder downsampled an original 120×120 input via 3 stride-2 stages down to 15×15. Based on the above data, answer the given subquestions.
      What is the output height after [[IMAGE:cf41183a839418fe_6_18]] ? The width is the same; enter one integer only.
      Source diagram or notationSource diagram or notationSource diagram or notation

        A published solution is not available for this question yet.

        Question 7 NAT · 1.0 marks

        [[IMAGE:cf41183a839418fe_5_16]] Recall: [[IMAGE:cf41183a839418fe_5_17]] for transpose Convolution ). The corresponding encoder to the above decoder downsampled an original 120×120 input via 3 stride-2 stages down to 15×15. Based on the above data, answer the given subquestions.
        What is the output height after [[IMAGE:cf41183a839418fe_6_19]] ? The width is the same; enter one integer only.
        Source diagram or notationSource diagram or notationSource diagram or notation

          A published solution is not available for this question yet.

          Question 8 MCQ · 2.0 marks

          [[IMAGE:cf41183a839418fe_5_16]] Recall: [[IMAGE:cf41183a839418fe_5_17]] for transpose Convolution ). The corresponding encoder to the above decoder downsampled an original 120×120 input via 3 stride-2 stages down to 15×15. Based on the above data, answer the given subquestions.
          Which layer is responsible for the decoder failing to reach the target 120×120 output, and why?
          Source diagram or notationSource diagram or notation
          1. [[IMAGE:cf41183a839418fe_6_20]] , because its kernel size is too small.
            Source diagram or notation
          2. [[IMAGE:cf41183a839418fe_6_21]] , because it uses the wrong number of input channels.
            Source diagram or notation
          3. [[IMAGE:cf41183a839418fe_7_22]] , because it uses [[IMAGE:cf41183a839418fe_7_23]] instead of [[IMAGE:cf41183a839418fe_7_24]] , so it fails to double the spatial resolution.
            Source diagram or notationSource diagram or notationSource diagram or notation
          4. None of the layers — the decoder already produces a 120×120 output.

          A published solution is not available for this question yet.

          Question 9 NAT · 1.0 marks

          [[IMAGE:cf41183a839418fe_5_16]] Recall: [[IMAGE:cf41183a839418fe_5_17]] for transpose Convolution ). The corresponding encoder to the above decoder downsampled an original 120×120 input via 3 stride-2 stages down to 15×15. Based on the above data, answer the given subquestions.
          Suppose instead the decoder's input [[IMAGE:cf41183a839418fe_7_25]] had spatial size 30×30 (instead of 15×15). After applying the correction identified in previous question, what would the final output height be? The width is the same; enter one integer only.
          Source diagram or notationSource diagram or notationSource diagram or notation

            A published solution is not available for this question yet.

            Question 10 NAT · 1.0 marks

            [[IMAGE:cf41183a839418fe_7_26]] Consider an input of size **32×32×32** (32 channels), a standard [[IMAGE:cf41183a839418fe_7_27]] , and the [[IMAGE:cf41183a839418fe_8_28]] module above. Ignore bias terms throughout, and assume spatial size is preserved (32×32 output) in both cases. Based on the above data, answer the given subquestions.
            How many learnable parameters does the standard [[IMAGE:cf41183a839418fe_8_29]] have?
            Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

              A published solution is not available for this question yet.

              Question 11 NAT · 1.0 marks

              [[IMAGE:cf41183a839418fe_7_26]] Consider an input of size **32×32×32** (32 channels), a standard [[IMAGE:cf41183a839418fe_7_27]] , and the [[IMAGE:cf41183a839418fe_8_28]] module above. Ignore bias terms throughout, and assume spatial size is preserved (32×32 output) in both cases. Based on the above data, answer the given subquestions.
              How many total learnable parameters does [[IMAGE:cf41183a839418fe_8_30]] have (depthwise + pointwise combined)?
              Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

                A published solution is not available for this question yet.

                Question 12 NAT · 1.0 marks

                [[IMAGE:cf41183a839418fe_7_26]] Consider an input of size **32×32×32** (32 channels), a standard [[IMAGE:cf41183a839418fe_7_27]] , and the [[IMAGE:cf41183a839418fe_8_28]] module above. Ignore bias terms throughout, and assume spatial size is preserved (32×32 output) in both cases. Based on the above data, answer the given subquestions.
                What is the parameter compression ratio (standard ÷ depthwise-separable), rounded to 2 decimal places?
                Source diagram or notationSource diagram or notationSource diagram or notation

                  A published solution is not available for this question yet.

                  Question 13 MSQ · 2.0 marks

                  [[IMAGE:cf41183a839418fe_7_26]] Consider an input of size **32×32×32** (32 channels), a standard [[IMAGE:cf41183a839418fe_7_27]] , and the [[IMAGE:cf41183a839418fe_8_28]] module above. Ignore bias terms throughout, and assume spatial size is preserved (32×32 output) in both cases. Based on the above data, answer the given subquestions.
                  Which of the following statements about [[IMAGE:cf41183a839418fe_9_31]] are correct?
                  Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
                  1. The depthwise layer must use [[IMAGE:cf41183a839418fe_9_32]] so that each input channel is convolved with its own separate filter.
                    Source diagram or notation
                  2. The pointwise layer uses a 1×1 kernel to mix information across channels after the depthwise step.
                  3. Replacing a standard convolution with a depthwise-separable convolution always improves model accuracy.
                  4. The compression ratio achieved depends on both the kernel size and the number of output channels.

                  A published solution is not available for this question yet.

                  Question 14 MSQ · 1.0 marks

                  Assume the [[IMAGE:cf41183a839418fe_9_33]] architecture itself has already been corrected so that its classifier input size matches the feature output. Focus only on the training/checkpoint/inference pipeline below. [[IMAGE:cf41183a839418fe_10_34]] Based on the above data, answer the given subquestions.
                  Which of the following statements correctly identify a real bug in the given code ?
                  Source diagram or notationSource diagram or notation
                  1. [[IMAGE:cf41183a839418fe_11_35]] is never called, so gradients accumulate across batches instead of being reset before each step.
                    Source diagram or notation
                  2. [[IMAGE:cf41183a839418fe_11_36]] saves the full model object, but the reload code calls [[IMAGE:cf41183a839418fe_11_37]] , which expects a state dict rather than a model instance.
                    Source diagram or notationSource diagram or notation
                  3. [[IMAGE:cf41183a839418fe_11_38]] is never moved to [[IMAGE:cf41183a839418fe_11_39]] , so [[IMAGE:cf41183a839418fe_11_40]] will raise a device- mismatch error whenever [[IMAGE:cf41183a839418fe_11_41]] is a GPU.
                    Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
                  4. [[IMAGE:cf41183a839418fe_11_42]] is the wrong loss function for a 10-class classification problem.
                    Source diagram or notation

                  A published solution is not available for this question yet.

                  Question 15 MCQ · 1.0 marks

                  Assume the [[IMAGE:cf41183a839418fe_9_33]] architecture itself has already been corrected so that its classifier input size matches the feature output. Focus only on the training/checkpoint/inference pipeline below. [[IMAGE:cf41183a839418fe_10_34]] Based on the above data, answer the given subquestions.
                  Suppose the device-mismatch bug is fixed, but [[IMAGE:cf41183a839418fe_11_43]] is still never called before inference. Given that [[IMAGE:cf41183a839418fe_11_44]] uses BatchNorm, what is the most likely consequence?
                  Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
                  1. No effect — BatchNorm behaves identically in train and eval mode.
                  2. BatchNorm continues to use statistics from the current inference batch instead of the stored running statistics, so predictions can differ from proper evaluation-mode inference and can depend on batch composition.
                  3. The model automatically switches to eval mode whenever [[IMAGE:cf41183a839418fe_11_45]] is used.
                    Source diagram or notation
                  4. The prediction will be correct but computed twice as slowly.

                  A published solution is not available for this question yet.

                  Question 16 NAT · 1.0 marks

                  Assume the [[IMAGE:cf41183a839418fe_9_33]] architecture itself has already been corrected so that its classifier input size matches the feature output. Focus only on the training/checkpoint/inference pipeline below. [[IMAGE:cf41183a839418fe_10_34]] Based on the above data, answer the given subquestions.
                  In the fully corrected pipeline, if [[IMAGE:cf41183a839418fe_11_46]] yields 20 batches per epoch over 5 epochs, how many total times should [[IMAGE:cf41183a839418fe_11_47]] be called?
                  Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

                    A published solution is not available for this question yet.

                    Question 17 NAT · 2.0 marks

                    A convolutional layer receives a feature map with **16 input channels** and applies **32 filters**, each of spatial size **5 × 5**. Following the "generic level of CNN" view (a filter spans all input channels), calculate the total number of **weights (excluding bias)** in this layer.

                      A published solution is not available for this question yet.

                      Question 18 NAT · 2.0 marks

                      A convolutional layer is applied to a **32 × 32** input using a **5 × 5** filter with **zero-padding of 2** and **stride 2**. Using the standard convolution output-size formula, calculate the **output width (one** **spatial dimension)**. The height is the same; enter one integer only.

                        A published solution is not available for this question yet.

                        Question 19 NAT · 2.0 marks

                        The final classification head applies a **softmax** over the pre-activation scores (logits). For a 3-class problem, the logits produced for a particular image are **[3, 1, 0]**. Compute the softmax **probability** **assigned to the first class** (round to 2 decimal places).

                          A published solution is not available for this question yet.

                          Question 20 NAT · 2.0 marks

                          Following the VGG design philosophy of replacing a single large-kernel convolution with a cascade of small **3 × 3, stride-1** convolutions, how many such 3 × 3 layers must be **stacked** so that the resulting stack has the same effective (one-dimensional) receptive field as a **single 11 × 11** **convolution** (the kernel size used in AlexNet's first layer)?

                            A published solution is not available for this question yet.

                            Question 21 NAT · 2.0 marks

                            A YOLOv1-style detector divides the input image into an **S = 14** grid. Each grid cell predicts **B = 3** bounding boxes (each with [[IMAGE:cf41183a839418fe_13_48]] ) and a single set of class scores over **C =** **25** classes. Compute the **total number of scalar values** in the final output tensor of the network.
                            Source diagram or notation

                              A published solution is not available for this question yet.

                              Question 22 NAT · 2.0 marks

                              The decoder of a depth-estimation network upsamples feature maps using a **transposed** **convolution** ("up-convolution") with **kernel size 4, stride 2, and padding 1**. If the input feature map has spatial size **30 × 30**, what is the output height? The width is the same; enter one integer only.

                                A published solution is not available for this question yet.

                                Question 23 NAT · 2.0 marks

                                In the original U-Net used for dense prediction (and adaptable to depth estimation), each encoder stage applies **two successive 3×3 convolutions with no padding (stride 1)**, followed by a **2×2** **max-pooling with stride 2**. If the input feature map to one such encoder stage has spatial size **140 × 140**, what is the output height of the feature map **after both convolutions and the max-** **pooling** of that stage? The width is the same; enter one integer only.

                                  A published solution is not available for this question yet.

                                  Question 24 NAT · 2.0 marks

                                  The SRGAN discriminator is a stack of **8 convolutional layers**, all using **3 × 3 kernels with** **padding 1**, and strides alternating as **s = 1, 2, 1, 2, 1, 2, 1, 2** across the eight layers (the stride-1 layers preserve spatial size; the stride-2 layers downsample). If the input image is **96 × 96**, what is the output height of the feature map **after the last (8th)** **convolutional layer**, just before the dense layers? The width is the same; enter one integer only.

                                    A published solution is not available for this question yet.

                                    Question 25 NAT · 2.0 marks

                                    Efficient restoration networks (e.g., LaKDNet) replace standard convolutions with **depth-wise** **separable convolutions** (a depth-wise convolution followed by a point-wise 1×1 convolution). An input feature map of size **32 × 32 × 16** (16 channels) is processed by a depth-wise separable convolution that uses **3 × 3** depth-wise kernels and produces **32 output channels**, with stride 1 and padding 1 (so the spatial size is preserved). Calculate the **total number of multiplications** for the full depth-wise separable operation (depth- wise + point-wise). Give the answer in millions (e.g., if 1,234,567 write 1.23).

                                      A published solution is not available for this question yet.

                                      Question 26 MSQ · 2.0 marks

                                      Compared with the sigmoid activation, which of the following are **correct** reasons for preferring the **ReLU** activation in deep CNNs?
                                      1. For positive inputs, ReLU does not saturate, so it avoids the gradient-killing that occurs in the flat tails of the sigmoid.
                                      2. ReLU is preferred because it bounds every activation to the interval [0, 1], preventing activations from growing large.
                                      3. ReLU is cheaper to evaluate — a simple thresholding at zero — whereas sigmoid/tanh require expensive exponentials.
                                      4. The sigmoid's derivative peaks at only **0.25**, so multiplying many such small factors across layers pushes gradients toward zero (vanishing gradient).

                                      A published solution is not available for this question yet.

                                      Question 27 MSQ · 2.0 marks

                                      VGGNet replaces large convolution kernels with stacks of small 3 × 3 convolutions. Which of the following statements correctly describe the **advantages** of this design choice?
                                      1. A stack of three 3 × 3 conv layers (stride 1) has the same effective receptive field as one 7 × 7 conv layer.
                                      2. For C channels per layer, three stacked 3 × 3 layers use fewer parameters than a single 7 × 7 layer.
                                      3. A single 3 × 3 convolution by itself already spans the same receptive field as a 7 × 7 convolution, so stacking is unnecessary.
                                      4. Inserting a ReLU non-linearity between successive 3 × 3 layers makes the composite function more discriminative than a single linear projection over the same region.

                                      A published solution is not available for this question yet.

                                      Question 28 MSQ · 2.0 marks

                                      Which of the following statements correctly describe the **Region Proposal Network (RPN)** in Faster R-CNN?
                                      1. The RPN is fully convolutional and **shares** its convolutional feature maps with the downstream detection network.
                                      2. The RPN relies on an **external Selective Search** module to generate its region proposals.
                                      3. At each sliding-window location, the RPN predicts objectness scores and box refinements **relative to k predefined anchor boxes**.
                                      4. Anchors of multiple **scales and aspect ratios** allow the RPN to propose regions for objects of very different sizes and shapes.

                                      A published solution is not available for this question yet.

                                      Question 29 MSQ · 2.0 marks

                                      Consider the multi-task loss used to train the original **YOLOv1** model. Which of the following statements about it are correct?
                                      1. The loss predicts the **square roots** of width and height so that a fixed absolute error is penalized **more** for small boxes than for large boxes.
                                      2. The hyperparameters are set to **λ_coord = 5** (upweighting localization error) and **λ_noobj = 0.5** (downweighting confidence error for cells with no object).
                                      3. The **classification** term for a cell is only penalized when an object is actually present in that grid cell.
                                      4. For the predictor "responsible" for an object, the target **confidence** value is fixed at 1, independent of the box's IoU with the ground truth.

                                      A published solution is not available for this question yet.

                                      Question 30 MSQ · 2.0 marks

                                      Regarding the **channel-attention module in CBAM**, which of the following are correct?
                                      1. It exploits the **inter-channel** relationship of features, learning "what" is meaningful by re-weighting feature channels.
                                      2. It produces a **spatial** attention map of size H × W × 1 that highlights where to focus, discarding all channel information.
                                      3. The spatial dimensions of the input feature map are squeezed using **both** **average-pooling and max-pooling**, producing two channel descriptors.
                                      4. The pooled descriptors are passed through a **shared MLP**, combined, and a sigmoid produces the channel attention weights that rescale the input feature map.

                                      A published solution is not available for this question yet.

                                      Question 31 MSQ · 2.0 marks

                                      SRGAN performs photo-realistic super-resolution using a GAN. Which of the following statements are correct?
                                      1. Its **perceptual loss** combines a **content loss** and an **adversarial loss**, rather than relying on pixel-wise MSE alone.
                                      2. The **content loss** is computed on **feature maps of a pre-trained VGG** **network** (perceptual similarity), not purely in pixel space.
                                      3. Purely MSE-optimized solutions tend to be **overly smooth** because they approximate a pixel-wise average of many plausible HR solutions, losing high-frequency texture.
                                      4. The adversarial term makes SRGAN maximize PSNR, which is why it always achieves a **higher PSNR** than the MSE-based SRResNet.

                                      A published solution is not available for this question yet.

                                      Question 32 MCQ · 2.0 marks

                                      In the **multi-scale deep network** for single-image depth estimation (Eigen et al.), the architecture is split into a **coarse network** and a **fine network**. Which statement best explains their respective roles?
                                      1. The coarse network uses a large receptive field (via deeper layers and fully- connected layers) to predict the global scene depth, while the fine network refines this prediction locally using input-image details.
                                      2. The fine network predicts the global scene structure, and the coarse network only sharpens edges.
                                      3. Both networks are architecturally identical and their outputs are simply averaged (ensembled).
                                      4. The coarse network processes each pixel independently and therefore ignores global scene context.

                                      A published solution is not available for this question yet.