MauryaHub PYQ Practice

cs3004_2026T2_Q2_NA.pdf

Deep Learning · Quiz 2 · May 2026

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 2 MCQ · 3.0 marks

[[IMAGE:897d6601db10d671_2_2]]
Source diagram or notation
  1. Bias will decrease significantly, while variance remains approximately the same.
  2. Variance will tend to decrease, while bias remains approximately the same.
  3. Both bias and variance will eventually become zero.
  4. Increasing the amount of training data has no effect on either the bias or the variance.

A published solution is not available for this question yet.

Question 3 MCQ · 3.0 marks

[[IMAGE:897d6601db10d671_2_3]]
Source diagram or notation
  1. [[IMAGE:897d6601db10d671_3_4]] only
    Source diagram or notation
  2. [[IMAGE:897d6601db10d671_3_5]]
    Source diagram or notation
  3. [[IMAGE:897d6601db10d671_3_6]]
    Source diagram or notation
  4. At all iterations

A published solution is not available for this question yet.

Question 4 NAT · 3.0 marks

[[IMAGE:897d6601db10d671_3_7]] Based on the above data, answer the given subquestions.
Compute the effective learning rate [[IMAGE:897d6601db10d671_3_8]] used at iteration [[IMAGE:897d6601db10d671_3_9]] . Enter the answer correct to four decimal places.
Source diagram or notationSource diagram or notationSource diagram or notation

    A published solution is not available for this question yet.

    Question 5 MCQ · 3.0 marks

    [[IMAGE:897d6601db10d671_3_7]] Based on the above data, answer the given subquestions.
    [[IMAGE:897d6601db10d671_4_10]]
    Source diagram or notationSource diagram or notation
    1. It would increase from iteration 1 to iteration 3.
    2. It would decrease from iteration 1 to iteration 3.
    3. It would stay constant at [[IMAGE:897d6601db10d671_4_11]] .
      Source diagram or notation
    4. It would decrease toward zero.

    A published solution is not available for this question yet.

    Question 6 MCQ · 3.0 marks

    [[IMAGE:897d6601db10d671_4_12]] Based on the above data, answer the given subquestions.
    [[IMAGE:897d6601db10d671_4_13]]
    Source diagram or notationSource diagram or notation
    1. [[IMAGE:897d6601db10d671_5_14]]
      Source diagram or notation
    2. [[IMAGE:897d6601db10d671_5_15]]
      Source diagram or notation
    3. [[IMAGE:897d6601db10d671_5_16]]
      Source diagram or notation
    4. [[IMAGE:897d6601db10d671_5_17]]
      Source diagram or notation

    A published solution is not available for this question yet.

    Question 7 MCQ · 3.0 marks

    [[IMAGE:897d6601db10d671_4_12]] Based on the above data, answer the given subquestions.
    [[IMAGE:897d6601db10d671_5_18]]
    Source diagram or notationSource diagram or notation
    1. 0
    2. [[IMAGE:897d6601db10d671_5_19]]
      Source diagram or notation
    3. [[IMAGE:897d6601db10d671_5_20]]
      Source diagram or notation
    4. It diverges to [[IMAGE:897d6601db10d671_5_21]]
      Source diagram or notation

    A published solution is not available for this question yet.

    Question 8 NAT · 3.0 marks

    A fully connected hidden layer has [[IMAGE:897d6601db10d671_5_22]] units and uses dropout with retention probability [[IMAGE:897d6601db10d671_5_23]] for each unit (units are dropped independently). The sub-questions are independent of each other. Based on the above data, answer the given subquestions.
    What is the expected number of units retained (active) in a single forward pass during training?
    Source diagram or notationSource diagram or notation

      A published solution is not available for this question yet.

      Question 9 MCQ · 3.0 marks

      A fully connected hidden layer has [[IMAGE:897d6601db10d671_5_22]] units and uses dropout with retention probability [[IMAGE:897d6601db10d671_5_23]] for each unit (units are dropped independently). The sub-questions are independent of each other. Based on the above data, answer the given subquestions.
      How many distinct ''thinned'' sub-networks are theoretically possible for this layer?
      Source diagram or notationSource diagram or notation
      1. [[IMAGE:897d6601db10d671_6_24]]
        Source diagram or notation
      2. [[IMAGE:897d6601db10d671_6_25]]
        Source diagram or notation
      3. [[IMAGE:897d6601db10d671_6_26]]
        Source diagram or notation
      4. [[IMAGE:897d6601db10d671_6_27]]
        Source diagram or notation

      A published solution is not available for this question yet.

      Question 10 NAT · 2.0 marks

      [[IMAGE:897d6601db10d671_6_28]]
      Compute the pre-activation value [[IMAGE:897d6601db10d671_7_29]] .
      Source diagram or notationSource diagram or notation

        A published solution is not available for this question yet.

        Question 11 NAT · 2.0 marks

        [[IMAGE:897d6601db10d671_6_28]]
        Given your value of [[IMAGE:897d6601db10d671_7_30]] from the previous part, compute [[IMAGE:897d6601db10d671_7_31]] .Enter the answer correct to one decimal place.
        Source diagram or notationSource diagram or notationSource diagram or notation

          A published solution is not available for this question yet.

          Question 12 MCQ · 2.0 marks

          [[IMAGE:897d6601db10d671_6_28]]
          Which of the following is the primary motivation for using Leaky ReLU instead of standard ReLU in this setting?
          Source diagram or notation
          1. To make the activation function computationally cheaper than ReLU.
          2. To bound the output strictly within [[IMAGE:897d6601db10d671_7_32]] .
            Source diagram or notation
          3. To allow a small, non-zero gradient to flow through when [[IMAGE:897d6601db10d671_7_33]] , avoiding permanently dead neurons.
            Source diagram or notation
          4. To make the activation zero-centered exactly like [[IMAGE:897d6601db10d671_7_34]] .
            Source diagram or notation

          A published solution is not available for this question yet.

          Question 13 MCQ · 2.0 marks

          [[IMAGE:897d6601db10d671_8_35]]
          For a training example where [[IMAGE:897d6601db10d671_8_36]] , which of the following best describes [[IMAGE:897d6601db10d671_8_37]] ?
          Source diagram or notationSource diagram or notationSource diagram or notation
          1. [[IMAGE:897d6601db10d671_8_38]] ; the largest value [[IMAGE:897d6601db10d671_8_39]] can attain
            Source diagram or notationSource diagram or notation
          2. [[IMAGE:897d6601db10d671_8_40]] ; the neuron has saturated
            Source diagram or notation
          3. [[IMAGE:897d6601db10d671_8_41]]
            Source diagram or notation
          4. [[IMAGE:897d6601db10d671_8_42]] is undefined at this point
            Source diagram or notation

          A published solution is not available for this question yet.

          Question 14 MCQ · 2.0 marks

          [[IMAGE:897d6601db10d671_8_35]]
          Suppose instead the same layer used ReLU, [[IMAGE:897d6601db10d671_8_43]] , and [[IMAGE:897d6601db10d671_8_44]] for most training examples but a large negative bias update later drives [[IMAGE:897d6601db10d671_8_45]] for all future inputs. What is [[IMAGE:897d6601db10d671_8_46]] in that regime, and what training consequence follows?
          Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
          1. [[IMAGE:897d6601db10d671_8_47]] ; the neuron continues to update normally
            Source diagram or notation
          2. [[IMAGE:897d6601db10d671_8_48]] ; the neuron stops updating and may stay ''dead'' permanently
            Source diagram or notation
          3. [[IMAGE:897d6601db10d671_9_49]] oscillates between 0 and 1 depending on the input
            Source diagram or notation
          4. [[IMAGE:897d6601db10d671_9_50]] ; the neuron updates at half the normal rate
            Source diagram or notation

          A published solution is not available for this question yet.

          Question 15 MSQ · 2.0 marks

          [[IMAGE:897d6601db10d671_9_51]] Based on the above data, answer the given subquestions.
          Which filter weights can receive a non-zero gradient during backpropagation? (Select all that apply.)
          Source diagram or notation
          1. [[IMAGE:897d6601db10d671_9_52]]
            Source diagram or notation
          2. [[IMAGE:897d6601db10d671_9_53]]
            Source diagram or notation
          3. [[IMAGE:897d6601db10d671_9_54]]
            Source diagram or notation
          4. [[IMAGE:897d6601db10d671_9_55]]
            Source diagram or notation

          A published solution is not available for this question yet.

          Question 16 NAT · 3.0 marks

          [[IMAGE:897d6601db10d671_9_51]] Based on the above data, answer the given subquestions.
          [[IMAGE:897d6601db10d671_10_56]] Suppose at each of the four output cells, and the current weights satisfy [[IMAGE:897d6601db10d671_10_57]] . [[IMAGE:897d6601db10d671_10_58]] Compute .
          Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

            A published solution is not available for this question yet.

            Question 17 NAT · 3.0 marks

            The input volume to a convolutional layer of a CNN has shape [[IMAGE:897d6601db10d671_10_59]] , where [[IMAGE:897d6601db10d671_10_60]] is the depth. Based on the above data, answer the given subquestions.
            You want a convolutional layer with a [[IMAGE:897d6601db10d671_10_61]] kernel and stride [[IMAGE:897d6601db10d671_10_62]] to produce an output whose width and height equal those of the input. What should be the padding [[IMAGE:897d6601db10d671_10_63]] ?
            Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

              A published solution is not available for this question yet.

              Question 18 MCQ · 2.0 marks

              The input volume to a convolutional layer of a CNN has shape [[IMAGE:897d6601db10d671_10_59]] , where [[IMAGE:897d6601db10d671_10_60]] is the depth. Based on the above data, answer the given subquestions.
              Suppose you instead convolve the [[IMAGE:897d6601db10d671_11_64]] input with a [[IMAGE:897d6601db10d671_11_65]] kernel using padding [[IMAGE:897d6601db10d671_11_66]] , stride [[IMAGE:897d6601db10d671_11_67]] , and [[IMAGE:897d6601db10d671_11_68]] filters. Which of these is the shape of the output volume?
              Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
              1. [[IMAGE:897d6601db10d671_11_69]]
                Source diagram or notation
              2. [[IMAGE:897d6601db10d671_11_70]]
                Source diagram or notation
              3. [[IMAGE:897d6601db10d671_11_71]]
                Source diagram or notation
              4. [[IMAGE:897d6601db10d671_11_72]]
                Source diagram or notation
              5. [[IMAGE:897d6601db10d671_11_73]]
                Source diagram or notation

              A published solution is not available for this question yet.

              Question 19 NAT · 2.0 marks

              The input volume to a convolutional layer of a CNN has shape [[IMAGE:897d6601db10d671_10_59]] , where [[IMAGE:897d6601db10d671_10_60]] is the depth. Based on the above data, answer the given subquestions.
              Find the total number of parameters (weights [[IMAGE:897d6601db10d671_11_74]] biases) associated with the convolutional layer of previous question.
              Source diagram or notationSource diagram or notationSource diagram or notation

                A published solution is not available for this question yet.

                Question 20 MSQ · 4.0 marks

                Which of the following statements are TRUE about the adaptive optimizers? Select all that apply.
                1. In AdaGrad, the effective learning rate [[IMAGE:897d6601db10d671_12_75]] decays slowly for the dense features.
                  Source diagram or notation
                2. In AdaGrad, the effective learning rate [[IMAGE:897d6601db10d671_12_76]] decays slowly for the sparse features.
                  Source diagram or notation
                3. In RMSProp, the effective learning rate [[IMAGE:897d6601db10d671_12_77]] is guaranteed to be non- increasing across iterations.
                  Source diagram or notation
                4. [[IMAGE:897d6601db10d671_12_78]]
                  Source diagram or notation
                5. [[IMAGE:897d6601db10d671_12_79]]
                  Source diagram or notation

                A published solution is not available for this question yet.