MauryaHub PYQ Practice

cs3004_2026T2_ET_FN.pdf

Deep Learning · End Term · May 2026 FN

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 2 NAT · 2.0 marks

A neuron takes an input vector [[IMAGE:8b451edb9ea00ccd_2_2]] , weight vector [[IMAGE:8b451edb9ea00ccd_2_3]] , and bias [[IMAGE:8b451edb9ea00ccd_2_4]] . The neuron uses Parametric ReLU activation function parameterized by [[IMAGE:8b451edb9ea00ccd_2_5]] : [[IMAGE:8b451edb9ea00ccd_2_6]] Calculate the exact activation output [[IMAGE:8b451edb9ea00ccd_2_7]] for this neuron. Enter the answer correct to one decimal place.
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

    A published solution is not available for this question yet.

    Question 3 NAT · 2.0 marks

    A satellite image of size [[IMAGE:8b451edb9ea00ccd_2_8]] is given as input to a convolutional layer. Suppose the layer uses [[IMAGE:8b451edb9ea00ccd_2_9]] kernels, each of size [[IMAGE:8b451edb9ea00ccd_2_10]] , and produces an output of size [[IMAGE:8b451edb9ea00ccd_2_11]] . Assume zero padding ( [[IMAGE:8b451edb9ea00ccd_3_12]] ). What is the value of the stride [[IMAGE:8b451edb9ea00ccd_3_13]] ?
    Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

      A published solution is not available for this question yet.

      Question 4 NAT · 3.0 marks

      Suppose we want to minimize the following function: [[IMAGE:8b451edb9ea00ccd_3_14]] Let the initial point [[IMAGE:8b451edb9ea00ccd_3_15]] . Based on the above data, answer the given subquestions.
      Using vanilla gradient descent with [[IMAGE:8b451edb9ea00ccd_3_16]] , compute the iterate [[IMAGE:8b451edb9ea00ccd_3_17]] after one update, and enter your answer as [[IMAGE:8b451edb9ea00ccd_3_18]] . Enter the answer correct to one decimal place.
      Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

        A published solution is not available for this question yet.

        Question 5 MCQ · 2.0 marks

        Suppose we want to minimize the following function: [[IMAGE:8b451edb9ea00ccd_3_14]] Let the initial point [[IMAGE:8b451edb9ea00ccd_3_15]] . Based on the above data, answer the given subquestions.
        Along which coordinate does gradient descent with [[IMAGE:8b451edb9ea00ccd_4_19]] fail to converge?
        Source diagram or notationSource diagram or notationSource diagram or notation
        1. [[IMAGE:8b451edb9ea00ccd_4_20]] -coordinate
          Source diagram or notation
        2. [[IMAGE:8b451edb9ea00ccd_4_21]] -coordinate
          Source diagram or notation
        3. Both coordinates diverge equally.
        4. Neither; both converge at the same rate.

        A published solution is not available for this question yet.

        Question 6 NAT · 3.0 marks

        Suppose we want to minimize the following function: [[IMAGE:8b451edb9ea00ccd_3_14]] Let the initial point [[IMAGE:8b451edb9ea00ccd_3_15]] . Based on the above data, answer the given subquestions.
        [[IMAGE:8b451edb9ea00ccd_4_22]]
        Source diagram or notationSource diagram or notationSource diagram or notation

          A published solution is not available for this question yet.

          Question 7 MCQ · 2.0 marks

          Suppose we want to minimize the following function: [[IMAGE:8b451edb9ea00ccd_3_14]] Let the initial point [[IMAGE:8b451edb9ea00ccd_3_15]] . Based on the above data, answer the given subquestions.
          How does AdaGrad address the behavior observed with vanilla gradient descent?
          Source diagram or notationSource diagram or notation
          1. The learning rate is reduced 25 times from the initial learning rate 0.1 along the [[IMAGE:8b451edb9ea00ccd_5_23]] -direction.
            Source diagram or notation
          2. The learning rate is reduced 25 times from the initial learning rate 0.1 along the [[IMAGE:8b451edb9ea00ccd_5_24]] -direction.
            Source diagram or notation
          3. The learning rate is reduced 5 times from the initial learning rate 0.1 along the [[IMAGE:8b451edb9ea00ccd_5_25]] -direction.
            Source diagram or notation
          4. The learning rate is reduced [[IMAGE:8b451edb9ea00ccd_5_26]] times from the initial learning rate [[IMAGE:8b451edb9ea00ccd_5_27]] along both the [[IMAGE:8b451edb9ea00ccd_5_28]] and [[IMAGE:8b451edb9ea00ccd_5_29]] directions.
            Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

          A published solution is not available for this question yet.

          Question 8 MSQ · 2.0 marks

          Consider a feedforward neural network with [[IMAGE:8b451edb9ea00ccd_5_30]] hidden layers. The pre-activation at layer [[IMAGE:8b451edb9ea00ccd_5_31]] is [[IMAGE:8b451edb9ea00ccd_5_32]] and the activation is [[IMAGE:8b451edb9ea00ccd_5_33]] , where [[IMAGE:8b451edb9ea00ccd_5_34]] is the activation function. Suppose every hidden-layer activation is the identity function [[IMAGE:8b451edb9ea00ccd_5_35]] , and the output activation is linear. Which of the following statements are correct?
          Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
          1. The entire network computes a function of the form [[IMAGE:8b451edb9ea00ccd_5_36]] for some single matrix [[IMAGE:8b451edb9ea00ccd_5_37]] and vector [[IMAGE:8b451edb9ea00ccd_5_38]] , regardless of the number of hidden layers.
            Source diagram or notationSource diagram or notationSource diagram or notation
          2. Adding more hidden layers with identity activations increases the representational capacity beyond that of a single-layer linear model.
          3. Replacing the identity activation with a nonlinear function such as sigmoid enables the network to learn nonlinear decision boundaries.
          4. With identity activations, the network can still learn arbitrary decision boundaries provided enough hidden neurons are used.

          A published solution is not available for this question yet.

          Question 9 NAT · 3.0 marks

          A multi-head attention block uses [[IMAGE:8b451edb9ea00ccd_6_39]] , [[IMAGE:8b451edb9ea00ccd_6_40]] heads, and [[IMAGE:8b451edb9ea00ccd_6_41]] . Counting only the weight matrices [[IMAGE:8b451edb9ea00ccd_6_42]] (each of dimension [[IMAGE:8b451edb9ea00ccd_6_43]] ) for every head, plus the output projection [[IMAGE:8b451edb9ea00ccd_6_44]] (dimension [[IMAGE:8b451edb9ea00ccd_6_45]] ), and ignoring all biases, compute the total number of parameters in this multi-head attention block.
          Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

            A published solution is not available for this question yet.

            Question 10 MCQ · 3.0 marks

            In a self-attention mechanism, each input vector has dimension [[IMAGE:8b451edb9ea00ccd_6_46]] . The queries, keys, and values are computed using separate linear transformations, each mapping from [[IMAGE:8b451edb9ea00ccd_6_47]] . Suppose there are [[IMAGE:8b451edb9ea00ccd_6_48]] input vectors. Based on the above data, answer the given subquestions.
            How many learnable parameters (including both weights and biases) are used to compute the queries, keys, and values?
            Source diagram or notationSource diagram or notationSource diagram or notation
            1. [[IMAGE:8b451edb9ea00ccd_7_49]]
              Source diagram or notation
            2. [[IMAGE:8b451edb9ea00ccd_7_50]]
              Source diagram or notation
            3. [[IMAGE:8b451edb9ea00ccd_7_51]]
              Source diagram or notation
            4. [[IMAGE:8b451edb9ea00ccd_7_52]]
              Source diagram or notation

            A published solution is not available for this question yet.

            Question 11 NAT · 2.0 marks

            In a self-attention mechanism, each input vector has dimension [[IMAGE:8b451edb9ea00ccd_6_46]] . The queries, keys, and values are computed using separate linear transformations, each mapping from [[IMAGE:8b451edb9ea00ccd_6_47]] . Suppose there are [[IMAGE:8b451edb9ea00ccd_6_48]] input vectors. Based on the above data, answer the given subquestions.
            How many attention weights [[IMAGE:8b451edb9ea00ccd_7_53]] are computed in total if we have [[IMAGE:8b451edb9ea00ccd_7_54]] and [[IMAGE:8b451edb9ea00ccd_7_55]] ?
            Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

              A published solution is not available for this question yet.

              Question 12 MCQ · 2.0 marks

              In a self-attention mechanism, each input vector has dimension [[IMAGE:8b451edb9ea00ccd_6_46]] . The queries, keys, and values are computed using separate linear transformations, each mapping from [[IMAGE:8b451edb9ea00ccd_6_47]] . Suppose there are [[IMAGE:8b451edb9ea00ccd_6_48]] input vectors. Based on the above data, answer the given subquestions.
              How many learnable parameters (including both weights and biases) are there in a fully connected shallow network that maps all [[IMAGE:8b451edb9ea00ccd_7_56]] inputs to all [[IMAGE:8b451edb9ea00ccd_7_57]] outputs?
              Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
              1. [[IMAGE:8b451edb9ea00ccd_7_58]]
                Source diagram or notation
              2. [[IMAGE:8b451edb9ea00ccd_7_59]]
                Source diagram or notation
              3. [[IMAGE:8b451edb9ea00ccd_7_60]]
                Source diagram or notation
              4. [[IMAGE:8b451edb9ea00ccd_7_61]]
                Source diagram or notation

              A published solution is not available for this question yet.

              Question 13 NAT · 2.0 marks

              [[IMAGE:8b451edb9ea00ccd_8_62]] Based on the above data, answer the given subquestions.
              Which token receives the lowest dot-product score [[IMAGE:8b451edb9ea00ccd_8_63]] from query [[IMAGE:8b451edb9ea00ccd_8_64]] ? Enter the token index (1, 2, or 3).
              Source diagram or notationSource diagram or notationSource diagram or notation

                A published solution is not available for this question yet.

                Question 14 NAT · 3.0 marks

                [[IMAGE:8b451edb9ea00ccd_8_62]] Based on the above data, answer the given subquestions.
                Compute the attention weight [[IMAGE:8b451edb9ea00ccd_9_65]] (the weight that query [[IMAGE:8b451edb9ea00ccd_9_66]] places on token 3). Enter the answer correct to three decimal places.
                Source diagram or notationSource diagram or notationSource diagram or notation

                  A published solution is not available for this question yet.

                  Question 15 NAT · 3.0 marks

                  [[IMAGE:8b451edb9ea00ccd_8_62]] Based on the above data, answer the given subquestions.
                  [[IMAGE:8b451edb9ea00ccd_9_67]]
                  Source diagram or notationSource diagram or notation

                    A published solution is not available for this question yet.

                    Question 16 MSQ · 3.0 marks

                    In a vanilla (non-attention) encoder-decoder model, the encoder compresses the entire input sequence into a single fixed-dimensional vector [[IMAGE:8b451edb9ea00ccd_10_68]] , such as [[IMAGE:8b451edb9ea00ccd_10_69]] , where [[IMAGE:8b451edb9ea00ccd_10_70]] is the final hidden state of an RNN encoder. The decoder uses this same vector to generate every output token. In practice, the translation quality of such a model tends to degrade as the length of the source sentence increases. Which of the following best explains this degradation? Select all that apply
                    Source diagram or notationSource diagram or notationSource diagram or notation
                    1. The decoder's softmax layer has a fixed vocabulary size and therefore cannot generate longer sentences.
                    2. The fixed-dimensional vector [[IMAGE:8b451edb9ea00ccd_10_71]] becomes an information bottleneck as the source sentence length increases.
                      Source diagram or notation
                    3. Teacher forcing is valid only for source sentences shorter than the maximum training length.
                    4. Longer source sentences necessarily require a higher-dimensional one-hot representation, which increases the input noise.

                    A published solution is not available for this question yet.

                    Question 17 NAT · 3.0 marks

                    Consider a Skip-gram model with negative sampling. For a positive word-context pair [[IMAGE:8b451edb9ea00ccd_10_72]] , suppose the loss for one training example is [[IMAGE:8b451edb9ea00ccd_10_73]] where [[IMAGE:8b451edb9ea00ccd_10_74]] is a logistic function. Suppose the initial vectors are [[IMAGE:8b451edb9ea00ccd_10_75]] The word vector w is updated using the gradient descent algorithm using a learning rate of [[IMAGE:8b451edb9ea00ccd_10_76]] Based on the above data, answer the given subquestions.
                    Calculate the loss [[IMAGE:8b451edb9ea00ccd_11_77]] for the given positive word-context pair [[IMAGE:8b451edb9ea00ccd_11_78]] and negative sample. Give your answer correct to three decimal places.
                    Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

                      A published solution is not available for this question yet.

                      Question 18 NAT · 3.0 marks

                      Consider a Skip-gram model with negative sampling. For a positive word-context pair [[IMAGE:8b451edb9ea00ccd_10_72]] , suppose the loss for one training example is [[IMAGE:8b451edb9ea00ccd_10_73]] where [[IMAGE:8b451edb9ea00ccd_10_74]] is a logistic function. Suppose the initial vectors are [[IMAGE:8b451edb9ea00ccd_10_75]] The word vector w is updated using the gradient descent algorithm using a learning rate of [[IMAGE:8b451edb9ea00ccd_10_76]] Based on the above data, answer the given subquestions.
                      The word vector w is updated using the gradient descent algorithm. Using gradient descent, obtain the updated word vector [[IMAGE:8b451edb9ea00ccd_11_79]] , and calculate the cosine similarity between [[IMAGE:8b451edb9ea00ccd_11_80]] and [[IMAGE:8b451edb9ea00ccd_11_81]] . Enter the answer correct to three decimal places.
                      Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

                        A published solution is not available for this question yet.

                        Question 19 NAT · 3.0 marks

                        Consider a corpus containing [[IMAGE:8b451edb9ea00ccd_11_82]] total words. Based on this, answer the given subquestions:
                        A word [[IMAGE:8b451edb9ea00ccd_12_83]] occurs [[IMAGE:8b451edb9ea00ccd_12_84]] times, a context [[IMAGE:8b451edb9ea00ccd_12_85]] occurs [[IMAGE:8b451edb9ea00ccd_12_86]] times, and the pair [[IMAGE:8b451edb9ea00ccd_12_87]] co-occurs [[IMAGE:8b451edb9ea00ccd_12_88]] times. Compute [[IMAGE:8b451edb9ea00ccd_12_89]] . Use natural log and enter the answer correct to two decimal places.
                        Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

                          A published solution is not available for this question yet.

                          Question 20 NAT · 1.0 marks

                          Consider a corpus containing [[IMAGE:8b451edb9ea00ccd_11_82]] total words. Based on this, answer the given subquestions:
                          A second pair [[IMAGE:8b451edb9ea00ccd_12_90]] has [[IMAGE:8b451edb9ea00ccd_12_91]] Suppose its PMI is [[IMAGE:8b451edb9ea00ccd_12_92]] . Under the PPMI (positive PMI) transformation, what value is stored for this pair?
                          Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

                            A published solution is not available for this question yet.

                            Question 21 MCQ · 3.0 marks

                            A neural network is trained using mini-batch gradient descent. Suppose the mini-batch size is increased from [[IMAGE:8b451edb9ea00ccd_13_93]] to [[IMAGE:8b451edb9ea00ccd_13_94]] , while the learning rate and training dataset remain unchanged. Which of the following statements is correct?
                            Source diagram or notationSource diagram or notation
                            1. Each update uses a gradient estimate with higher variance, while the number of updates per epoch decreases.
                            2. Each update uses a gradient estimate with lower variance, while the number of updates per epoch decreases.
                            3. The number of parameter updates per epoch increases by a factor of four, leading to faster convergence.
                            4. The larger batch size guarantees convergence to a better local minimum.

                            A published solution is not available for this question yet.