MauryaHub PYQ Practice

cs3004_2025T3_ET_FN.pdf

Deep Learning · End Term · Sep 2025 FN

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 2 MCQ · 2.0 marks

What is the primary advantage of word embeddings compared to one-hot encoding?
  1. Simpler implementation
  2. Captures semantic relationships
  3. They automatically correct spelling errors
  4. Word embeddings are sparser than one-hot encoding

A published solution is not available for this question yet.

Question 3 MCQ · 2.0 marks

What is the primary advantage of Transformer encoders over RNN encoders for long sequences?
  1. Lower memory usage
  2. Parallelization during training
  3. Smaller model size
  4. Simpler architecture

A published solution is not available for this question yet.

Question 4 MSQ · 3.0 marks

Suppose that we implement an OR function with two boolean inputs using the McCulloch Pitts (MP) neuron. Assume that the neuron does not have any inhibitory input. Which of the following option(s) represents correct [[IMAGE:64b1fdfafe1feffe_3_2]] and the Number of Correctly Classified (NCC) data points for various values of threshold [[IMAGE:64b1fdfafe1feffe_3_3]] . [[IMAGE:64b1fdfafe1feffe_3_4]]
Source diagram or notationSource diagram or notationSource diagram or notation
  1. [[IMAGE:64b1fdfafe1feffe_3_5]] , NCC = 4
    Source diagram or notation
  2. [[IMAGE:64b1fdfafe1feffe_3_6]] , NCC = 4
    Source diagram or notation
  3. [[IMAGE:64b1fdfafe1feffe_3_7]] , NCC = 3
    Source diagram or notation
  4. [[IMAGE:64b1fdfafe1feffe_3_8]] , NCC = 1
    Source diagram or notation
  5. [[IMAGE:64b1fdfafe1feffe_3_9]] , NCC = 2
    Source diagram or notation
  6. [[IMAGE:64b1fdfafe1feffe_3_10]] , NCC = 0
    Source diagram or notation

A published solution is not available for this question yet.

Question 5 MSQ · 2.0 marks

Which of the following optimizers use momentum (or momentum-like mechanisms) in updating gradients?
  1. NAG
  2. AdaGrad
  3. RMSProp
  4. AdaDelta
  5. ADAM

A published solution is not available for this question yet.

Question 6 MSQ · 2.0 marks

Which of the following statements are true for the saturated neurons problem? Select all that apply.
  1. ReLU activation function is preferred over tanh, as it is less likely to saturate.
  2. Tanh is preferred over sigmoid as it solves the problem of neuron saturation.
  3. Leaky ReLU mitigates the problem of neuron saturation compared to ReLU.
  4. If a neuron saturates in a network, the weights do not get updated further.

A published solution is not available for this question yet.

Question 7 MSQ · 3.0 marks

You are training a deep feedforward neural network (100 layers) for binary classification with a sigmoid output with tanh activation in some hidden layers and ReLU activations in some hidden layers. You notice that gradients in some layers vanish early, preventing those weights from updating, even though the network hasn’t converged. Which of the following fixes could help?
  1. Increase the size of your training set
  2. Replace ReLU activations with leaky ReLUs
  3. Use appropriate weight initialization

A published solution is not available for this question yet.

Question 8 NAT · 3.0 marks

A transformer with a single self-attention layer has batch size [[IMAGE:64b1fdfafe1feffe_4_11]] , context length (sequence length) [[IMAGE:64b1fdfafe1feffe_4_12]] , and number of heads [[IMAGE:64b1fdfafe1feffe_4_13]] . How many attention scores are computed in total across two heads and all sequences for one forward pass?
Source diagram or notationSource diagram or notationSource diagram or notation

    A published solution is not available for this question yet.

    Question 9 NAT · 3.0 marks

    [[IMAGE:64b1fdfafe1feffe_5_14]]
    Source diagram or notation

      A published solution is not available for this question yet.

      Question 10 NAT · 3.0 marks

      Given the input matrix [[IMAGE:64b1fdfafe1feffe_6_15]] and kernel [[IMAGE:64b1fdfafe1feffe_6_16]] : [[IMAGE:64b1fdfafe1feffe_6_17]] • Perform convolution of [[IMAGE:64b1fdfafe1feffe_6_18]] over [[IMAGE:64b1fdfafe1feffe_6_19]] with stride [[IMAGE:64b1fdfafe1feffe_6_20]] and no padding to obtain matrix [[IMAGE:64b1fdfafe1feffe_6_21]] . • Apply average pooling over [[IMAGE:64b1fdfafe1feffe_6_22]] of size [[IMAGE:64b1fdfafe1feffe_6_23]] with stride [[IMAGE:64b1fdfafe1feffe_6_24]] and no padding to obtain [[IMAGE:64b1fdfafe1feffe_6_25]] . • Apply the **ReLU activation** on [[IMAGE:64b1fdfafe1feffe_6_26]] to obtain the final output [[IMAGE:64b1fdfafe1feffe_6_27]] . Given that [[IMAGE:64b1fdfafe1feffe_6_28]] compute [[IMAGE:64b1fdfafe1feffe_6_29]] where [[IMAGE:64b1fdfafe1feffe_6_30]] is the **center element** of the kernel. **Submit your answer correct to two decimal places.**
      Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

        A published solution is not available for this question yet.

        Question 11 NAT · 2.0 marks

        You are designing a Skip-Gram model to learn word embeddings. In this iteration, the model focuses on a specific **center word** and attempts to predict the surrounding **context words** within a window size of [[IMAGE:64b1fdfafe1feffe_7_31]] . The sentence is: **Deep neural networks learn complex data patterns.** We are currently processing the center word: **networks**. Vocabulary = {0: complex, 1: data, 2: deep, 3: learn, 4: models, 5: networks, 6: neural, 7: patterns } The one-hot vector for the center word ( [[IMAGE:64b1fdfafe1feffe_7_32]] ) is a column vector of size [[IMAGE:64b1fdfafe1feffe_7_33]] . [[IMAGE:64b1fdfafe1feffe_7_34]] The model's weight matrices are given below: **Input/Center Weight Matrix (** [[IMAGE:64b1fdfafe1feffe_7_35]] **):** [[IMAGE:64b1fdfafe1feffe_7_36]] (Note: Row 5 corresponds to word index 5 `networks') **Output/Context Weight Matrix (** [[IMAGE:64b1fdfafe1feffe_7_37]] **):** [[IMAGE:64b1fdfafe1feffe_7_38]] For the center word [[IMAGE:64b1fdfafe1feffe_7_39]] , the hidden layer representation is given by: [[IMAGE:64b1fdfafe1feffe_7_40]] (Or simply selecting the row corresponding to the center word). The unnormalized score vector [[IMAGE:64b1fdfafe1feffe_7_41]] for all words in the vocabulary is calculated as: [[IMAGE:64b1fdfafe1feffe_7_42]] Based on the above data, answer the given subquestions.
        In the Skip-gram model with a window size of [[IMAGE:64b1fdfafe1feffe_8_43]] (on each side), how many **unique** input and output pairs will be generated for the following sentence. The sentence is: **deep neural networks** **learn complex data patterns.**
        Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

          A published solution is not available for this question yet.

          Question 12 MCQ · 1.0 marks

          You are designing a Skip-Gram model to learn word embeddings. In this iteration, the model focuses on a specific **center word** and attempts to predict the surrounding **context words** within a window size of [[IMAGE:64b1fdfafe1feffe_7_31]] . The sentence is: **Deep neural networks learn complex data patterns.** We are currently processing the center word: **networks**. Vocabulary = {0: complex, 1: data, 2: deep, 3: learn, 4: models, 5: networks, 6: neural, 7: patterns } The one-hot vector for the center word ( [[IMAGE:64b1fdfafe1feffe_7_32]] ) is a column vector of size [[IMAGE:64b1fdfafe1feffe_7_33]] . [[IMAGE:64b1fdfafe1feffe_7_34]] The model's weight matrices are given below: **Input/Center Weight Matrix (** [[IMAGE:64b1fdfafe1feffe_7_35]] **):** [[IMAGE:64b1fdfafe1feffe_7_36]] (Note: Row 5 corresponds to word index 5 `networks') **Output/Context Weight Matrix (** [[IMAGE:64b1fdfafe1feffe_7_37]] **):** [[IMAGE:64b1fdfafe1feffe_7_38]] For the center word [[IMAGE:64b1fdfafe1feffe_7_39]] , the hidden layer representation is given by: [[IMAGE:64b1fdfafe1feffe_7_40]] (Or simply selecting the row corresponding to the center word). The unnormalized score vector [[IMAGE:64b1fdfafe1feffe_7_41]] for all words in the vocabulary is calculated as: [[IMAGE:64b1fdfafe1feffe_7_42]] Based on the above data, answer the given subquestions.
          Calculate the hidden layer representation vector [[IMAGE:64b1fdfafe1feffe_8_44]] for the center word "networks".
          Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
          1. [[IMAGE:64b1fdfafe1feffe_8_45]]
            Source diagram or notation
          2. [[IMAGE:64b1fdfafe1feffe_8_46]]
            Source diagram or notation
          3. [[IMAGE:64b1fdfafe1feffe_8_47]]
            Source diagram or notation
          4. [[IMAGE:64b1fdfafe1feffe_8_48]]
            Source diagram or notation

          A published solution is not available for this question yet.

          Question 13 NAT · 2.0 marks

          You are designing a Skip-Gram model to learn word embeddings. In this iteration, the model focuses on a specific **center word** and attempts to predict the surrounding **context words** within a window size of [[IMAGE:64b1fdfafe1feffe_7_31]] . The sentence is: **Deep neural networks learn complex data patterns.** We are currently processing the center word: **networks**. Vocabulary = {0: complex, 1: data, 2: deep, 3: learn, 4: models, 5: networks, 6: neural, 7: patterns } The one-hot vector for the center word ( [[IMAGE:64b1fdfafe1feffe_7_32]] ) is a column vector of size [[IMAGE:64b1fdfafe1feffe_7_33]] . [[IMAGE:64b1fdfafe1feffe_7_34]] The model's weight matrices are given below: **Input/Center Weight Matrix (** [[IMAGE:64b1fdfafe1feffe_7_35]] **):** [[IMAGE:64b1fdfafe1feffe_7_36]] (Note: Row 5 corresponds to word index 5 `networks') **Output/Context Weight Matrix (** [[IMAGE:64b1fdfafe1feffe_7_37]] **):** [[IMAGE:64b1fdfafe1feffe_7_38]] For the center word [[IMAGE:64b1fdfafe1feffe_7_39]] , the hidden layer representation is given by: [[IMAGE:64b1fdfafe1feffe_7_40]] (Or simply selecting the row corresponding to the center word). The unnormalized score vector [[IMAGE:64b1fdfafe1feffe_7_41]] for all words in the vocabulary is calculated as: [[IMAGE:64b1fdfafe1feffe_7_42]] Based on the above data, answer the given subquestions.
          We want to predict the immediate **next word** to the word "networks" in the sentence, which is "learn". Calculate the unnormalized score ( [[IMAGE:64b1fdfafe1feffe_8_49]] ) for this specific target word.
          Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

            A published solution is not available for this question yet.

            Question 14 NAT · 2.0 marks

            You are designing a Skip-Gram model to learn word embeddings. In this iteration, the model focuses on a specific **center word** and attempts to predict the surrounding **context words** within a window size of [[IMAGE:64b1fdfafe1feffe_7_31]] . The sentence is: **Deep neural networks learn complex data patterns.** We are currently processing the center word: **networks**. Vocabulary = {0: complex, 1: data, 2: deep, 3: learn, 4: models, 5: networks, 6: neural, 7: patterns } The one-hot vector for the center word ( [[IMAGE:64b1fdfafe1feffe_7_32]] ) is a column vector of size [[IMAGE:64b1fdfafe1feffe_7_33]] . [[IMAGE:64b1fdfafe1feffe_7_34]] The model's weight matrices are given below: **Input/Center Weight Matrix (** [[IMAGE:64b1fdfafe1feffe_7_35]] **):** [[IMAGE:64b1fdfafe1feffe_7_36]] (Note: Row 5 corresponds to word index 5 `networks') **Output/Context Weight Matrix (** [[IMAGE:64b1fdfafe1feffe_7_37]] **):** [[IMAGE:64b1fdfafe1feffe_7_38]] For the center word [[IMAGE:64b1fdfafe1feffe_7_39]] , the hidden layer representation is given by: [[IMAGE:64b1fdfafe1feffe_7_40]] (Or simply selecting the row corresponding to the center word). The unnormalized score vector [[IMAGE:64b1fdfafe1feffe_7_41]] for all words in the vocabulary is calculated as: [[IMAGE:64b1fdfafe1feffe_7_42]] Based on the above data, answer the given subquestions.
            Identify the word in the vocabulary that has the highest probability of being predicted as a context word given the center word 'networks'. Provide the corresponding **integer index (key)** from the vocabulary list as your answer.
            Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

              A published solution is not available for this question yet.

              Question 15 NAT · 2.0 marks

              You are designing a Skip-Gram model to learn word embeddings. In this iteration, the model focuses on a specific **center word** and attempts to predict the surrounding **context words** within a window size of [[IMAGE:64b1fdfafe1feffe_7_31]] . The sentence is: **Deep neural networks learn complex data patterns.** We are currently processing the center word: **networks**. Vocabulary = {0: complex, 1: data, 2: deep, 3: learn, 4: models, 5: networks, 6: neural, 7: patterns } The one-hot vector for the center word ( [[IMAGE:64b1fdfafe1feffe_7_32]] ) is a column vector of size [[IMAGE:64b1fdfafe1feffe_7_33]] . [[IMAGE:64b1fdfafe1feffe_7_34]] The model's weight matrices are given below: **Input/Center Weight Matrix (** [[IMAGE:64b1fdfafe1feffe_7_35]] **):** [[IMAGE:64b1fdfafe1feffe_7_36]] (Note: Row 5 corresponds to word index 5 `networks') **Output/Context Weight Matrix (** [[IMAGE:64b1fdfafe1feffe_7_37]] **):** [[IMAGE:64b1fdfafe1feffe_7_38]] For the center word [[IMAGE:64b1fdfafe1feffe_7_39]] , the hidden layer representation is given by: [[IMAGE:64b1fdfafe1feffe_7_40]] (Or simply selecting the row corresponding to the center word). The unnormalized score vector [[IMAGE:64b1fdfafe1feffe_7_41]] for all words in the vocabulary is calculated as: [[IMAGE:64b1fdfafe1feffe_7_42]] Based on the above data, answer the given subquestions.
              How many total weights from Win and Wout combined will be updated?
              Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

                A published solution is not available for this question yet.

                Question 16 NAT · 2.0 marks

                Suppose you are given three encoder hidden states at time [[IMAGE:64b1fdfafe1feffe_10_50]] : [[IMAGE:64b1fdfafe1feffe_10_51]] The previous decoder hidden state is: [[IMAGE:64b1fdfafe1feffe_10_52]] Given the attention score function: [[IMAGE:64b1fdfafe1feffe_10_53]] where the hyperbolic tangent function is defined as: [[IMAGE:64b1fdfafe1feffe_10_54]] and [[IMAGE:64b1fdfafe1feffe_10_55]] Based on the above data, answer the given subquestions.
                Compute the attention score for hidden state [[IMAGE:64b1fdfafe1feffe_10_56]] using the given function. Submit the final answer correct to two decimal places.
                Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

                  A published solution is not available for this question yet.

                  Question 17 NAT · 4.0 marks

                  Suppose you are given three encoder hidden states at time [[IMAGE:64b1fdfafe1feffe_10_50]] : [[IMAGE:64b1fdfafe1feffe_10_51]] The previous decoder hidden state is: [[IMAGE:64b1fdfafe1feffe_10_52]] Given the attention score function: [[IMAGE:64b1fdfafe1feffe_10_53]] where the hyperbolic tangent function is defined as: [[IMAGE:64b1fdfafe1feffe_10_54]] and [[IMAGE:64b1fdfafe1feffe_10_55]] Based on the above data, answer the given subquestions.
                  Normalize the attention scores using the softmax function to obtain the attention weights [[IMAGE:64b1fdfafe1feffe_11_57]] . Submit [[IMAGE:64b1fdfafe1feffe_11_58]] (i.e., first element of the [[IMAGE:64b1fdfafe1feffe_11_59]] vector). Submit the final answer correct to two decimal places. [[IMAGE:64b1fdfafe1feffe_11_60]]
                  Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

                    A published solution is not available for this question yet.

                    Question 18 NAT · 2.0 marks

                    Suppose you are given three encoder hidden states at time [[IMAGE:64b1fdfafe1feffe_10_50]] : [[IMAGE:64b1fdfafe1feffe_10_51]] The previous decoder hidden state is: [[IMAGE:64b1fdfafe1feffe_10_52]] Given the attention score function: [[IMAGE:64b1fdfafe1feffe_10_53]] where the hyperbolic tangent function is defined as: [[IMAGE:64b1fdfafe1feffe_10_54]] and [[IMAGE:64b1fdfafe1feffe_10_55]] Based on the above data, answer the given subquestions.
                    Calculate the context vector [[IMAGE:64b1fdfafe1feffe_11_61]] : [[IMAGE:64b1fdfafe1feffe_11_62]] Provide the sum of all the elements of [[IMAGE:64b1fdfafe1feffe_11_63]] . Submit the final answer correct to two decimal places.
                    Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

                      A published solution is not available for this question yet.