MauryaHub PYQ Practice

cs3004_2026T1_ET_FN.pdf

Deep Learning · End Term · Jan 2026 FN

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 2 MCQ · 2.0 marks

Consider a two-dimensional image of size [[IMAGE:0bc72138559e0d4e_2_2]] and a square filter of size [[IMAGE:0bc72138559e0d4e_2_3]] . How much padding would you need on each side to keep the output height and width equal to the input height and width? Assume a stride of 1.
Source diagram or notationSource diagram or notation
  1. [[IMAGE:0bc72138559e0d4e_2_4]]
    Source diagram or notation
  2. [[IMAGE:0bc72138559e0d4e_2_5]]
    Source diagram or notation
  3. [[IMAGE:0bc72138559e0d4e_2_6]]
    Source diagram or notation
  4. [[IMAGE:0bc72138559e0d4e_2_7]]
    Source diagram or notation

A published solution is not available for this question yet.

Question 3 MCQ · 2.0 marks

Consider a vanilla Recurrent Neural Network (RNN) and an LSTM network with the same input dimension, hidden dimension, and output dimension. **Statement 1:** The number of trainable parameters in a vanilla RNN is higher than that in an LSTM network. **Statement 2:** An LSTM has a higher number of parameters because it contains multiple gates (input gate, forget gate, output gate, and candidate state), each having separate weight matrices and biases. Choose the correct option from the following.
  1. Statement 1 is true, and Statement 2 is the correct reason.
  2. Statement 1 is true, but Statement 2 is false.
  3. Statement 1 is false, but Statement 2 is true.
  4. Statement 1 is false, and Statement 2 is false.

A published solution is not available for this question yet.

Question 4 MCQ · 3.0 marks

Which of the functions given below has the steepest slope at the point [[IMAGE:0bc72138559e0d4e_3_8]] ?
Source diagram or notation
  1. [[IMAGE:0bc72138559e0d4e_3_9]]
    Source diagram or notation
  2. [[IMAGE:0bc72138559e0d4e_3_10]]
    Source diagram or notation
  3. [[IMAGE:0bc72138559e0d4e_3_11]]
    Source diagram or notation
  4. [[IMAGE:0bc72138559e0d4e_3_12]]
    Source diagram or notation

A published solution is not available for this question yet.

Question 5 MCQ · 3.0 marks

Consider a XOR function for input [[IMAGE:0bc72138559e0d4e_3_13]] , where the points [[IMAGE:0bc72138559e0d4e_3_14]] belong to one class and [[IMAGE:0bc72138559e0d4e_3_15]] belong to another class. It is not possible to separate these two classes by using a linear decision boundary. We consider a neural network with: • Input [[IMAGE:0bc72138559e0d4e_3_16]] • One hidden layer with 2 neurons and ReLU activation • Output layer: one neuron with linear activation [[IMAGE:0bc72138559e0d4e_4_17]] The network is defined as: [[IMAGE:0bc72138559e0d4e_4_18]] Consider the weight matrices: [[IMAGE:0bc72138559e0d4e_4_19]] Can this network correctly classify XOR using a threshold at [[IMAGE:0bc72138559e0d4e_4_20]] ?
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
  1. Yes
  2. No

A published solution is not available for this question yet.

Question 6 MCQ · 3.0 marks

Match the attention mechanism in Column I with the correct description in Column II. [[IMAGE:0bc72138559e0d4e_4_21]]
Source diagram or notation
  1. [[IMAGE:0bc72138559e0d4e_4_22]]
    Source diagram or notation
  2. [[IMAGE:0bc72138559e0d4e_5_23]]
    Source diagram or notation
  3. [[IMAGE:0bc72138559e0d4e_5_24]]
    Source diagram or notation
  4. [[IMAGE:0bc72138559e0d4e_5_25]]
    Source diagram or notation

A published solution is not available for this question yet.

Question 7 NAT · 2.0 marks

Consider a neural network with the following architecture: • Input: [[IMAGE:0bc72138559e0d4e_5_26]] , Output: [[IMAGE:0bc72138559e0d4e_5_27]] . • [[IMAGE:0bc72138559e0d4e_5_28]] hidden layers with sigmoid(logistic) as the activation function, where each hidden layer has [[IMAGE:0bc72138559e0d4e_5_29]] neurons. One neuron in the output layer with sigmoid as the activation function. • All the weights are initialized as zero. Ignore the biases. Based on the above data, answer the given subquestions.
What will be the output [[IMAGE:0bc72138559e0d4e_5_30]] ? Enter the answer correct to one decimal place.
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

    A published solution is not available for this question yet.

    Question 8 NAT · 2.0 marks

    Consider a neural network with the following architecture: • Input: [[IMAGE:0bc72138559e0d4e_5_26]] , Output: [[IMAGE:0bc72138559e0d4e_5_27]] . • [[IMAGE:0bc72138559e0d4e_5_28]] hidden layers with sigmoid(logistic) as the activation function, where each hidden layer has [[IMAGE:0bc72138559e0d4e_5_29]] neurons. One neuron in the output layer with sigmoid as the activation function. • All the weights are initialized as zero. Ignore the biases. Based on the above data, answer the given subquestions.
    What will be the output [[IMAGE:0bc72138559e0d4e_6_31]] if we change the activation function in the outer layer to linear?
    Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

      A published solution is not available for this question yet.

      Question 9 MSQ · 3.0 marks

      Consider a neural network with the following architecture: • Input: [[IMAGE:0bc72138559e0d4e_5_26]] , Output: [[IMAGE:0bc72138559e0d4e_5_27]] . • [[IMAGE:0bc72138559e0d4e_5_28]] hidden layers with sigmoid(logistic) as the activation function, where each hidden layer has [[IMAGE:0bc72138559e0d4e_5_29]] neurons. One neuron in the output layer with sigmoid as the activation function. • All the weights are initialized as zero. Ignore the biases. Based on the above data, answer the given subquestions.
      Suppose you perform the backpropagation using the original configuration of the network and update the parameters of the network using stochastic gradient descent. Let [[IMAGE:0bc72138559e0d4e_6_32]] be the weight associated with the first hidden layer to input. Then, which among the following will be true for [[IMAGE:0bc72138559e0d4e_6_33]] ?
      Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
      1. It remains the same after the first update using stochastic gradient descent.
      2. Different neurons in the first hidden layer will receive different gradient updates.
      3. All the elements in [[IMAGE:0bc72138559e0d4e_6_34]] will be zero after the first update.
        Source diagram or notation
      4. Entries in [[IMAGE:0bc72138559e0d4e_6_35]] is either all positive or all negative.
        Source diagram or notation

      A published solution is not available for this question yet.

      Question 10 NAT · 2.0 marks

      For a binary classification task, consider an image [[IMAGE:0bc72138559e0d4e_6_36]] of size [[IMAGE:0bc72138559e0d4e_6_37]] passed through a toy convolutional network that has the following architecture: • **ConvLayer:** filter size = [[IMAGE:0bc72138559e0d4e_7_38]] , and number of filter/channels = 1, with stride 1 and no padding. • Convolved output is passed through a ReLU activation function. • **Max-Pool Layer:** filter size = [[IMAGE:0bc72138559e0d4e_7_39]] , with stride 2 and padding not applied. • Output of Max-pool layer is taken as [[IMAGE:0bc72138559e0d4e_7_40]] . • Ignore the bias. Following are the image and the kernel matrix [[IMAGE:0bc72138559e0d4e_7_41]] : [[IMAGE:0bc72138559e0d4e_7_42]] Based on the above data, answer the given subquestions.
      Find the total number of parameters in the network.
      Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

        A published solution is not available for this question yet.

        Question 11 NAT · 3.0 marks

        For a binary classification task, consider an image [[IMAGE:0bc72138559e0d4e_6_36]] of size [[IMAGE:0bc72138559e0d4e_6_37]] passed through a toy convolutional network that has the following architecture: • **ConvLayer:** filter size = [[IMAGE:0bc72138559e0d4e_7_38]] , and number of filter/channels = 1, with stride 1 and no padding. • Convolved output is passed through a ReLU activation function. • **Max-Pool Layer:** filter size = [[IMAGE:0bc72138559e0d4e_7_39]] , with stride 2 and padding not applied. • Output of Max-pool layer is taken as [[IMAGE:0bc72138559e0d4e_7_40]] . • Ignore the bias. Following are the image and the kernel matrix [[IMAGE:0bc72138559e0d4e_7_41]] : [[IMAGE:0bc72138559e0d4e_7_42]] Based on the above data, answer the given subquestions.
        [[IMAGE:0bc72138559e0d4e_7_44]] [[IMAGE:0bc72138559e0d4e_7_45]] The loss is defined as [[IMAGE:0bc72138559e0d4e_7_43]] . If then compute .
        Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

          A published solution is not available for this question yet.

          Question 12 MSQ · 2.0 marks

          After training a neural network, the training error is observed to be 10%. It gives the test error to be 56%. Which of the following methods can be used to reduce the test error?
          1. Adam
          2. ReLu activation
          3. Injecting noise at input
          4. Maxout
          5. L2 regularization

          A published solution is not available for this question yet.

          Question 13 MSQ · 3.0 marks

          Consider a transformer model using scaled dot-product attention: [[IMAGE:0bc72138559e0d4e_8_46]] Assume a fixed input size and embedding dimension [[IMAGE:0bc72138559e0d4e_8_47]] . The model may use either single-head or multi-head attention. Which of the following statements are correct?
          Source diagram or notationSource diagram or notation
          1. Scaling by [[IMAGE:0bc72138559e0d4e_8_48]] helps prevent the dot-product values from becoming too large when [[IMAGE:0bc72138559e0d4e_8_49]] is high.
            Source diagram or notationSource diagram or notation
          2. In multi-head attention, all heads share the same learned projections of [[IMAGE:0bc72138559e0d4e_8_50]] .
            Source diagram or notation
          3. If all attention heads learn identical projection matrices, multi-head attention behaves equivalently to single-head attention.
          4. For a fixed embedding dimension, increasing the number of heads always increases the total number of parameters in the query, key, and value projections.

          A published solution is not available for this question yet.

          Question 14 NAT · 3.0 marks

          A text corpus contains the following three sentences: 1. students read the books in library 2. teachers read the books after class 3. students and teachers discuss in library After removing the stop words "the'', "in", and "and'', the vocabulary becomes: [[IMAGE:0bc72138559e0d4e_9_51]] Assume the context window size is [[IMAGE:0bc72138559e0d4e_9_52]] (one word to the left and one word to the right). Based on the above data, answer the given subquestions.
          Construct the co-occurrence matrix [[IMAGE:0bc72138559e0d4e_9_53]] , where rows represent target words and columns represent context words. Let [[IMAGE:0bc72138559e0d4e_9_54]] denote the element in the [[IMAGE:0bc72138559e0d4e_9_55]] -th row and [[IMAGE:0bc72138559e0d4e_9_56]] -th column. Enter the value of [[IMAGE:0bc72138559e0d4e_9_57]] .
          Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

            A published solution is not available for this question yet.

            Question 15 MCQ · 3.0 marks

            A text corpus contains the following three sentences: 1. students read the books in library 2. teachers read the books after class 3. students and teachers discuss in library After removing the stop words "the'', "in", and "and'', the vocabulary becomes: [[IMAGE:0bc72138559e0d4e_9_51]] Assume the context window size is [[IMAGE:0bc72138559e0d4e_9_52]] (one word to the left and one word to the right). Based on the above data, answer the given subquestions.
            Based on the co-occurrence matrix constructed previously, determine whether [[IMAGE:0bc72138559e0d4e_10_58]] is positive, zero, or negative. (Use natural log).
            Source diagram or notationSource diagram or notationSource diagram or notation
            1. Positive
            2. Zero
            3. Negative

            A published solution is not available for this question yet.

            Question 16 MSQ · 3.0 marks

            A text corpus contains the following three sentences: 1. students read the books in library 2. teachers read the books after class 3. students and teachers discuss in library After removing the stop words "the'', "in", and "and'', the vocabulary becomes: [[IMAGE:0bc72138559e0d4e_9_51]] Assume the context window size is [[IMAGE:0bc72138559e0d4e_9_52]] (one word to the left and one word to the right). Based on the above data, answer the given subquestions.
            Which of the following problems of the co-occurrence matrix are addressed by applying Singular Value Decomposition (SVD)? Choose all the correct options.
            Source diagram or notationSource diagram or notation
            1. High dimensional representation
            2. Sparse matrix with many zeros
            3. Captures latent semantic similarity between words
            4. Increases vocabulary size
            5. Produces dense low-dimensional embeddings

            A published solution is not available for this question yet.

            Question 17 NAT · 2.0 marks

            You are given a transformer attention setup with the following configuration: • Sequence Length : [[IMAGE:0bc72138559e0d4e_10_59]] • Number of Heads : [[IMAGE:0bc72138559e0d4e_10_60]] • Embedding dimension : [[IMAGE:0bc72138559e0d4e_10_61]] • Input [[IMAGE:0bc72138559e0d4e_10_62]] • [[IMAGE:0bc72138559e0d4e_11_63]] • [[IMAGE:0bc72138559e0d4e_11_64]] • [[IMAGE:0bc72138559e0d4e_11_65]] • [[IMAGE:0bc72138559e0d4e_11_66]] • [[IMAGE:0bc72138559e0d4e_11_67]] Based on the above data, answer the given subquestions.
            Suppose [[IMAGE:0bc72138559e0d4e_11_68]] , [[IMAGE:0bc72138559e0d4e_11_69]] , [[IMAGE:0bc72138559e0d4e_11_70]] and [[IMAGE:0bc72138559e0d4e_11_71]] . The scaled dot-product attention operation for a single head, given by: [[IMAGE:0bc72138559e0d4e_11_72]] Compute the total number of elements in the resulting attention output.
            Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

              A published solution is not available for this question yet.

              Question 18 NAT · 2.0 marks

              You are given a transformer attention setup with the following configuration: • Sequence Length : [[IMAGE:0bc72138559e0d4e_10_59]] • Number of Heads : [[IMAGE:0bc72138559e0d4e_10_60]] • Embedding dimension : [[IMAGE:0bc72138559e0d4e_10_61]] • Input [[IMAGE:0bc72138559e0d4e_10_62]] • [[IMAGE:0bc72138559e0d4e_11_63]] • [[IMAGE:0bc72138559e0d4e_11_64]] • [[IMAGE:0bc72138559e0d4e_11_65]] • [[IMAGE:0bc72138559e0d4e_11_66]] • [[IMAGE:0bc72138559e0d4e_11_67]] Based on the above data, answer the given subquestions.
              Suppose [[IMAGE:0bc72138559e0d4e_11_73]] , [[IMAGE:0bc72138559e0d4e_11_74]] , [[IMAGE:0bc72138559e0d4e_11_75]] and [[IMAGE:0bc72138559e0d4e_11_76]] . What will be the number of parameters in the Multihead Attention ? (ignore the bias)
              Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

                A published solution is not available for this question yet.

                Question 19 NAT · 2.0 marks

                You are given a transformer attention setup with the following configuration: • Sequence Length : [[IMAGE:0bc72138559e0d4e_10_59]] • Number of Heads : [[IMAGE:0bc72138559e0d4e_10_60]] • Embedding dimension : [[IMAGE:0bc72138559e0d4e_10_61]] • Input [[IMAGE:0bc72138559e0d4e_10_62]] • [[IMAGE:0bc72138559e0d4e_11_63]] • [[IMAGE:0bc72138559e0d4e_11_64]] • [[IMAGE:0bc72138559e0d4e_11_65]] • [[IMAGE:0bc72138559e0d4e_11_66]] • [[IMAGE:0bc72138559e0d4e_11_67]] Based on the above data, answer the given subquestions.
                Assume a Feed-Forward Network (FFN) follows the Multi-Head Attention layer in the encoder. The FFN consists of two linear transformations: [[IMAGE:0bc72138559e0d4e_12_77]] Calculate the total number of parameters in the FFN layer (including bias terms), where [[IMAGE:0bc72138559e0d4e_12_78]] and [[IMAGE:0bc72138559e0d4e_12_79]] .
                Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

                  A published solution is not available for this question yet.

                  Question 20 NAT · 2.0 marks

                  Assume that your CBOW model outputs a probability distribution over a vocabulary of 20,000 words for a given context. If the correct target word is word number 150, and the model's predicted probability for this word is 0.2, calculate the cross-entropy loss for this prediction. Enter the answer correct to two decimal places. (use natural log)

                    A published solution is not available for this question yet.

                    Question 21 NAT · 3.0 marks

                    In the Skip-gram model with a window size of 2 (on each side), how many **unique** pairs of target and context words will be generated for the following sentence. Note: Ignore punctuation in your calculation '**Learning from data helps build intelligent systems for society**'

                      A published solution is not available for this question yet.