MauryaHub PYQ Practice

da5004_2026T2_Q1_NA.pdf

Large Language Models · Quiz 1 · May 2026

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 2 MSQ · 4.0 marks

Which statements are true about traditional attention used in sequence-to-sequence models?
  1. The decoder decides which encoder states are important.
  2. Attention weights are computed over all decoder hidden states.
  3. The context vector is a weighted sum of encoder states.
  4. Encoder hidden states are ignored after encoding.

A published solution is not available for this question yet.

Question 3 NAT · 3.0 marks

Consider the text "learn easy math", consisting of three tokens: {learn,easy,math}. The token embeddings are arranged in the input matrix [[IMAGE:7ca18a647a21b6db_2_2]] : [[IMAGE:7ca18a647a21b6db_2_3]] The query, key, and value projection matrices are given by: [[IMAGE:7ca18a647a21b6db_2_4]] [[IMAGE:7ca18a647a21b6db_3_5]] [[IMAGE:7ca18a647a21b6db_3_6]] [[IMAGE:7ca18a647a21b6db_3_7]] Based on the above data, answer the given subquestions.
Compute the [[IMAGE:7ca18a647a21b6db_3_8]] of the final scaled dot-product attention and submit the sum of the diagonal elements of [[IMAGE:7ca18a647a21b6db_3_9]] matrix.
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

    A published solution is not available for this question yet.

    Question 4 MCQ · 3.0 marks

    Consider the text "learn easy math", consisting of three tokens: {learn,easy,math}. The token embeddings are arranged in the input matrix [[IMAGE:7ca18a647a21b6db_2_2]] : [[IMAGE:7ca18a647a21b6db_2_3]] The query, key, and value projection matrices are given by: [[IMAGE:7ca18a647a21b6db_2_4]] [[IMAGE:7ca18a647a21b6db_3_5]] [[IMAGE:7ca18a647a21b6db_3_6]] [[IMAGE:7ca18a647a21b6db_3_7]] Based on the above data, answer the given subquestions.
    Choose the token pair with the least attention score.
    Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
    1. {"learn", "easy"}
    2. {"learn", "math"}
    3. {"easy", "math"}
    4. {"learn", "learn"}
    5. {"easy" , "easy"}

    A published solution is not available for this question yet.

    Question 5 NAT · 4.0 marks

    A Transformer model processes a sequence containing 6 tokens using Multi-Head Attention. The model initially uses 4 attention heads. If the number of attention heads is doubled to 8, while the sequence length remains unchanged, how many **additional** attention scores are computed across all heads?

      A published solution is not available for this question yet.

      Question 6 NAT · 4.0 marks

      A batch of 8 sentences is passed through a Transformer Encoder consisting of 6 encoder layers and 8 attention heads. After padding, each sentence has a maximum length of 32 tokens. The model uses an embedding dimension of [[IMAGE:7ca18a647a21b6db_4_10]] . What is the total number of elements (volume) in the output tensor produced by the final encoder layer for the entire batch?
      Source diagram or notation

        A published solution is not available for this question yet.

        Question 7 NAT · 4.0 marks

        Consider a masked multi-head attention layer in a transformer decoder. The source vocabulary is of size [[IMAGE:7ca18a647a21b6db_4_11]] , [[IMAGE:7ca18a647a21b6db_4_12]] and context length [[IMAGE:7ca18a647a21b6db_4_13]] . To enforce autoregressive generation, a causal mask is applied to the raw attention score before applying softmax. How many entries in the causal mask are set to [[IMAGE:7ca18a647a21b6db_4_14]] ?
        Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

          A published solution is not available for this question yet.

          Question 8 NAT · 4.0 marks

          Consider a Transformer model with a hidden state dimension [[IMAGE:7ca18a647a21b6db_5_15]] . The positional encoding for a given position, [[IMAGE:7ca18a647a21b6db_5_16]] and dimension [[IMAGE:7ca18a647a21b6db_5_17]] is defined by: [[IMAGE:7ca18a647a21b6db_5_18]] [[IMAGE:7ca18a647a21b6db_5_19]] Compute the squared euclidean norm for the vector corresponding to [[IMAGE:7ca18a647a21b6db_5_20]] and enter the value.
          Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

            A published solution is not available for this question yet.

            Question 9 MCQ · 4.0 marks

            Consider the following statements regarding the use of Teacher Forcing while training autoregressive sequence-to-sequence models and select the appropriate option: Statement 1: Teacher Forcing involves passing the input token to a larger Teacher model whose output is then input as the next token Statement 2: Teacher Forcing helps to prevent compounding of early incorrect predictions that can destabilize training
            1. Statement 1 is True but Statement 2 is False
            2. Statement 1 is False but Statement 2 is True
            3. Both the statements are True
            4. Both the statements are False

            A published solution is not available for this question yet.

            Question 10 NAT · 3.0 marks

            [[IMAGE:7ca18a647a21b6db_6_21]] Based on the above data, answer the given subquestions.
            Compute the joint probability of the sentence : [[IMAGE:7ca18a647a21b6db_6_22]] (Provide answer correct upto 2 digits after the decimal)
            Source diagram or notationSource diagram or notation

              A published solution is not available for this question yet.

              Question 11 MCQ · 3.0 marks

              [[IMAGE:7ca18a647a21b6db_6_21]] Based on the above data, answer the given subquestions.
              Choose the option corresponding to the sentence with the highest joint probability under the given language model.
              Source diagram or notation
              1. "i eat apples"
              2. "apples like You"
              3. "i eat bananas"
              4. "you like bananas
              5. "you like apples

              A published solution is not available for this question yet.

              Question 12 NAT · 2.0 marks

              [[IMAGE:7ca18a647a21b6db_6_21]] Based on the above data, answer the given subquestions.
              Using the given probability tree calculate the conditional probability of predicting "apples" given "like" [i.e P(apples | like)]. (Submit -1 if the provided information is insufficient)
              Source diagram or notation

                A published solution is not available for this question yet.

                Question 13 MSQ · 3.0 marks

                Which of the following are components found within a GPT decoder layer?
                1. Causal (Masked) Multi-Head Self-Attention
                2. Position-wise Feed-Forward Network
                3. Multi Head (Cross) Attention
                4. Recurrent Neural Network (RNN) Layer
                5. Add & Norm Layer

                A published solution is not available for this question yet.

                Question 14 NAT · 2.0 marks

                The table below represents the conditional probability distribution over the vocabulary. Each column corresponds to a decoding timestep and the values represent the conditional probability of selecting a token at the timestep given the previously generated tokens [[IMAGE:7ca18a647a21b6db_8_23]] For all calculations, consider only the tokens mentioned in the first column as the Vocabulary Based on the above data, answer the given subquestions.
                Use greedy decoding to determine the most likely sequence of 6 tokens and compute the probability of generating that sequence. Enter the value rounded off to 3 decimal places
                Source diagram or notation

                  A published solution is not available for this question yet.

                  Question 15 NAT · 2.0 marks

                  The table below represents the conditional probability distribution over the vocabulary. Each column corresponds to a decoding timestep and the values represent the conditional probability of selecting a token at the timestep given the previously generated tokens [[IMAGE:7ca18a647a21b6db_8_23]] For all calculations, consider only the tokens mentioned in the first column as the Vocabulary Based on the above data, answer the given subquestions.
                  Use exhaustive search strategy to identify the most probable sequence consisting of 4 tokens. How many decoder runs are required in total to identify the most probable sequence containing 4 tokens?
                  Source diagram or notation

                    A published solution is not available for this question yet.

                    Question 16 NAT · 3.0 marks

                    The table below represents the conditional probability distribution over the vocabulary. Each column corresponds to a decoding timestep and the values represent the conditional probability of selecting a token at the timestep given the previously generated tokens [[IMAGE:7ca18a647a21b6db_8_23]] For all calculations, consider only the tokens mentioned in the first column as the Vocabulary Based on the above data, answer the given subquestions.
                    Use Top-k sampling at timestep t = 4 and k = 3. Compute the renormalized probability of the most probable token from the candidate tokens and enter the value rounded off to three decimal places
                    Source diagram or notation

                      A published solution is not available for this question yet.

                      Question 17 NAT · 2.0 marks

                      The table below represents the conditional probability distribution over the vocabulary. Each column corresponds to a decoding timestep and the values represent the conditional probability of selecting a token at the timestep given the previously generated tokens [[IMAGE:7ca18a647a21b6db_8_23]] For all calculations, consider only the tokens mentioned in the first column as the Vocabulary Based on the above data, answer the given subquestions.
                      Consider the sequence "the castle was abandoned". What is the minimum beam width required to guarantee that this sequence is retained during beam search at timestep t = 4 ?
                      Source diagram or notation

                        A published solution is not available for this question yet.