MauryaHub PYQ Practice

da5004_2025T3_Q1_NA.pdf

Large Language Models · Quiz 1 · Sep 2025

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 173 MCQ · 2.0 marks

[[IMAGE:30260f7a668f5932_2_1]]
Source diagram or notation
  1. To increase the numerical precision of the attention scores.
  2. To reduce the overall computational cost of the attention mechanism.
  3. To ensure that the sum of the attention scores for each query equals 1.
  4. To avoid numerical issues and loss of gradients during training.

A published solution is not available for this question yet.

Question 174 MCQ · 2.0 marks

Which of the following is true regarding sinusoidal encoding?
  1. The encoding vector of words present at even positions in the given text sequence uses sine, and the encoding vector of words present at odd positions uses cosine.
  2. The encoding vector of words present at odd positions in the given text sequence uses sine, and the encoding vector of words present at even positions uses cosine.
  3. The encoding vector for each word position contains values computed using both sine (for even dimensions) and cosine (for odd dimensions).

A published solution is not available for this question yet.

Question 175 MCQ · 2.0 marks

[[IMAGE:30260f7a668f5932_3_2]]
Source diagram or notation
  1. ‘content’
  2. ‘je’
  3. ‘suis’
  4. ‘heureux’

A published solution is not available for this question yet.

Question 176 MCQ · 2.0 marks

[[IMAGE:30260f7a668f5932_3_3]]
Source diagram or notation
  1. [[IMAGE:30260f7a668f5932_3_4]]
    Source diagram or notation
  2. [[IMAGE:30260f7a668f5932_3_5]]
    Source diagram or notation
  3. [[IMAGE:30260f7a668f5932_3_6]]
    Source diagram or notation
  4. [[IMAGE:30260f7a668f5932_3_7]]
    Source diagram or notation
  5. [[IMAGE:30260f7a668f5932_3_8]]
    Source diagram or notation

A published solution is not available for this question yet.

Question 177 MCQ · 2.0 marks

In the context of language model training, teacher forcing is a technique where:
  1. The correct target tokens from the training data are used as input for the next time step, regardless of the model’s prediction.
  2. The model’s predictions are used as the input for the next time step, even if they are incorrect.
  3. The model is forced to learn without any pre-trained weights.
  4. The model is forced to correct the input data.

A published solution is not available for this question yet.

Question 178 MCQ · 2.0 marks

[[IMAGE:30260f7a668f5932_4_9]]
Source diagram or notation
  1. [[IMAGE:30260f7a668f5932_4_10]]
    Source diagram or notation
  2. [[IMAGE:30260f7a668f5932_4_11]]
    Source diagram or notation
  3. [[IMAGE:30260f7a668f5932_4_12]]
    Source diagram or notation
  4. [[IMAGE:30260f7a668f5932_4_13]]
    Source diagram or notation

A published solution is not available for this question yet.

Question 179 MCQ · 3.0 marks

[[IMAGE:30260f7a668f5932_5_14]]
Source diagram or notation
  1. [[IMAGE:30260f7a668f5932_5_15]]
    Source diagram or notation
  2. [[IMAGE:30260f7a668f5932_5_16]]
    Source diagram or notation
  3. [[IMAGE:30260f7a668f5932_5_17]]
    Source diagram or notation
  4. [[IMAGE:30260f7a668f5932_5_18]]
    Source diagram or notation

A published solution is not available for this question yet.

Question 180 MCQ · 3.0 marks

[[IMAGE:30260f7a668f5932_5_19]]
Source diagram or notation
  1. [[IMAGE:30260f7a668f5932_5_20]]
    Source diagram or notation
  2. [[IMAGE:30260f7a668f5932_5_21]]
    Source diagram or notation
  3. [[IMAGE:30260f7a668f5932_5_22]]
    Source diagram or notation
  4. [[IMAGE:30260f7a668f5932_5_23]]
    Source diagram or notation
  5. [[IMAGE:30260f7a668f5932_5_24]]
    Source diagram or notation

A published solution is not available for this question yet.

Question 181 MCQ · 3.0 marks

For a vocabulary size of 10, how many beams will there be in greedy search, beam search with beam size 3, and exhaustive search at step 1 and step 3, respectively?
  1. Step 1: Greedy: 1, Beam: 1, Exhaustive: 1 Step 3: Greedy: 1, Beam: 1, Exhaustive: 1
  2. Step 1: Greedy: 1, Beam: 3, Exhaustive: 10 Step 3: Greedy: 1, Beam: 3, Exhaustive: 10
  3. Step 1: Greedy: 1, Beam: 3, Exhaustive: 1 Step 3: Greedy: 1, Beam: 3, Exhaustive: 100
  4. Step 1: Greedy: 1, Beam: 3, Exhaustive: 10 Step 3: Greedy: 1, Beam: 3, Exhaustive: 1000

A published solution is not available for this question yet.

Question 182 MSQ · 2.0 marks

Which of the following decoding strategies is/are inherently non-deterministic?
  1. Beam
  2. Greedy
  3. Nucleus
  4. Exhaustive

A published solution is not available for this question yet.

Question 183 MSQ · 2.0 marks

Which of the following statements about multiple attention heads in a Transformer are true?
  1. They allow learning diverse local/global relations.
  2. They reduce memory usage compared to one head.
  3. Increasing the number of heads while keeping dmodel fixed, decreases the dimensionality handled by each head.
  4. All heads necessarily learn completely independent information about the input.

A published solution is not available for this question yet.

Question 184 MSQ · 4.0 marks

Which of the following statements about the Transformer architecture are true?
  1. Masking (look-ahead mask) is applied in every decoder self-attention layer, not just the first one.
  2. Token embeddings are used in the encoder but not in the decoder.
  3. Positional encoding is required only in the encoder, not in the decoder.
  4. Token embeddings are used in both the encoder and the decoder.
  5. Positional encodings are added to the input embeddings in both the encoder and the decoder.

A published solution is not available for this question yet.

Question 185 MSQ · 4.0 marks

[[IMAGE:30260f7a668f5932_7_25]]
Source diagram or notation
  1. Beam search with beam size 2 will keep either A or B as per probability whereas Top-2 sampling will keep both A and B.
  2. Beam search with beam size 2 will keep both A and B whereas Top-2 sampling will keep either A or B.
  3. Beam search with beam size 2 will randomly choose any two tokens, whereas Top-2 sampling will always pick the two least probable tokens.
  4. Beam search with beam size 2 will always keep the top-2 most probable tokens, whereas Top-2 sampling will restrict choices to top-2 tokens and then sample one probabilistically.

A published solution is not available for this question yet.

Question 186 NAT · 3.0 marks

[[IMAGE:30260f7a668f5932_8_26]]
Source diagram or notation

    A published solution is not available for this question yet.

    Question 187 NAT · 3.0 marks

    [[IMAGE:30260f7a668f5932_8_27]]
    Source diagram or notation

      A published solution is not available for this question yet.

      Question 188 NAT · 2.0 marks

      [[IMAGE:30260f7a668f5932_9_28]]
      Source diagram or notation

        A published solution is not available for this question yet.

        Question 189 NAT · 3.0 marks

        [[IMAGE:30260f7a668f5932_9_29]]
        How many next-token prediction targets are generated during training a full batch?
        Source diagram or notation

          A published solution is not available for this question yet.

          Question 190 NAT · 2.0 marks

          [[IMAGE:30260f7a668f5932_9_29]]
          For one sequence, how many non-zero attention scores remain in the masked attention score matrix (per head)?
          Source diagram or notation

            A published solution is not available for this question yet.

            Question 191 NAT · 2.0 marks

            [[IMAGE:30260f7a668f5932_9_29]]
            [[IMAGE:30260f7a668f5932_10_30]]
            Source diagram or notationSource diagram or notation

              A published solution is not available for this question yet.