MauryaHub PYQ Practice

da5004_2025T2_Q1_NA.pdf

Large Language Models · Quiz 1 · May 2025

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 133 MCQ · 2.0 marks

We continue the notation from the previous question. Suppose the randomly assigned weights lead to the good event, that is, there is an unique perfect matching in G with minimum weight, and say this minimum weight is r. What can you say about the determinant of Z in this case?
  1. [[IMAGE:a8f64888d57debe8_1_3]]
    Source diagram or notation
  2. [[IMAGE:a8f64888d57debe8_1_4]]
    Source diagram or notation
  3. [[IMAGE:a8f64888d57debe8_1_5]]
    Source diagram or notation
  4. [[IMAGE:a8f64888d57debe8_1_6]]       **LLM** **Section Id :** 64065391701 **Section Number :** 8 **Section type :** Online **Mandatory or Optional :** Mandatory **Number of Questions :** 16 **Number of Questions to be attempted :** 16 **Section Marks :** 50 **Display Number Panel :** Yes **Section Negative Marks :** 0 **Group All Questions :** No **Enable Mark as Answered Mark for Review and** No **Clear Response :** **Section Maximum Duration :** 0 **Section Minimum Duration :** 0 **Section Time In :** Minutes **Maximum Instruction Time :** 0
    Source diagram or notation

A published solution is not available for this question yet.

Question 135 MCQ · 2.0 marks

Transformers process input tokens:
  1. One at a time (sequentially)
  2. In reverse order
  3. All at once (in parallel)
  4. Only after seeing the full input

A published solution is not available for this question yet.

Question 136 MCQ · 2.0 marks

What is the purpose of the softmax function in the attention mechanism?
  1. Normalize attention scores to a probability distribution
  2. Add non-linearity to the model
  3. To predict the correct class for the loss function
  4. Remove redundant features from the input

A published solution is not available for this question yet.

Question 137 MCQ · 2.0 marks

What is the main difference between GPT and BERT pre-training objectives?
  1. GPT uses Masked Language Modeling, BERT uses Causal Language Modeling
  2. GPT uses Causal Language Modeling, BERT uses Masked Language Modeling
  3. Both use Masked Language Modeling
  4. Both use Causal Language Modeling

A published solution is not available for this question yet.

Question 138 MCQ · 2.0 marks

In Top-K sampling for language generation, increasing the value of K typically has which of the following effects?
  1. It makes the output more deterministic and repetitive.
  2. It reduces the probability of selecting high-frequency words.
  3. It increases the diversity of the generated text but may reduce coherence if K is too large.
  4. It guarantees grammatical correctness by focusing on top-ranked tokens only.

A published solution is not available for this question yet.

Question 139 MCQ · 2.0 marks

Why are residual connections important in transformer architectures?
  1. They reduce memory consumption.
  2. They add extra cost of compute by adding batch normalization.
  3. They help in training deep networks by enabling gradient flow.
  4. They remove the need for layer normalization.

A published solution is not available for this question yet.

Question 140 MSQ · 3.0 marks

Which of the following statements are true regarding causal language modeling (CLM)?
  1. The model only attends to past and current tokens during training.
  2. The model is trained by predicting the next token in a sequence.
  3. The model uses bidirectional context.
  4. The CLM objective is commonly used for encoder-only models.

A published solution is not available for this question yet.

Question 141 MSQ · 3.0 marks

Which of the following are valid reasons why transformer-based large language models are widely used in natural language processing?
  1. Transformers process input sequences in parallel, enabling faster training.
  2. They use recurrence to remember long-term dependencies more effectively than LSTMs.
  3. Transformers are limited to short text inputs due to their architecture.
  4. Large transformer models can be fine-tuned for various NLP tasks using a single pre-trained model.

A published solution is not available for this question yet.

Question 142 MSQ · 3.0 marks

Which elements are included in BERT’s input representation for Next Sentence Prediction?
  1. [[IMAGE:a8f64888d57debe8_5_7]]
    Source diagram or notation
  2. [[IMAGE:a8f64888d57debe8_5_8]]
    Source diagram or notation
  3. [[IMAGE:a8f64888d57debe8_5_9]]
    Source diagram or notation
  4. [[IMAGE:a8f64888d57debe8_5_10]]
    Source diagram or notation

A published solution is not available for this question yet.

Question 143 MSQ · 3.0 marks

How does Top-p (nucleus) sampling differ from Top-K sampling in language generation?
  1. It samples only from a fixed number of tokens at each step.
  2. It samples from the smallest set of tokens whose cumulative probability exceeds p.
  3. It always selects the top-p tokens with equal probability.
  4. It guarantees diversity by selecting all low-probability tokens.

A published solution is not available for this question yet.

Question 144 NAT · 3.0 marks

[[IMAGE:a8f64888d57debe8_6_11]]
Source diagram or notation

    A published solution is not available for this question yet.

    Question 145 NAT · 3.0 marks

    [[IMAGE:a8f64888d57debe8_7_12]]
    Source diagram or notation

      A published solution is not available for this question yet.

      Question 146 NAT · 3.0 marks

      [[IMAGE:a8f64888d57debe8_7_13]]
      Source diagram or notation

        A published solution is not available for this question yet.

        Question 147 MCQ · 3.0 marks

        [[IMAGE:a8f64888d57debe8_9_14]] Based on the above data, answer the given subquestions.
        Select the scaled dot-product attention for the first head:
        Source diagram or notation
        1. [[IMAGE:a8f64888d57debe8_10_15]]
          Source diagram or notation
        2. [[IMAGE:a8f64888d57debe8_10_16]]
          Source diagram or notation
        3. [[IMAGE:a8f64888d57debe8_10_17]]
          Source diagram or notation
        4. [[IMAGE:a8f64888d57debe8_10_18]]
          Source diagram or notation

        A published solution is not available for this question yet.

        Question 148 MCQ · 3.0 marks

        [[IMAGE:a8f64888d57debe8_9_14]] Based on the above data, answer the given subquestions.
        Select the scaled dot-product attention for the second head:
        Source diagram or notation
        1. [[IMAGE:a8f64888d57debe8_10_19]]
          Source diagram or notation
        2. [[IMAGE:a8f64888d57debe8_10_20]]
          Source diagram or notation
        3. [[IMAGE:a8f64888d57debe8_10_21]]
          Source diagram or notation
        4. [[IMAGE:a8f64888d57debe8_10_22]]
          Source diagram or notation

        A published solution is not available for this question yet.

        Question 149 MCQ · 2.0 marks

        [[IMAGE:a8f64888d57debe8_9_14]] Based on the above data, answer the given subquestions.
        Concatenate the outputs from both the attention heads, then apply the output projection matrix Wo to produce the final output of the multi-head attention mechanism. Select the correct result of this operation.
        Source diagram or notation
        1. [[IMAGE:a8f64888d57debe8_11_23]]
          Source diagram or notation
        2. [[IMAGE:a8f64888d57debe8_11_24]]
          Source diagram or notation
        3. [[IMAGE:a8f64888d57debe8_11_25]]
          Source diagram or notation
        4. [[IMAGE:a8f64888d57debe8_11_26]]
          Source diagram or notation

        A published solution is not available for this question yet.

        Question 150 NAT · 1.0 marks

        [[IMAGE:a8f64888d57debe8_12_27]] Based on the above data, answer the given subquestions.
        For the given input matrix X and multihead attention output MHA(X), apply a residual connection and store the result in matrix R, and finally compute the sum of all elements in matrix R.
        Source diagram or notation

          A published solution is not available for this question yet.

          Question 151 NAT · 2.0 marks

          [[IMAGE:a8f64888d57debe8_12_27]] Based on the above data, answer the given subquestions.
          [[IMAGE:a8f64888d57debe8_13_28]]
          Source diagram or notationSource diagram or notation

            A published solution is not available for this question yet.

            Question 152 NAT · 3.0 marks

            [[IMAGE:a8f64888d57debe8_12_27]] Based on the above data, answer the given subquestions.
            [[IMAGE:a8f64888d57debe8_13_29]]
            Source diagram or notationSource diagram or notation

              A published solution is not available for this question yet.

              Question 153 NAT · 2.0 marks

              The table presents the **conditional probability distribution** over vocabulary tokens at each timestep during sequence generation. Each **column** corresponds to a timestep (from 1 to 5), and the values represent the probability of selecting each token given the tokens chosen in all previous timesteps. For timestep t, the values in the column represent: [[IMAGE:a8f64888d57debe8_14_30]] Based on the above data, answer the given subquestions.
              In exhaustive search, at timestep t=1, we run the decoder once to obtain probability distributions over all tokens in the vocabulary. Given the 6 tokens available in the table in the main question, how many times must we run the decoder at timestep t=4?
              Source diagram or notation

                A published solution is not available for this question yet.

                Question 154 NAT · 1.0 marks

                The table presents the **conditional probability distribution** over vocabulary tokens at each timestep during sequence generation. Each **column** corresponds to a timestep (from 1 to 5), and the values represent the probability of selecting each token given the tokens chosen in all previous timesteps. For timestep t, the values in the column represent: [[IMAGE:a8f64888d57debe8_14_30]] Based on the above data, answer the given subquestions.
                How many total sequences of exactly length 5 are possible according to exhaustive search?
                Source diagram or notation

                  A published solution is not available for this question yet.

                  Question 155 NAT · 2.0 marks

                  The table presents the **conditional probability distribution** over vocabulary tokens at each timestep during sequence generation. Each **column** corresponds to a timestep (from 1 to 5), and the values represent the probability of selecting each token given the tokens chosen in all previous timesteps. For timestep t, the values in the column represent: [[IMAGE:a8f64888d57debe8_14_30]] Based on the above data, answer the given subquestions.
                  If we use top-k sampling with k=2 at timestep 1, what is the normalized probability of selecting token “sky” at the timestep=1?
                  Source diagram or notation

                    A published solution is not available for this question yet.