MauryaHub PYQ Practice

da5004_2025T1_Q1_NA.pdf

Large Language Models · Quiz 1 · Jan 2025

← Course papers · Start practice / exam

This page contains the reliably extracted subset, not the complete original paper.

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 103 MCQ · 0.0 marks

General Note: • If you find the given information is insufficient in any of the Numerical Answer Type (NAT) questions, enter -1 as the answer. • In all your calculations, take at least two digits. For example, if the intermediate result is 1.860875, then take it as 1.86 to the next step.
  1. Instructions has been mentioned above.
  2. This Instructions is just for a reference & not for an evaluation.

A published solution is not available for this question yet.

Question 104 MCQ · 2.0 marks

[[IMAGE:4019a127e45cedea_2_0]]
Source diagram or notation
  1. [[IMAGE:4019a127e45cedea_2_1]]
    Source diagram or notation
  2. [[IMAGE:4019a127e45cedea_2_2]]
    Source diagram or notation
  3. [[IMAGE:4019a127e45cedea_2_3]]
    Source diagram or notation
  4. [[IMAGE:4019a127e45cedea_2_4]]
    Source diagram or notation

A published solution is not available for this question yet.

Question 105 MCQ · 2.0 marks

Consider following assertion and reason pair: **Assertion**: A transformer model cannot be trained in autogregressive mode for machine translation task. **Reason**: Teacher forcing improves the convergence rate. Choose the correct statements
  1. Assertion and Reason are both true and Reason is a correct explanation of Assertion.
  2. Assertion and Reason are both true and Reason is not a correct explanation of Assertion.
  3. Assertion is true but Reason is false.
  4. Assertion is false but Reason is true.

A published solution is not available for this question yet.

Question 106 MCQ · 3.0 marks

Suppose we use a pre-trained model for text generation with the given prompt “I am going to”. Which of the following decoding strategies can be used such that the pre-trained model generates different text completion each time it is executed
  1. Beam search with beam size 1
  2. Greedy approach
  3. Top-K with k = 2
  4. None of these

A published solution is not available for this question yet.

Question 107 MSQ · 2.0 marks

Which of the following models, in general, struggle to encode a context when translating long sentences?
  1. An RNN model
  2. An RNN model with attention mechanism
  3. A transformer model

A published solution is not available for this question yet.

Question 108 MSQ · 2.0 marks

Choose the correct statements
  1. All the parameters of the GPT model are randomly initialized during pre- training
  2. The model minimizes the CLM objective during fine-tuning to improve the performance
  3. All parameters of the GPT model are randomly initialized during fine-tuning
  4. In general, fine-tuning requires a dataset with labels

A published solution is not available for this question yet.

Question 109 MSQ · 3.0 marks

[[IMAGE:4019a127e45cedea_4_5]]
Source diagram or notation
  1. [[IMAGE:4019a127e45cedea_4_6]]
    Source diagram or notation
  2. [[IMAGE:4019a127e45cedea_4_7]]
    Source diagram or notation
  3. [[IMAGE:4019a127e45cedea_4_8]]
    Source diagram or notation
  4. [[IMAGE:4019a127e45cedea_4_9]]
    Source diagram or notation

A published solution is not available for this question yet.

Question 110 MSQ · 3.0 marks

Which of the following decoding strategies is (are) appropriate for machine translation tasks? Assume we have infinite compute and the values for k ≥ 3, p > 0.35 where required.
  1. Top-k sampling
  2. Beam search with the beam size K = 2
  3. Top-p sampling
  4. Exhaustive search
  5. None of these

A published solution is not available for this question yet.

Question 111 NAT · 2.0 marks

[[IMAGE:4019a127e45cedea_5_10]] Based on the above data, answer the given subquestions.
Suppose the number of learnable parameters in the source input-embedding layer is 3200, how many parameters are there in the positional embedding layer of the source language?
Source diagram or notation

    A published solution is not available for this question yet.

    Question 112 NAT · 3.0 marks

    [[IMAGE:4019a127e45cedea_5_10]] Based on the above data, answer the given subquestions.
    How many parameters are there in the multi-head attention layer of the encoder (exclude the parameters in the WO matrix used for linear transformation and FFN layer)?
    Source diagram or notation

      A published solution is not available for this question yet.

      Question 113 NAT · 3.0 marks

      [[IMAGE:4019a127e45cedea_5_10]] Based on the above data, answer the given subquestions.
      How many parameters does the matrix WQ have in anyone head of the encoder?
      Source diagram or notation

        A published solution is not available for this question yet.

        Question 114 NAT · 1.0 marks

        [[IMAGE:4019a127e45cedea_6_11]] Based on the above data, answer the given subquestions.
        Say matrix F is computed by applying batch normalization on X, what will be sum of every element in first / topmost row of F?
        Source diagram or notation

          A published solution is not available for this question yet.

          Question 115 NAT · 1.0 marks

          [[IMAGE:4019a127e45cedea_6_11]] Based on the above data, answer the given subquestions.
          Say matrix F is computed by applying layer normalization on X, what will be sum of every element in first / leftmost column of F?
          Source diagram or notation

            A published solution is not available for this question yet.

            Question 116 NAT · 2.0 marks

            [[IMAGE:4019a127e45cedea_6_11]] Based on the above data, answer the given subquestions.
            [[IMAGE:4019a127e45cedea_7_12]]
            Source diagram or notationSource diagram or notation

              A published solution is not available for this question yet.

              Question 117 NAT · 2.0 marks

              [[IMAGE:4019a127e45cedea_6_11]] Based on the above data, answer the given subquestions.
              [[IMAGE:4019a127e45cedea_8_13]]
              Source diagram or notationSource diagram or notation

                A published solution is not available for this question yet.

                Question 118 NAT · 2.0 marks

                [[IMAGE:4019a127e45cedea_8_14]]
                [[IMAGE:4019a127e45cedea_9_15]]
                Source diagram or notationSource diagram or notation

                  A published solution is not available for this question yet.

                  Question 119 NAT · 2.0 marks

                  [[IMAGE:4019a127e45cedea_8_14]]
                  [[IMAGE:4019a127e45cedea_9_16]]
                  Source diagram or notationSource diagram or notation

                    A published solution is not available for this question yet.

                    Question 120 NAT · 3.0 marks

                    [[IMAGE:4019a127e45cedea_8_14]]
                    [[IMAGE:4019a127e45cedea_9_17]]
                    Source diagram or notationSource diagram or notation

                      A published solution is not available for this question yet.

                      Question 121 MCQ · 2.0 marks

                      [[IMAGE:4019a127e45cedea_8_14]]
                      Once the model is sufficiently trained, which of the following tokens will have the highest probability to be the next token, if the input is “Sachin has”:
                      Source diagram or notation
                      1. broken
                      2. acted
                      3. highest
                      4. None of these       **i-NLP** **Section Id :** 64065379955 **Section Number :** 8 **Section type :** Online **Mandatory or Optional :** Mandatory **Number of Questions :** 31 **Number of Questions to be attempted :** 31 **Section Marks :** 50 **Display Number Panel :** Yes **Section Negative Marks :** 0 **Group All Questions :** No **Enable Mark as Answered Mark for Review and** No **Clear Response :** **Section Maximum Duration :** 0 **Section Minimum Duration :** 0 **Section Time In :** Minutes **Maximum Instruction Time :** 0

                      A published solution is not available for this question yet.