MauryaHub PYQ Practice

da5004_2024T1_Q1_NA.pdf

Large Language Models · Quiz 1 · Jan 2024

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 181 NAT · 1.0 marks

Enter the correct answer for blank (b) __________.

    A published solution is not available for this question yet.

    Question 183 MCQ · 0.0 marks

    [[IMAGE:f9418c188bfff09a_3_0]]
    Source diagram or notation
    1. Useful Data has been mentioned above.
    2. This data attachment is just for a reference & not for an evaluation.

    A published solution is not available for this question yet.

    Question 184 NAT · 3.0 marks

    [[IMAGE:f9418c188bfff09a_4_1]] Based on the above data, answer the given subquestions.
    [[IMAGE:f9418c188bfff09a_4_2]]
    Source diagram or notationSource diagram or notation

      A published solution is not available for this question yet.

      Question 185 NAT · 2.0 marks

      [[IMAGE:f9418c188bfff09a_4_1]] Based on the above data, answer the given subquestions.
      [[IMAGE:f9418c188bfff09a_5_3]]
      Source diagram or notationSource diagram or notation

        A published solution is not available for this question yet.

        Question 186 NAT · 3.0 marks

        [[IMAGE:f9418c188bfff09a_4_1]] Based on the above data, answer the given subquestions.
        [[IMAGE:f9418c188bfff09a_5_4]] [[IMAGE:f9418c188bfff09a_5_5]]
        Source diagram or notationSource diagram or notationSource diagram or notation

          A published solution is not available for this question yet.

          Question 187 MCQ · 2.0 marks

          [[IMAGE:f9418c188bfff09a_4_1]] Based on the above data, answer the given subquestions.
          [[IMAGE:f9418c188bfff09a_6_6]] Which token in the sequence was given the highest score?
          Source diagram or notationSource diagram or notation
          1. G
          2. C
          3. T
          4. A

          A published solution is not available for this question yet.

          Question 188 NAT · 3.0 marks

          Consider the following configuration for the Vannila transformer architecture with one encoder layer and one decoder layer. [[IMAGE:f9418c188bfff09a_7_7]] Based on the above data, answer the given subquestions.
          Suppose the number of learnable parameters in the source input embedding layer is 1600, how many parameters are there in the positional embedding layer of the source language? Assume the positional embeddings are learnable.
          Source diagram or notation

            A published solution is not available for this question yet.

            Question 189 NAT · 3.0 marks

            Consider the following configuration for the Vannila transformer architecture with one encoder layer and one decoder layer. [[IMAGE:f9418c188bfff09a_7_7]] Based on the above data, answer the given subquestions.
            How many parameters are there in the multi-head attention layer of the encoder (exclude the parameters in the WO matrix used for linear transformation)?
            Source diagram or notation

              A published solution is not available for this question yet.

              Question 190 MCQ · 4.0 marks

              [[IMAGE:f9418c188bfff09a_8_8]]
              Source diagram or notation
              1. [[IMAGE:f9418c188bfff09a_9_9]]
                Source diagram or notation
              2. [[IMAGE:f9418c188bfff09a_9_10]]
                Source diagram or notation
              3. [[IMAGE:f9418c188bfff09a_9_11]]
                Source diagram or notation
              4. [[IMAGE:f9418c188bfff09a_9_12]]
                Source diagram or notation

              A published solution is not available for this question yet.

              Question 191 MCQ · 4.0 marks

              Suppose we run the pre-trained GPT model in an autoregressive fashion for generating a text sequence. Then, the statement that “ We do not need to add the positional embedding to the predicted tokens at every time step” is
              1. TRUE
              2. FALSE

              A published solution is not available for this question yet.

              Question 192 MSQ · 3.0 marks

              Suppose we have a dataset for machine translation tasks with thousands of samples. Suppose a team considers training the transformer model. The model could be trained in two approaches A  Autoregressive training B  Teacher forcing Choose the correct statements
              1. Approach A helps the model to converge faster than approach B
              2. Approach B helps the model converge faster than approach A
              3. One can start the training with approach B first and then switch to approach A after some training steps
              4. Once the training starts with approach B and then switching to approach A after some training steps can not be done

              A published solution is not available for this question yet.

              Question 193 MSQ · 4.0 marks

              [[IMAGE:f9418c188bfff09a_11_13]]
              Source diagram or notation
              1. 512
              2. 1024
              3. 1536
              4. 2048

              A published solution is not available for this question yet.

              Question 194 MSQ · 4.0 marks

              Choose the correct statements
              1. All the parameters of the GPT model are randomly initialized during pre- training
              2. The model minimizes the CLM objective during fine-tuning to improve the performance
              3. All parameters of the GPT model are randomly initialized during fine-tuning
              4. In general, fine-tuning requires a dataset with labels

              A published solution is not available for this question yet.

              Question 195 MSQ · 4.0 marks

              [[IMAGE:f9418c188bfff09a_12_14]]
              Source diagram or notation
              1. Top-k sampling
              2. Beam search with the beam size K = 2
              3. Top-p sampling
              4. Exhaustive search
              5. None of these

              A published solution is not available for this question yet.

              Question 196 MCQ · 4.0 marks

              Consider a GPT model used for Causal language modelling. We feed the input sentence “This is a cool idea” to the model by appending special staring [BOS] and ending [EOS] tokens ( that is, “[BOS] This is a cool idea [EOS]”). Assume the context length of the model is 7. The attention matrix computed in one of the attention layers is given below [[IMAGE:f9418c188bfff09a_13_15]] Based on the above data, answer the given subquestions.
              The attention score matrix given in Table 1 is appropriate for the causal language modelling task
              Source diagram or notation
              1. True
              2. False
              3. Insufficient information

              A published solution is not available for this question yet.

              Question 197 NAT · 4.0 marks

              Consider a GPT model used for Causal language modelling. We feed the input sentence “This is a cool idea” to the model by appending special staring [BOS] and ending [EOS] tokens ( that is, “[BOS] This is a cool idea [EOS]”). Assume the context length of the model is 7. The attention matrix computed in one of the attention layers is given below [[IMAGE:f9418c188bfff09a_13_15]] Based on the above data, answer the given subquestions.
              Assume the time step starts from t = 0 and ends at t = 6. Suppose the model is at time step t = 4, what is the attention value assigned for the word “is”?
              Source diagram or notation

                A published solution is not available for this question yet.

                Question 198 MCQ · 3.0 marks

                The statement that “the Next Sentence Prediction (NSP) task requires the BERT model to run autoregressively given the first sentence( or segment)” is
                1. TRUE
                2. FALSE       **BDBN** **Section Id :** 64065351462 **Section Number :** 13

                A published solution is not available for this question yet.