da5004_2025T1_Q1_NA.pdf
Large Language Models · Quiz 1 · Jan 2025
← Course papers · Start practice / exam
This page contains the reliably extracted subset, not the complete original paper.
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 103 MCQ · 0.0 marks
General Note:
• If you find the given information is insufficient in any of the Numerical Answer Type (NAT)
questions, enter -1 as the answer.
• In all your calculations, take at least two digits. For example, if the intermediate result is
1.860875, then take it as 1.86 to the next step.
Instructions has been mentioned above.
This Instructions is just for a reference & not for an evaluation.
A published solution is not available for this question yet.
Question 104 MCQ · 2.0 marks
[[IMAGE:4019a127e45cedea_2_0]]

[[IMAGE:4019a127e45cedea_2_1]]

[[IMAGE:4019a127e45cedea_2_2]]

[[IMAGE:4019a127e45cedea_2_3]]

[[IMAGE:4019a127e45cedea_2_4]]

A published solution is not available for this question yet.
Question 105 MCQ · 2.0 marks
Consider following assertion and reason pair:
**Assertion**: A transformer model cannot be trained in autogregressive mode for machine
translation task.
**Reason**: Teacher forcing improves the convergence rate.
Choose the correct statements
Assertion and Reason are both true and Reason is a correct explanation of
Assertion.
Assertion and Reason are both true and Reason is not a correct explanation of
Assertion.
Assertion is true but Reason is false.
Assertion is false but Reason is true.
A published solution is not available for this question yet.
Question 106 MCQ · 3.0 marks
Suppose we use a pre-trained model for text generation with the given prompt “I am going to”.
Which of the following decoding strategies can be used such that the pre-trained model generates
different text completion each time it is executed
Beam search with beam size 1
Greedy approach
Top-K with k = 2
None of these
A published solution is not available for this question yet.
Question 107 MSQ · 2.0 marks
Which of the following models, in general, struggle to encode a context when translating long
sentences?
An RNN model
An RNN model with attention mechanism
A transformer model
A published solution is not available for this question yet.
Question 108 MSQ · 2.0 marks
Choose the correct statements
All the parameters of the GPT model are randomly initialized during pre-
training
The model minimizes the CLM objective during fine-tuning to improve the
performance
All parameters of the GPT model are randomly initialized during fine-tuning
In general, fine-tuning requires a dataset with labels
A published solution is not available for this question yet.
Question 109 MSQ · 3.0 marks
[[IMAGE:4019a127e45cedea_4_5]]

[[IMAGE:4019a127e45cedea_4_6]]

[[IMAGE:4019a127e45cedea_4_7]]

[[IMAGE:4019a127e45cedea_4_8]]

[[IMAGE:4019a127e45cedea_4_9]]

A published solution is not available for this question yet.
Question 110 MSQ · 3.0 marks
Which of the following decoding strategies is (are) appropriate for machine translation tasks?
Assume we have infinite compute and the values for k ≥ 3, p > 0.35 where required.
Top-k sampling
Beam search with the beam size K = 2
Top-p sampling
Exhaustive search
None of these
A published solution is not available for this question yet.
Question 111 NAT · 2.0 marks
[[IMAGE:4019a127e45cedea_5_10]]
Based on the above data, answer the given subquestions.
Suppose the number of learnable parameters in the source input-embedding layer is 3200, how
many parameters are there in the positional embedding layer of the source language?

A published solution is not available for this question yet.
Question 112 NAT · 3.0 marks
[[IMAGE:4019a127e45cedea_5_10]]
Based on the above data, answer the given subquestions.
How many parameters are there in the multi-head attention layer of the encoder (exclude the
parameters in the WO matrix used for linear transformation and FFN layer)?

A published solution is not available for this question yet.
Question 113 NAT · 3.0 marks
[[IMAGE:4019a127e45cedea_5_10]]
Based on the above data, answer the given subquestions.
How many parameters does the matrix WQ have in anyone head of the encoder?

A published solution is not available for this question yet.
Question 114 NAT · 1.0 marks
[[IMAGE:4019a127e45cedea_6_11]]
Based on the above data, answer the given subquestions.
Say matrix F is computed by applying batch normalization on X, what will be sum of every element
in first / topmost row of F?

A published solution is not available for this question yet.
Question 115 NAT · 1.0 marks
[[IMAGE:4019a127e45cedea_6_11]]
Based on the above data, answer the given subquestions.
Say matrix F is computed by applying layer normalization on X, what will be sum of every element
in first / leftmost column of F?

A published solution is not available for this question yet.
Question 116 NAT · 2.0 marks
[[IMAGE:4019a127e45cedea_6_11]]
Based on the above data, answer the given subquestions.
[[IMAGE:4019a127e45cedea_7_12]]


A published solution is not available for this question yet.
Question 117 NAT · 2.0 marks
[[IMAGE:4019a127e45cedea_6_11]]
Based on the above data, answer the given subquestions.
[[IMAGE:4019a127e45cedea_8_13]]


A published solution is not available for this question yet.
Question 118 NAT · 2.0 marks
[[IMAGE:4019a127e45cedea_8_14]]
[[IMAGE:4019a127e45cedea_9_15]]


A published solution is not available for this question yet.
Question 119 NAT · 2.0 marks
[[IMAGE:4019a127e45cedea_8_14]]
[[IMAGE:4019a127e45cedea_9_16]]


A published solution is not available for this question yet.
Question 120 NAT · 3.0 marks
[[IMAGE:4019a127e45cedea_8_14]]
[[IMAGE:4019a127e45cedea_9_17]]


A published solution is not available for this question yet.
Question 121 MCQ · 2.0 marks
[[IMAGE:4019a127e45cedea_8_14]]
Once the model is sufficiently trained, which of the following tokens will have the highest
probability to be the next token, if the input is “Sachin has”:

broken
acted
highest
None of these
**i-NLP**
**Section Id :** 64065379955
**Section Number :** 8
**Section type :** Online
**Mandatory or Optional :** Mandatory
**Number of Questions :** 31
**Number of Questions to be attempted :** 31
**Section Marks :** 50
**Display Number Panel :** Yes
**Section Negative Marks :** 0
**Group All Questions :** No
**Enable Mark as Answered Mark for Review and**
No
**Clear Response :**
**Section Maximum Duration :** 0
**Section Minimum Duration :** 0
**Section Time In :** Minutes
**Maximum Instruction Time :** 0
A published solution is not available for this question yet.