da5004_2024T1_Q1_NA.pdf
Large Language Models · Quiz 1 · Jan 2024
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 181 NAT · 1.0 marks
Enter the correct answer for blank (b) __________.
A published solution is not available for this question yet.
Question 183 MCQ · 0.0 marks
[[IMAGE:f9418c188bfff09a_3_0]]

Useful Data has been mentioned above.
This data attachment is just for a reference & not for an evaluation.
A published solution is not available for this question yet.
Question 184 NAT · 3.0 marks
[[IMAGE:f9418c188bfff09a_4_1]]
Based on the above data, answer the given subquestions.
[[IMAGE:f9418c188bfff09a_4_2]]


A published solution is not available for this question yet.
Question 185 NAT · 2.0 marks
[[IMAGE:f9418c188bfff09a_4_1]]
Based on the above data, answer the given subquestions.
[[IMAGE:f9418c188bfff09a_5_3]]


A published solution is not available for this question yet.
Question 186 NAT · 3.0 marks
[[IMAGE:f9418c188bfff09a_4_1]]
Based on the above data, answer the given subquestions.
[[IMAGE:f9418c188bfff09a_5_4]]
[[IMAGE:f9418c188bfff09a_5_5]]



A published solution is not available for this question yet.
Question 187 MCQ · 2.0 marks
[[IMAGE:f9418c188bfff09a_4_1]]
Based on the above data, answer the given subquestions.
[[IMAGE:f9418c188bfff09a_6_6]]
Which token in the sequence was given the highest score?


G
C
T
A
A published solution is not available for this question yet.
Question 188 NAT · 3.0 marks
Consider the following configuration for the Vannila transformer architecture with one encoder
layer and one decoder layer.
[[IMAGE:f9418c188bfff09a_7_7]]
Based on the above data, answer the given subquestions.
Suppose the number of learnable parameters in the source input embedding layer is 1600, how
many parameters are there in the positional embedding layer of the source language? Assume the
positional embeddings are learnable.

A published solution is not available for this question yet.
Question 189 NAT · 3.0 marks
Consider the following configuration for the Vannila transformer architecture with one encoder
layer and one decoder layer.
[[IMAGE:f9418c188bfff09a_7_7]]
Based on the above data, answer the given subquestions.
How many parameters are there in the multi-head attention layer of the encoder (exclude the
parameters in the WO matrix used for linear transformation)?

A published solution is not available for this question yet.
Question 190 MCQ · 4.0 marks
[[IMAGE:f9418c188bfff09a_8_8]]

[[IMAGE:f9418c188bfff09a_9_9]]

[[IMAGE:f9418c188bfff09a_9_10]]

[[IMAGE:f9418c188bfff09a_9_11]]

[[IMAGE:f9418c188bfff09a_9_12]]

A published solution is not available for this question yet.
Question 191 MCQ · 4.0 marks
Suppose we run the pre-trained GPT model in an autoregressive fashion for generating a text
sequence. Then, the statement that “ We do not need to add the positional embedding to the
predicted tokens at every time step” is
TRUE
FALSE
A published solution is not available for this question yet.
Question 192 MSQ · 3.0 marks
Suppose we have a dataset for machine translation tasks with thousands of samples. Suppose a
team considers training the transformer model. The model could be trained in two approaches
A Autoregressive training
B Teacher forcing
Choose the correct statements
Approach A helps the model to converge faster than approach B
Approach B helps the model converge faster than approach A
One can start the training with approach B first and then switch to approach A
after some training steps
Once the training starts with approach B and then switching to approach A
after some training steps can not be done
A published solution is not available for this question yet.
Question 193 MSQ · 4.0 marks
[[IMAGE:f9418c188bfff09a_11_13]]

512
1024
1536
2048
A published solution is not available for this question yet.
Question 194 MSQ · 4.0 marks
Choose the correct statements
All the parameters of the GPT model are randomly initialized during pre-
training
The model minimizes the CLM objective during fine-tuning to improve the
performance
All parameters of the GPT model are randomly initialized during fine-tuning
In general, fine-tuning requires a dataset with labels
A published solution is not available for this question yet.
Question 195 MSQ · 4.0 marks
[[IMAGE:f9418c188bfff09a_12_14]]

Top-k sampling
Beam search with the beam size K = 2
Top-p sampling
Exhaustive search
None of these
A published solution is not available for this question yet.
Question 196 MCQ · 4.0 marks
Consider a GPT model used for Causal language modelling. We feed the input sentence “This is a
cool idea” to the model by appending special staring [BOS] and ending [EOS] tokens ( that is,
“[BOS] This is a cool idea [EOS]”). Assume the context length of the model is 7. The attention matrix
computed in one of the attention layers is given below
[[IMAGE:f9418c188bfff09a_13_15]]
Based on the above data, answer the given subquestions.
The attention score matrix given in Table 1 is appropriate for the causal language modelling task

True
False
Insufficient information
A published solution is not available for this question yet.
Question 197 NAT · 4.0 marks
Consider a GPT model used for Causal language modelling. We feed the input sentence “This is a
cool idea” to the model by appending special staring [BOS] and ending [EOS] tokens ( that is,
“[BOS] This is a cool idea [EOS]”). Assume the context length of the model is 7. The attention matrix
computed in one of the attention layers is given below
[[IMAGE:f9418c188bfff09a_13_15]]
Based on the above data, answer the given subquestions.
Assume the time step starts from t = 0 and ends at t = 6. Suppose the model is at time step t = 4,
what is the attention value assigned for the word “is”?

A published solution is not available for this question yet.
Question 198 MCQ · 3.0 marks
The statement that “the Next Sentence Prediction (NSP) task requires the BERT model to run
autoregressively given the first sentence( or segment)” is
TRUE
FALSE
**BDBN**
**Section Id :** 64065351462
**Section Number :** 13
A published solution is not available for this question yet.