da5004_2026T2_Q1_NA.pdf
Large Language Models · Quiz 1 · May 2026
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 MSQ · 4.0 marks
Which statements are true about traditional attention used in sequence-to-sequence models?
The decoder decides which encoder states are important.
Attention weights are computed over all decoder hidden states.
The context vector is a weighted sum of encoder states.
Encoder hidden states are ignored after encoding.
A published solution is not available for this question yet.
Question 3 NAT · 3.0 marks
Consider the text "learn easy math", consisting of three tokens: {learn,easy,math}.
The token embeddings are arranged in the input matrix [[IMAGE:7ca18a647a21b6db_2_2]] :
[[IMAGE:7ca18a647a21b6db_2_3]]
The query, key, and value projection matrices are given by:
[[IMAGE:7ca18a647a21b6db_2_4]]
[[IMAGE:7ca18a647a21b6db_3_5]]
[[IMAGE:7ca18a647a21b6db_3_6]]
[[IMAGE:7ca18a647a21b6db_3_7]]
Based on the above data, answer the given subquestions.
Compute the [[IMAGE:7ca18a647a21b6db_3_8]] of the final scaled dot-product attention and submit the sum of the diagonal
elements of [[IMAGE:7ca18a647a21b6db_3_9]] matrix.








A published solution is not available for this question yet.
Question 4 MCQ · 3.0 marks
Consider the text "learn easy math", consisting of three tokens: {learn,easy,math}.
The token embeddings are arranged in the input matrix [[IMAGE:7ca18a647a21b6db_2_2]] :
[[IMAGE:7ca18a647a21b6db_2_3]]
The query, key, and value projection matrices are given by:
[[IMAGE:7ca18a647a21b6db_2_4]]
[[IMAGE:7ca18a647a21b6db_3_5]]
[[IMAGE:7ca18a647a21b6db_3_6]]
[[IMAGE:7ca18a647a21b6db_3_7]]
Based on the above data, answer the given subquestions.
Choose the token pair with the least attention score.






{"learn", "easy"}
{"learn", "math"}
{"easy", "math"}
{"learn", "learn"}
{"easy" , "easy"}
A published solution is not available for this question yet.
Question 5 NAT · 4.0 marks
A Transformer model processes a sequence containing 6 tokens using Multi-Head Attention. The
model initially uses 4 attention heads. If the number of attention heads is doubled to 8, while the
sequence length remains unchanged, how many **additional** attention scores are computed across
all heads?
A published solution is not available for this question yet.
Question 6 NAT · 4.0 marks
A batch of 8 sentences is passed through a Transformer Encoder consisting of 6 encoder layers
and 8 attention heads. After padding, each sentence has a maximum length of 32 tokens. The
model uses an embedding dimension of [[IMAGE:7ca18a647a21b6db_4_10]] . What is the total number of elements
(volume) in the output tensor produced by the final encoder layer for the entire batch?

A published solution is not available for this question yet.
Question 7 NAT · 4.0 marks
Consider a masked multi-head attention layer in a transformer decoder. The source vocabulary is
of size [[IMAGE:7ca18a647a21b6db_4_11]] , [[IMAGE:7ca18a647a21b6db_4_12]] and context length [[IMAGE:7ca18a647a21b6db_4_13]] . To enforce autoregressive
generation, a causal mask is applied to the raw attention score before applying softmax. How
many entries in the causal mask are set to [[IMAGE:7ca18a647a21b6db_4_14]] ?




A published solution is not available for this question yet.
Question 8 NAT · 4.0 marks
Consider a Transformer model with a hidden state dimension [[IMAGE:7ca18a647a21b6db_5_15]] . The positional
encoding for a given position, [[IMAGE:7ca18a647a21b6db_5_16]] and dimension [[IMAGE:7ca18a647a21b6db_5_17]] is defined by:
[[IMAGE:7ca18a647a21b6db_5_18]]
[[IMAGE:7ca18a647a21b6db_5_19]]
Compute the squared euclidean norm for the vector corresponding to [[IMAGE:7ca18a647a21b6db_5_20]] and enter the
value.






A published solution is not available for this question yet.
Question 9 MCQ · 4.0 marks
Consider the following statements regarding the use of Teacher Forcing while training
autoregressive sequence-to-sequence models and select the appropriate option:
Statement 1: Teacher Forcing involves passing the input token to a larger Teacher model whose
output is then input as the next token
Statement 2: Teacher Forcing helps to prevent compounding of early incorrect predictions that
can destabilize training
Statement 1 is True but Statement 2 is False
Statement 1 is False but Statement 2 is True
Both the statements are True
Both the statements are False
A published solution is not available for this question yet.
Question 10 NAT · 3.0 marks
[[IMAGE:7ca18a647a21b6db_6_21]]
Based on the above data, answer the given subquestions.
Compute the joint probability of the sentence :
[[IMAGE:7ca18a647a21b6db_6_22]]
(Provide answer correct upto 2 digits after the decimal)


A published solution is not available for this question yet.
Question 11 MCQ · 3.0 marks
[[IMAGE:7ca18a647a21b6db_6_21]]
Based on the above data, answer the given subquestions.
Choose the option corresponding to the sentence with the highest joint probability under the
given language model.

"i eat apples"
"apples like You"
"i eat bananas"
"you like bananas
"you like apples
A published solution is not available for this question yet.
Question 12 NAT · 2.0 marks
[[IMAGE:7ca18a647a21b6db_6_21]]
Based on the above data, answer the given subquestions.
Using the given probability tree calculate the conditional probability of predicting "apples" given
"like" [i.e P(apples | like)]. (Submit -1 if the provided information is insufficient)

A published solution is not available for this question yet.
Question 13 MSQ · 3.0 marks
Which of the following are components found within a GPT decoder layer?
Causal (Masked) Multi-Head Self-Attention
Position-wise Feed-Forward Network
Multi Head (Cross) Attention
Recurrent Neural Network (RNN) Layer
Add & Norm Layer
A published solution is not available for this question yet.
Question 14 NAT · 2.0 marks
The table below represents the conditional probability distribution over the vocabulary. Each
column corresponds to a decoding timestep and the values represent the conditional probability
of selecting a token at the timestep given the previously generated tokens
[[IMAGE:7ca18a647a21b6db_8_23]]
For all calculations, consider only the tokens mentioned in the first column as the Vocabulary
Based on the above data, answer the given subquestions.
Use greedy decoding to determine the most likely sequence of 6 tokens and compute the
probability of generating that sequence. Enter the value rounded off to 3 decimal places

A published solution is not available for this question yet.
Question 15 NAT · 2.0 marks
The table below represents the conditional probability distribution over the vocabulary. Each
column corresponds to a decoding timestep and the values represent the conditional probability
of selecting a token at the timestep given the previously generated tokens
[[IMAGE:7ca18a647a21b6db_8_23]]
For all calculations, consider only the tokens mentioned in the first column as the Vocabulary
Based on the above data, answer the given subquestions.
Use exhaustive search strategy to identify the most probable sequence consisting of 4 tokens.
How many decoder runs are required in total to identify the most probable sequence containing 4
tokens?

A published solution is not available for this question yet.
Question 16 NAT · 3.0 marks
The table below represents the conditional probability distribution over the vocabulary. Each
column corresponds to a decoding timestep and the values represent the conditional probability
of selecting a token at the timestep given the previously generated tokens
[[IMAGE:7ca18a647a21b6db_8_23]]
For all calculations, consider only the tokens mentioned in the first column as the Vocabulary
Based on the above data, answer the given subquestions.
Use Top-k sampling at timestep t = 4 and k = 3. Compute the renormalized probability of the most
probable token from the candidate tokens and enter the value rounded off to three decimal places

A published solution is not available for this question yet.
Question 17 NAT · 2.0 marks
The table below represents the conditional probability distribution over the vocabulary. Each
column corresponds to a decoding timestep and the values represent the conditional probability
of selecting a token at the timestep given the previously generated tokens
[[IMAGE:7ca18a647a21b6db_8_23]]
For all calculations, consider only the tokens mentioned in the first column as the Vocabulary
Based on the above data, answer the given subquestions.
Consider the sequence "the castle was abandoned". What is the minimum beam width required to
guarantee that this sequence is retained during beam search at timestep t = 4 ?

A published solution is not available for this question yet.