da5004_2025T3_Q1_NA.pdf
Large Language Models · Quiz 1 · Sep 2025
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 173 MCQ · 2.0 marks
[[IMAGE:30260f7a668f5932_2_1]]

To increase the numerical precision of the attention scores.
To reduce the overall computational cost of the attention mechanism.
To ensure that the sum of the attention scores for each query equals 1.
To avoid numerical issues and loss of gradients during training.
A published solution is not available for this question yet.
Question 174 MCQ · 2.0 marks
Which of the following is true regarding sinusoidal encoding?
The encoding vector of words present at even positions in the given text
sequence uses sine, and the encoding vector of words present at odd positions uses cosine.
The encoding vector of words present at odd positions in the given text
sequence uses sine, and the encoding vector of words present at even positions uses cosine.
The encoding vector for each word position contains values computed using
both sine (for even dimensions) and cosine (for odd dimensions).
A published solution is not available for this question yet.
Question 175 MCQ · 2.0 marks
[[IMAGE:30260f7a668f5932_3_2]]

‘content’
‘je’
‘suis’
‘heureux’
A published solution is not available for this question yet.
Question 176 MCQ · 2.0 marks
[[IMAGE:30260f7a668f5932_3_3]]

[[IMAGE:30260f7a668f5932_3_4]]

[[IMAGE:30260f7a668f5932_3_5]]

[[IMAGE:30260f7a668f5932_3_6]]

[[IMAGE:30260f7a668f5932_3_7]]

[[IMAGE:30260f7a668f5932_3_8]]

A published solution is not available for this question yet.
Question 177 MCQ · 2.0 marks
In the context of language model training, teacher forcing is a technique where:
The correct target tokens from the training data are used as input for the next
time step, regardless of the model’s prediction.
The model’s predictions are used as the input for the next time step, even if
they are incorrect.
The model is forced to learn without any pre-trained weights.
The model is forced to correct the input data.
A published solution is not available for this question yet.
Question 178 MCQ · 2.0 marks
[[IMAGE:30260f7a668f5932_4_9]]

[[IMAGE:30260f7a668f5932_4_10]]

[[IMAGE:30260f7a668f5932_4_11]]

[[IMAGE:30260f7a668f5932_4_12]]

[[IMAGE:30260f7a668f5932_4_13]]

A published solution is not available for this question yet.
Question 179 MCQ · 3.0 marks
[[IMAGE:30260f7a668f5932_5_14]]

[[IMAGE:30260f7a668f5932_5_15]]

[[IMAGE:30260f7a668f5932_5_16]]

[[IMAGE:30260f7a668f5932_5_17]]

[[IMAGE:30260f7a668f5932_5_18]]

A published solution is not available for this question yet.
Question 180 MCQ · 3.0 marks
[[IMAGE:30260f7a668f5932_5_19]]

[[IMAGE:30260f7a668f5932_5_20]]

[[IMAGE:30260f7a668f5932_5_21]]

[[IMAGE:30260f7a668f5932_5_22]]

[[IMAGE:30260f7a668f5932_5_23]]

[[IMAGE:30260f7a668f5932_5_24]]

A published solution is not available for this question yet.
Question 181 MCQ · 3.0 marks
For a vocabulary size of 10, how many beams will there be in greedy search, beam search with
beam size 3, and exhaustive search at step 1 and step 3, respectively?
Step 1:
Greedy: 1, Beam: 1, Exhaustive: 1
Step 3:
Greedy: 1, Beam: 1, Exhaustive: 1
Step 1:
Greedy: 1, Beam: 3, Exhaustive: 10
Step 3:
Greedy: 1, Beam: 3, Exhaustive: 10
Step 1:
Greedy: 1, Beam: 3, Exhaustive: 1
Step 3:
Greedy: 1, Beam: 3, Exhaustive: 100
Step 1:
Greedy: 1, Beam: 3, Exhaustive: 10
Step 3:
Greedy: 1, Beam: 3, Exhaustive: 1000
A published solution is not available for this question yet.
Question 182 MSQ · 2.0 marks
Which of the following decoding strategies is/are inherently non-deterministic?
Beam
Greedy
Nucleus
Exhaustive
A published solution is not available for this question yet.
Question 183 MSQ · 2.0 marks
Which of the following statements about multiple attention heads in a Transformer are true?
They allow learning diverse local/global relations.
They reduce memory usage compared to one head.
Increasing the number of heads while keeping dmodel fixed, decreases the
dimensionality handled by each head.
All heads necessarily learn completely independent information about the
input.
A published solution is not available for this question yet.
Question 184 MSQ · 4.0 marks
Which of the following statements about the Transformer architecture are true?
Masking (look-ahead mask) is applied in every decoder self-attention layer, not
just the first one.
Token embeddings are used in the encoder but not in the decoder.
Positional encoding is required only in the encoder, not in the decoder.
Token embeddings are used in both the encoder and the decoder.
Positional encodings are added to the input embeddings in both the encoder
and the decoder.
A published solution is not available for this question yet.
Question 185 MSQ · 4.0 marks
[[IMAGE:30260f7a668f5932_7_25]]

Beam search with beam size 2 will keep either A or B as per probability
whereas Top-2 sampling will keep both A and B.
Beam search with beam size 2 will keep both A and B whereas Top-2 sampling
will keep either A or B.
Beam search with beam size 2 will randomly choose any two tokens, whereas
Top-2 sampling will always pick the two least probable tokens.
Beam search with beam size 2 will always keep the top-2 most probable
tokens, whereas Top-2 sampling will restrict choices to top-2 tokens and then sample one
probabilistically.
A published solution is not available for this question yet.
Question 186 NAT · 3.0 marks
[[IMAGE:30260f7a668f5932_8_26]]

A published solution is not available for this question yet.
Question 187 NAT · 3.0 marks
[[IMAGE:30260f7a668f5932_8_27]]

A published solution is not available for this question yet.
Question 188 NAT · 2.0 marks
[[IMAGE:30260f7a668f5932_9_28]]

A published solution is not available for this question yet.
Question 189 NAT · 3.0 marks
[[IMAGE:30260f7a668f5932_9_29]]
How many next-token prediction targets are generated during training a full batch?

A published solution is not available for this question yet.
Question 190 NAT · 2.0 marks
[[IMAGE:30260f7a668f5932_9_29]]
For one sequence, how many non-zero attention scores remain in the masked attention score
matrix (per head)?

A published solution is not available for this question yet.
Question 191 NAT · 2.0 marks
[[IMAGE:30260f7a668f5932_9_29]]
[[IMAGE:30260f7a668f5932_10_30]]


A published solution is not available for this question yet.