da5004_2026T1_Q1_NA.pdf
Large Language Models · Quiz 1 · Jan 2026
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 MCQ · 2.0 marks
In a self-attention mechanism, let the input be a matrix [[IMAGE:36b97be09a59b1a1_3_2]] , where [[IMAGE:36b97be09a59b1a1_3_3]] is the sequence
length. If the dimension of the query projection matrix is [[IMAGE:36b97be09a59b1a1_3_4]] , what is the
dimension of the Query matrix [[IMAGE:36b97be09a59b1a1_3_5]] ?




[[IMAGE:36b97be09a59b1a1_3_6]]

[[IMAGE:36b97be09a59b1a1_3_7]]

[[IMAGE:36b97be09a59b1a1_3_8]]

[[IMAGE:36b97be09a59b1a1_3_9]]

A published solution is not available for this question yet.
Question 3 MCQ · 2.0 marks
[[IMAGE:36b97be09a59b1a1_3_10]]
In the scaled dot-product attention equation , what is
the primary reason for dividing the dot product by [[IMAGE:36b97be09a59b1a1_4_11]] ?


To ensure the matrix dimensions match for multiplication with [[IMAGE:36b97be09a59b1a1_4_12]]

To reduce the number of trainable parameters in the model
To prevent the dot product values from growing too large, which would push
the softmax function into regions with extremely small gradients
To normalize the embedding vectors to have a unit length
A published solution is not available for this question yet.
Question 4 MCQ · 2.0 marks
Consider a Transformer Encoder stack consisting of [[IMAGE:36b97be09a59b1a1_4_13]] identical layers. If the input to the first
layer is a sequence of word embeddings of shape [[IMAGE:36b97be09a59b1a1_4_14]] (where [[IMAGE:36b97be09a59b1a1_4_15]] and [[IMAGE:36b97be09a59b1a1_4_16]] ),
what is the shape of the output of the final (6th) layer?




[[IMAGE:36b97be09a59b1a1_4_17]] (A single context vector)

[[IMAGE:36b97be09a59b1a1_4_18]] (One vector per layer)

[[IMAGE:36b97be09a59b1a1_4_19]] (One vector per token, preserving sequence length)

[[IMAGE:36b97be09a59b1a1_4_20]] (Concatenation of all token vectors across layers)

A published solution is not available for this question yet.
Question 5 MCQ · 2.0 marks
In a standard Transformer encoder-decoder architecture, is masking typically applied in the cross-
attention layer?
Yes, both causal masking and padding masking are applied to prevent the
decoder from attending to future encoder positions
Yes, only causal masking is applied to maintain the autoregressive property
No causal masking is needed, but padding masking may be applied to ignore
padded positions in the source sequence
No masking is ever applied in cross-attention since the encoder processes the
complete input sequence
A published solution is not available for this question yet.
Question 6 MCQ · 2.0 marks
Consider the BERT architecture. During pre-training, it uses segment embeddings ( [[IMAGE:36b97be09a59b1a1_5_21]] ) to
distinguish between sentences. However, during fine-tuning for a **single-sentence classification**
**task** (e.g., Sentiment Analysis), how are these segment embeddings typically utilized?

Segment embeddings are not used; the embedding layer is disabled for
single-sentence tasks.
All tokens in the input sequence are assigned the same segment embedding
(e.g., [[IMAGE:36b97be09a59b1a1_5_22]] ).

The first half of the sentence gets [[IMAGE:36b97be09a59b1a1_5_23]] and the second half gets [[IMAGE:36b97be09a59b1a1_5_24]] to maintain
symmetry.


Segment embeddings are randomly initialized for every new single-sentence
input.
A published solution is not available for this question yet.
Question 7 MCQ · 2.0 marks
Consider a decoding step in a language model where the logits for the next token are
[[IMAGE:36b97be09a59b1a1_5_25]] . We apply a temperature scaling [[IMAGE:36b97be09a59b1a1_5_26]] before the softmax. As [[IMAGE:36b97be09a59b1a1_5_27]] (approaches
zero), what does the resulting probability distribution [[IMAGE:36b97be09a59b1a1_5_28]] look like?




It approaches a uniform distribution where all tokens have equal probability.
It approaches a "one-hot" distribution where the token with the highest logit
has probability 1.0 (Greedy Search).
It remains unchanged from the original softmax distribution.
It causes numerical instability and results in NaNs.
A published solution is not available for this question yet.
Question 8 MCQ · 2.0 marks
You are fine-tuning a pre-trained BERT model for a spam classification task. The input sequence is:
[[IMAGE:36b97be09a59b1a1_5_29]] . Which vector from the final layer output is typically
passed to the classification head ( [[IMAGE:36b97be09a59b1a1_5_30]] )?


The average pooling of all token vectors.
The final hidden state vector corresponding to the [[IMAGE:36b97be09a59b1a1_6_31]] token ( [[IMAGE:36b97be09a59b1a1_6_32]] ).


The final hidden state vector corresponding to the [[IMAGE:36b97be09a59b1a1_6_33]] token ( [[IMAGE:36b97be09a59b1a1_6_34]] ).


The concatenation of all hidden state vectors.
A published solution is not available for this question yet.
Question 9 MSQ · 3.0 marks
Select all statements that correctly describe the motivation and behavior of **Multi-Head Attention**
as compared to single-head attention.
It allows the model to jointly attend to information from different
representation subspaces at different positions.
It is mathematically similar to having multiple filters/kernels in a CNN to
capture different features.
It reduces the total number of parameters required compared to a single head
with the same total dimension.
Each head can theoretically learn to capture different linguistic relationships
(e.g., one head links "it" to "animal", another links "it" to "tired").
A published solution is not available for this question yet.
Question 10 MSQ · 3.0 marks
[[IMAGE:36b97be09a59b1a1_6_35]]
Given the vectorized self-attention calculation , select all true
statements regarding the matrix dimensions and operations. Assume [[IMAGE:36b97be09a59b1a1_6_36]] .


The matrix resulting from [[IMAGE:36b97be09a59b1a1_6_37]] has dimensions [[IMAGE:36b97be09a59b1a1_6_38]] .


The softmax operation is applied to the entire matrix at once (globally), not
row-wise.
The final output [[IMAGE:36b97be09a59b1a1_6_39]] has dimensions [[IMAGE:36b97be09a59b1a1_6_40]] .


The matrix [[IMAGE:36b97be09a59b1a1_7_41]] represents the raw affinity/score between every pair of
tokens in the sequence.

A published solution is not available for this question yet.
Question 11 MSQ · 2.0 marks
In the context of the Transformer architecture, which of the following statements correctly
describe the source of the Query ( [[IMAGE:36b97be09a59b1a1_7_42]] ), Key ( [[IMAGE:36b97be09a59b1a1_7_43]] ), and Value ( [[IMAGE:36b97be09a59b1a1_7_44]] ) vectors during the Cross-Attention
(Encoder-Decoder Attention) mechanism?



The Keys ( [[IMAGE:36b97be09a59b1a1_7_45]] ) and Values ( [[IMAGE:36b97be09a59b1a1_7_46]] ) are derived from the output of the Decoder.


The Queries ( [[IMAGE:36b97be09a59b1a1_7_47]] ) are generated from the previous hidden states of the
Decoder.

The Queries ( [[IMAGE:36b97be09a59b1a1_7_48]] ), Keys ( [[IMAGE:36b97be09a59b1a1_7_49]] ), and Values ( [[IMAGE:36b97be09a59b1a1_7_50]] ) all originate from the same input
sequence.



This mechanism allows the Decoder to focus on relevant parts of the input
sequence processed by the Encoder.
A published solution is not available for this question yet.
Question 12 NAT · 2.0 marks
Consider the positional encoding for a given position [[IMAGE:36b97be09a59b1a1_7_51]] and dimension [[IMAGE:36b97be09a59b1a1_7_52]] is defined by:
[[IMAGE:36b97be09a59b1a1_7_53]]
[[IMAGE:36b97be09a59b1a1_7_54]]
Consider a Transformer model with a hidden state dimension [[IMAGE:36b97be09a59b1a1_8_55]] . You are calculating
the positional embedding for the 4th word in a sequence (index [[IMAGE:36b97be09a59b1a1_8_56]] , using 0-based indexing).
Calculate the value of the 6th element (index [[IMAGE:36b97be09a59b1a1_8_57]] ) of the positional embedding vector for this
word. (use radians for angle)







A published solution is not available for this question yet.
Question 13 NAT · 1.0 marks
Suppose you are given three encoder hidden states at time [[IMAGE:36b97be09a59b1a1_8_58]] :
[[IMAGE:36b97be09a59b1a1_8_59]]
The previous decoder hidden state is:
[[IMAGE:36b97be09a59b1a1_8_60]]
Given the attention score function:
[[IMAGE:36b97be09a59b1a1_8_61]]
where the hyperbolic tangent function is defined as:
[[IMAGE:36b97be09a59b1a1_8_62]]
and
[[IMAGE:36b97be09a59b1a1_9_63]]
Based on the above data, answer the given subquestions.
Compute the attention score for hidden state [[IMAGE:36b97be09a59b1a1_9_64]] using the given function. Enter the final answer
correct to two decimal places.







A published solution is not available for this question yet.
Question 14 NAT · 3.0 marks
Suppose you are given three encoder hidden states at time [[IMAGE:36b97be09a59b1a1_8_58]] :
[[IMAGE:36b97be09a59b1a1_8_59]]
The previous decoder hidden state is:
[[IMAGE:36b97be09a59b1a1_8_60]]
Given the attention score function:
[[IMAGE:36b97be09a59b1a1_8_61]]
where the hyperbolic tangent function is defined as:
[[IMAGE:36b97be09a59b1a1_8_62]]
and
[[IMAGE:36b97be09a59b1a1_9_63]]
Based on the above data, answer the given subquestions.
Normalize the attention scores using the softmax function to obtain the attention weights [[IMAGE:36b97be09a59b1a1_9_65]] .
Submit [[IMAGE:36b97be09a59b1a1_9_66]] (i.e., the first element of the [[IMAGE:36b97be09a59b1a1_9_67]] vector). Enter the final answer correct to two decimal
places.
[[IMAGE:36b97be09a59b1a1_9_68]]










A published solution is not available for this question yet.
Question 15 NAT · 2.0 marks
Suppose you are given three encoder hidden states at time [[IMAGE:36b97be09a59b1a1_8_58]] :
[[IMAGE:36b97be09a59b1a1_8_59]]
The previous decoder hidden state is:
[[IMAGE:36b97be09a59b1a1_8_60]]
Given the attention score function:
[[IMAGE:36b97be09a59b1a1_8_61]]
where the hyperbolic tangent function is defined as:
[[IMAGE:36b97be09a59b1a1_8_62]]
and
[[IMAGE:36b97be09a59b1a1_9_63]]
Based on the above data, answer the given subquestions.
Calculate the context vector [[IMAGE:36b97be09a59b1a1_10_69]] :
[[IMAGE:36b97be09a59b1a1_10_70]]
Provide the **sum of all elements** of [[IMAGE:36b97be09a59b1a1_10_71]] . Enter the final answer correct to two decimal places.









A published solution is not available for this question yet.
Question 16 MCQ · 2.0 marks
Consider the following Configuration for a GPT model:
• Vocabulary size: 40,000 tokens
• Embedding dimension (d_model): 768
• Maximum sequence length: 512
• Number of transformer blocks: 12
• Number of attention heads per block: 12
• Feed-forward network hidden dimension: 3,072
• Activation function: GELU
Note: In multi-head attention, the model dimension is split equally among all heads.
Based on the above data, answer the given subquestions.
What is the number of parameters in the token embedding matrix?
30,720,000
40,768
393,216
15,728,640,000
A published solution is not available for this question yet.
Question 17 MCQ · 2.0 marks
Consider the following Configuration for a GPT model:
• Vocabulary size: 40,000 tokens
• Embedding dimension (d_model): 768
• Maximum sequence length: 512
• Number of transformer blocks: 12
• Number of attention heads per block: 12
• Feed-forward network hidden dimension: 3,072
• Activation function: GELU
Note: In multi-head attention, the model dimension is split equally among all heads.
Based on the above data, answer the given subquestions.
What is the total number of parameters in the positional embedding matrix?
512,768
393,216
3,932,160
39,321
A published solution is not available for this question yet.
Question 18 MCQ · 2.0 marks
Consider the following Configuration for a GPT model:
• Vocabulary size: 40,000 tokens
• Embedding dimension (d_model): 768
• Maximum sequence length: 512
• Number of transformer blocks: 12
• Number of attention heads per block: 12
• Feed-forward network hidden dimension: 3,072
• Activation function: GELU
Note: In multi-head attention, the model dimension is split equally among all heads.
Based on the above data, answer the given subquestions.
For ONE complete attention head, what is the total number of parameters (Q + K + V projections
combined)?
49,152
196,608
589,824
147,456
A published solution is not available for this question yet.
Question 19 MCQ · 2.0 marks
Consider the following Configuration for a GPT model:
• Vocabulary size: 40,000 tokens
• Embedding dimension (d_model): 768
• Maximum sequence length: 512
• Number of transformer blocks: 12
• Number of attention heads per block: 12
• Feed-forward network hidden dimension: 3,072
• Activation function: GELU
Note: In multi-head attention, the model dimension is split equally among all heads.
Based on the above data, answer the given subquestions.
After the multi-head attention computation, all head outputs are concatenated and projected back
to the model dimension.
How many parameters are in the output projection matrix (weights only, excluding bias)?
768
49,152
589,824
294,912
A published solution is not available for this question yet.
Question 20 MCQ · 2.0 marks
Consider the following Configuration for a GPT model:
• Vocabulary size: 40,000 tokens
• Embedding dimension (d_model): 768
• Maximum sequence length: 512
• Number of transformer blocks: 12
• Number of attention heads per block: 12
• Feed-forward network hidden dimension: 3,072
• Activation function: GELU
Note: In multi-head attention, the model dimension is split equally among all heads.
Based on the above data, answer the given subquestions.
What is the total number of parameters in the complete FFN (including both weight matrices and
bias vectors)?
2,359,296
4,718,592
4,722,432
9,437,184
A published solution is not available for this question yet.