da5004_2026T2_ET_FN.pdf
Large Language Models · End Term · May 2026 FN
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 NAT · 2.0 marks
Consider a Multi-Head Attention sub-layer with the following hyper-parameters:
• Embedding dimension ( [[IMAGE:2ef6735b833c260e_2_2]] ) = 256
• Number of heads ( [[IMAGE:2ef6735b833c260e_2_3]] ) = 4
• key/value dimension per head ( [[IMAGE:2ef6735b833c260e_2_4]] ) = [[IMAGE:2ef6735b833c260e_2_5]] .
Calculate the total number of trainable parameters contained **only** in the projection matrices
[[IMAGE:2ef6735b833c260e_2_6]] and [[IMAGE:2ef6735b833c260e_2_7]] . (Ignore biases). Enter the integer value.






A published solution is not available for this question yet.
Question 3 NAT · 2.0 marks
For T5-Base, the model dimension is 768 and FFN dimension is 3072. Ignore bias. Find the number
of parameters in the FFN of one encoder layer.
A published solution is not available for this question yet.
Question 4 NAT · 3.0 marks
A batch of 5 sentences is passed through a Transformer Encoder. The sequences are padded to a
maximum length of 20 tokens. The embedding dimension ( [[IMAGE:2ef6735b833c260e_3_8]] ) is 64. What is the total number
of elements (volume) in the output tensor produced by the final layer of the encoder for this
batch?

A published solution is not available for this question yet.
Question 5 NAT · 3.0 marks
Calculate the KV Cache memory size in Megabytes for a single request (batch size = 1) with the
following parameters:
• Number of layers: 24
• Embedding dimension (d_model): 1024
• Sequence length (n): 2048 tokens
• Precision: 16-bit (2 bytes per element)
• Assume standard Multi-Head Attention (MHA) where total KV dimension per layer is 2 x d_model
(2 for K+V)
A published solution is not available for this question yet.
Question 6 MCQ · 2.0 marks
In BERT's Next Sentence Prediction task, how are the negative (NotNext) examples created?
By taking the sentence that immediately precedes the current sentence.
By reversing the word order of the actual next sentence.
By picking a random sentence from the corpus (50% of the time).
By masking all verbs in the actual next sentence.
A published solution is not available for this question yet.
Question 7 MCQ · 2.0 marks
What happens when a very large model is trained for many epochs on a relatively small dataset in
the T5 scaling-law setting?
The performance continues to improve linearly with more training steps.
The performance degrades due to overfitting (memorization) of the training
data.
The model automatically learns to augment the data, preventing overfitting.
The performance plateaus but does not degrade.
A published solution is not available for this question yet.
Question 8 MCQ · 2.0 marks
**PagedAttention** is designed to address which specific inefficiency in Large Language Model (LLM)
serving?
High computational latency of matrix multiplication.
Memory fragmentation and waste in the KV cache due to pre-allocating
contiguous memory blocks for variable-length sequences.
Low bandwidth of the internet connection during API calls.
The difficulty of training models on long sequences.
A published solution is not available for this question yet.
Question 9 MCQ · 2.0 marks
The standard Self-Attention mechanism has a time complexity of [[IMAGE:2ef6735b833c260e_5_9]] , where [[IMAGE:2ef6735b833c260e_5_10]] is the
sequence length and [[IMAGE:2ef6735b833c260e_5_11]] is the embedding dimension. In Kernel-based Linear Attention methods
(like Linear Transformers), the associativity of matrix multiplication is exploited to compute
[[IMAGE:2ef6735b833c260e_5_12]] instead of [[IMAGE:2ef6735b833c260e_5_13]] .
What is the resulting time complexity of this Linear Attention mechanism?





[[IMAGE:2ef6735b833c260e_5_14]]

[[IMAGE:2ef6735b833c260e_5_15]]

[[IMAGE:2ef6735b833c260e_5_16]]

[[IMAGE:2ef6735b833c260e_5_17]]

A published solution is not available for this question yet.
Question 10 MCQ · 2.0 marks
In a standard Transformer layer utilizing Rotary Positional Embeddings (RoPE), at which specific
stage of the forward pass is the rotary transformation applied?
To the input embeddings [[IMAGE:2ef6735b833c260e_5_18]] immediately before the Linear projections for [[IMAGE:2ef6735b833c260e_5_19]] ,
[[IMAGE:2ef6735b833c260e_5_20]] , and [[IMAGE:2ef6735b833c260e_5_21]] .




To the [[IMAGE:2ef6735b833c260e_5_22]] and [[IMAGE:2ef6735b833c260e_5_23]] vectors simultaneously using a coupled weight matrix within
the Attention head.


To the [[IMAGE:2ef6735b833c260e_5_24]] and [[IMAGE:2ef6735b833c260e_5_25]] vectors independently, after their linear projections but
before the scaled dot-product calculation.


To the [[IMAGE:2ef6735b833c260e_5_26]] and [[IMAGE:2ef6735b833c260e_5_27]] vectors after the linear projections to ensure the entire
latent space is rotationally invariant.


A published solution is not available for this question yet.
Question 11 MCQ · 2.0 marks
In a Transformer using ALiBi, the attention score is modified as:
[[IMAGE:2ef6735b833c260e_6_28]]
Suppose a researcher proposes modifying ALiBi to:
[[IMAGE:2ef6735b833c260e_6_29]]
and claims that:
“The model will still learn to focus on nearby tokens through training.”
Which of the following is the most accurate critique of this proposal?


The modification will not affect attention behavior significantly because
softmax normalizes all scores, making additive biases irrelevant.
The modification will cause instability because the attention scores will
become negative for nearby tokens and positive for distant tokens.
The modification is equivalent to scaling (Q) and (K) vectors differently and
therefore does not fundamentally change the positional bias.
The modification will bias attention toward distant tokens, and due to the
exponential nature of softmax, this effect can dominate learned [[IMAGE:2ef6735b833c260e_6_30]] similarities, making it
difficult for training to recover locality.

A published solution is not available for this question yet.
Question 12 MSQ · 3.0 marks
Comparing the architectures of BERT and GPT, select all the **structural differences** that are true.
BERT uses bidirectional self-attention, allowing tokens to attend to both left
and right contexts.
GPT uses masked (causal) self-attention, allowing tokens to attend only to
previous tokens.
BERT is an autoregressive model, while GPT is an autoencoding model.
During pre-training, BERT predicts masked tokens, whereas GPT predicts the
next token in the sequence.
A published solution is not available for this question yet.
Question 13 MSQ · 3.0 marks
Based on the architectural comparison experiments (Encoder-Decoder vs. Decoder-only vs.
Encoder-only) discussed in the lectures:
Encoder-Decoder architectures generally perform best for sequence-to-
sequence tasks like Translation and Summarization.
A Decoder-only model with a **Prefix-LM** objective can match the performance
of an Encoder-Decoder model.
Sharing parameters between the encoder and decoder always improves
performance regardless of task.
Encoder-only models (like BERT) are typically less suitable for generation tasks
compared to Encoder-Decoder models.
A published solution is not available for this question yet.
Question 14 MSQ · 3.0 marks
Select all true statements regarding **KV Caching** during the autoregressive inference of LLMs.
It trades off increased memory usage for reduced computational latency.
It avoids recomputing the Key and Value vectors for tokens that have already
been processed in previous steps.
It is essential during the training phase to speed up backpropagation.
The memory required for the KV cache grows linearly with the sequence
length and batch size.
A published solution is not available for this question yet.
Question 15 MSQ · 2.0 marks
Select all correct statements regarding Decoding Strategies in language generation.
Beam Search with a beam width [[IMAGE:2ef6735b833c260e_8_31]] is mathematically equivalent to Greedy
Search.

Greedy search is deterministic and often leads to repetitive or degenerative
text loops.
Top- [[IMAGE:2ef6735b833c260e_8_32]] sampling guarantees that the generated text will always be
grammatically correct.

Top- [[IMAGE:2ef6735b833c260e_8_33]] (Nucleus) sampling allows for a dynamic vocabulary size at each step,
whereas Top- [[IMAGE:2ef6735b833c260e_8_34]] uses a fixed number of candidates.


A published solution is not available for this question yet.
Question 16 MSQ · 2.0 marks
In a feature-based approach using a pretrained transformer (like BERT) for a Multiple-Choice QA
task, which of the following statements regarding the methodology are correct? (Select all that
apply)
The weights of the pretrained transformer are updated via backpropagation
during the training phase.
The input must be formatted as multiple pairs, such as [CLS] Question [SEP]
Choice_n, creating a distinct representation for each candidate answer.
The final prediction is determined by passing the fixed output representations
through a trainable linear layer and applying a Softmax function.
The model generates the correct answer string token-by-token using an
autoregressive decoding strategy like Greedy Search.
A published solution is not available for this question yet.
Question 17 MCQ · 3.0 marks
A Transformer processes a sequence of length 5 using 6 identical self-attention layers.
Count the total number of allowed attention links across all layers and Choose the option
representing the correct pair
(A) BERT
(B) GPT
BERT: 150, GPT: 90
BERT: 90, GPT: 150
BERT: 150, GPT: 150
BERT: 90, GPT: 90
BERT: 30, GPT: 36
BERT: 36, GPT: 30
A published solution is not available for this question yet.
Question 18 MCQ · 3.0 marks
Consider the architecture of **BART**. Which of the following descriptions best match its structural
design?
A Decoder-only model similar to GPT, but with bidirectional attention in the
first layer.
An Encoder-only model similar to BERT, but trained with a causal masking
objective.
An encoder-decoder model where the encoder processes a corrupted input
bidirectionally and the decoder autoregressively reconstructs the original sequence.
A dual-encoder model where one encoder processes the context and another
processes the query.
A published solution is not available for this question yet.
Question 19 MCQ · 3.0 marks
Why is **Flash Attention** considered an "IO-aware" algorithm?
It reduces the number of parameters in the model to fit in GPU memory.
It compresses the input data using JPEG-like encoding before processing.
It optimizes the movement of data between the high-bandwidth memory
(HBM) and the faster on-chip SRAM to avoid memory bandwidth bottlenecks.
It writes all intermediate attention matrices to the hard disk to save RAM.
A published solution is not available for this question yet.
Question 20 MCQ · 3.0 marks
In a transformer with relative positional encoding (RPE), attention logits between a query at
position 4 and keys at positions 2 and 6 depend on relative distances [[IMAGE:2ef6735b833c260e_9_35]] where [[IMAGE:2ef6735b833c260e_9_36]] is the
position of query and [[IMAGE:2ef6735b833c260e_9_37]] is the position of key.
Suppose the learned scalar biases as per RPE are:
• [[IMAGE:2ef6735b833c260e_10_38]]
• [[IMAGE:2ef6735b833c260e_10_39]]
• [[IMAGE:2ef6735b833c260e_10_40]]
• [[IMAGE:2ef6735b833c260e_10_41]]
• [[IMAGE:2ef6735b833c260e_10_42]]
The base (content-only) attention logits are:
• [[IMAGE:2ef6735b833c260e_10_43]]
• [[IMAGE:2ef6735b833c260e_10_44]]
What are the final attention logits after incorporating RPE?










[[IMAGE:2ef6735b833c260e_10_45]]

[[IMAGE:2ef6735b833c260e_10_46]]

[[IMAGE:2ef6735b833c260e_10_47]]

[[IMAGE:2ef6735b833c260e_10_48]]

A published solution is not available for this question yet.
Question 21 MCQ · 3.0 marks
An engineer is designing a Transformer model capable of length extrapolation (handling
sequences longer than those seen during training) while minimizing the number of trainable
weights. They are evaluating Absolute Positional Encodings (APE), Relative Positional Encodings
(RPE), Rotary Positional Encodings (RoPE), and Attention with Linear Biases (ALiBi).Which of the
following statements correctly describes the parameterization and behavior of these techniques?
RoPE is considered a learned encoding because the model must optimize the
rotation angles [[IMAGE:2ef6735b833c260e_10_49]] as trainable parameters during the backpropagation phase.

The engineer should choose APE or RPE to minimize weights, as both rely
exclusively on fixed sinusoidal functions without requiring a lookup table.
RoPE and ALiBi are preferable for this use case because they utilize fixed
mathematical transformations, whereas standard APE and RPE typically require learning position-
specific embedding weights.
ALiBi is the only technique among the four that requires the model to learn a
unique "slope" parameter for each attention head to determine the rate of penalty decay.
A published solution is not available for this question yet.