da5004_2026T2_Q2_NA.pdf
Large Language Models · Quiz 2 · May 2026
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 SHORT_TEXT · 2.0 marks
Answer the subquestions for the given transformer architecture for one (N = 1) **Encoder Decoder**
**Block**
[[IMAGE:f854547ec11886a8_2_2]]
In the given Transformer architecture for one Encoder Decoder block, identify every **Add & Norm**
layer. Submit the corresponding component alphabets in **alphabetical (ascending) order** as a
single uppercase string without spaces, commas, or quotation marks.
Example: If the correct components are Q, G, and S, your answer should be GQS

A published solution is not available for this question yet.
Question 3 MCQ · 1.0 marks
Answer the subquestions for the given transformer architecture for one (N = 1) **Encoder Decoder**
**Block**
[[IMAGE:f854547ec11886a8_2_2]]
Identify the function of component **"L"** in the given Transformer architecture.

Positional Encoding
Feed-Forward Network
Masked Multi-Head Self-Attention
Multi-Head Cross-Attention
Linear Output Projection
A published solution is not available for this question yet.
Question 4 MSQ · 2.0 marks
Answer the subquestions for the given transformer architecture for one (N = 1) **Encoder Decoder**
**Block**
[[IMAGE:f854547ec11886a8_2_2]]
Based on the architecture shown in the given image, identify all components that appear more
than once. Assume the architecture contains exactly one Encoder block and one Decoder block.

Positional Encoding
Feed-Forward Network
Masked Multi-Head Self-Attention
Multi-Head Cross-Attention
tokenizer
Add & Norm
A published solution is not available for this question yet.
Question 5 MSQ · 3.0 marks
[[IMAGE:f854547ec11886a8_4_3]]

The approach violates the autoregressive property
The approach ensures autoregressive property is not violated by ensuring that
the final attention weights of future tokens are 0
The final attention weights in the new approach will be greater than or equal
to that of the original approach
The attention weights for the final token of the sequence will always be the
same for both the approaches
A published solution is not available for this question yet.
Question 6 MSQ · 3.0 marks
Which of the following options correctly reflect the internal state updates performed during a Byte
Pair Encoding (BPE) training iteration?
After a pair is chosen for merging, it is permanently added to the active
vocabulary list.
The algorithm recalculates the log probability of the entire corpus before the
next iteration.
The independent frequencies of the two constituent tokens that were part of
the merge are reduced by the frequency of the merged token.
The original constituent tokens of the merged pair are removed from the
vocabulary to optimize dictionary size.
A published solution is not available for this question yet.
Question 7 MSQ · 3.0 marks
A baseline T5 model with denoising objective on a span of corrupted tokens is used for
unsupervised pre-training. Consider the phrase **"life is like a box of chocolates"**. If the word "life"
and the span "box of chocolates" are selected for corruption, which of the following statements
correctly describe the training process?
The input sequence replaces the corrupted spans with unique sentinal tokens
as: "<X> is like a <Y>".
The output sequence generated takes the form "<X> life <Y> box of chocolates
<Z>", with the final sentinal token to indicate sequence completion.
The loss is calculated exclusively over the generated sentinal tokens and the
missing text.
The decoder reconstructs the entire original sequence autoregressively to
compute the total loss over all positions.
A published solution is not available for this question yet.
Question 8 MCQ · 3.0 marks
During the fine-tuning of GPT-1 for Textual Entailment tasks, what is the purpose of the delimiter
token ($)?
To trigger the softmax function.
To separate premise and hypothesis.
To mark the end of the sentence.
To initialize the linear head.
A published solution is not available for this question yet.
Question 9 MCQ · 3.0 marks
During the fine-tuning of a decoder-only Transformer for sequence classification, why is the
hidden representation of the final input token used as the input to the classification layer?
The first token's representation relies exclusively on absolute positional
encodings and lacks semantic context for classification.
The final token's representation is computed independently of self-attention,
making its gradients more stable during the fine-tuning process.
Selecting the final token aligns with the autoregressive objective, eliminating
the need for teacher forcing during the classification forward pass.
Causal self-attention ensures that only the final token's representation
aggregates information from all preceding tokens in the sequence.
A published solution is not available for this question yet.
Question 10 MCQ · 3.0 marks
In the decoding phase of a Language model, temperature scaling with [[IMAGE:f854547ec11886a8_6_4]] is applied to the
output logits before performing top- [[IMAGE:f854547ec11886a8_6_5]] (nucleus) sampling. Which of the following options best
describes the impact of the temperature scaling on the subsequent sampling step?


Increases the required number of candidate tokens thus expanding the
nucleus.
Decreases the required number of candidate tokens thus shrinking the
nucleus.
Truncates the tail of the distribution causing the sampling to behave similar to
top- [[IMAGE:f854547ec11886a8_6_6]] sampling.

Size of the nucleus remains unaffected by the temperature scaling as it only
depends on order of the tokens.
A published solution is not available for this question yet.
Question 11 MCQ · 3.0 marks
In the WordPiece tokenization algorithm, how does the scoring formula prioritize token merges?
It evaluates the log-probability of the pair and selects the highest scoring
transition.
It merges the pair with the highest absolute frequency in the corpus.
It prioritizes pairs that occur together frequently but rarely appear
independently in the corpus.
None of these
A published solution is not available for this question yet.
Question 12 MCQ · 3.0 marks
Consider the T5 "text-to-text" paradigm. How does the model process fundamentally different
tasks such as Semantic Similarity (continuous regression) and Sentiment Analysis (discrete
classification) during the fine-tuning phase?
The model uses the encoder to output the regression score and decoder for
the classification tokens switching based on the input prompt.
The decoder part of the model is appended with separate prediction heads for
regression and classification which are updated based on Mean Squared Error and Cross-Entropy
respectively.
The model treats both tasks as sequence generation objectives with task-
specific prefix added in the input and all output is generated as strings.
None of these.
A published solution is not available for this question yet.
Question 13 MCQ · 3.0 marks
In the context of the BART (Bidirectional AutoRegressive Transformer) architecture, how does the
model handle corrupted input sequences compared to its output?
The decoder ignores the encoder output and performs standard causal
language modeling.
The encoder processes the original sequence and the decoder predicts the
corrupted tokens.
The encoder processes the corrupted sequence and the decoder predicts the
entire original sequence.
Both the encoder and decoder process the corrupted sequence to calculate
masked loss.
A published solution is not available for this question yet.
Question 14 MCQ · 3.0 marks
When comparing the C4 dataset to an 'Unfiltered-C4' dataset that is 8 times larger, what was the
observed effect on downstream task performance?
Performance remained identical, suggesting a saturation point in data utility.
Performance improved only for high-resource tasks like translation.
Performance improved significantly due to the increased diversity of data.
Performance degraded across all tasks despite the larger scale.
A published solution is not available for this question yet.
Question 15 MCQ · 3.0 marks
Which of the following describes the 'Adapter Layers' fine-tuning strategy?
Randomly freezing 50% of the attention heads in each layer during training.
Adding small dense-ReLU-dense blocks after FFN layers and updating only
those and Layer Norm parameters.
Updating all parameters in the model but using a much smaller learning rate.
Increasing the dimension [[IMAGE:f854547ec11886a8_8_7]] of the existing Feed-Forward networks.

A published solution is not available for this question yet.
Question 16 MCQ · 3.0 marks
Which of the following is a reported negative effect of including significantly duplicated content in
LLM training data?
It decreases the model's overall inference speed during deployment.
It causes the model to generate repeated sequences much more frequently.
It permanently reduces the maximum supported context window length.
It causes the model to lose the ability to process multilingual input.
A published solution is not available for this question yet.
Question 17 MCQ · 3.0 marks
If you have a fixed compute budget ( [[IMAGE:f854547ec11886a8_8_8]] ) and choose to train an extraordinarily large model
(massive [[IMAGE:f854547ec11886a8_8_9]] ) on a relatively small dataset (small [[IMAGE:f854547ec11886a8_8_10]] ), what is the likely outcome according to scaling
laws?



The model will be over-trained and memorize the dataset.
The model will be under-trained and less capable than a smaller model trained
on more data with the same compute budget.
The model will perfectly memorize the small dataset and generalize flawlessly.
The model's inference speed will increase dramatically.
A published solution is not available for this question yet.
Question 18 MSQ · 3.0 marks
Consider the following corpus consisting of 4 words (ignore spaces and punctuation).
**there is no spoon**
The vocabulary is constructed using the whole words in this corpus and the individual characters
in this corpus. The vocabulary should now have 13 tokens. Use this vocabulary for the given
subquestions.
Using the vocabulary constructed, two new continuous strings: "prison" and "rhino" are to be
segmented. Which of the following proposed segmentations are valid?
prison
p, r, is, o, n
p, r, is, on
r, h, i, n, o
r, h, i, no
A published solution is not available for this question yet.
Question 19 MCQ · 3.0 marks
Consider the following corpus consisting of 4 words (ignore spaces and punctuation).
**there is no spoon**
The vocabulary is constructed using the whole words in this corpus and the individual characters
in this corpus. The vocabulary should now have 13 tokens. Use this vocabulary for the given
subquestions.
Using a unigram language model, identify the segment with the highest probability from the
given
n, o, i, s, e, s
no, is, e, s
nois, e, s
no, i, s, e, s
A published solution is not available for this question yet.