da5004_2026T1_Q2_NA.pdf
Large Language Models · Quiz 2 · Jan 2026
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 MCQ · 2.0 marks
In the context of Multi-Head Attention, if the model dimension [[IMAGE:a28d0ec52ebbc652_2_2]] and we employ
[[IMAGE:a28d0ec52ebbc652_2_3]] heads, what is the dimension of the concatenated output of the 8 heads before it is passed
through the final linear output projection [[IMAGE:a28d0ec52ebbc652_2_4]] ?



[[IMAGE:a28d0ec52ebbc652_2_5]]

[[IMAGE:a28d0ec52ebbc652_2_6]]

[[IMAGE:a28d0ec52ebbc652_2_7]]

[[IMAGE:a28d0ec52ebbc652_2_8]]

A published solution is not available for this question yet.
Question 3 MCQ · 2.0 marks
In the Masked Language Modeling (MLM) objective of BERT, 15% of tokens are chosen for
prediction. Of these chosen tokens, 80% are replaced with [[IMAGE:a28d0ec52ebbc652_2_9]] , 10% are replaced with a
random word, and 10% are left unchanged. What is the primary theoretical motivation for
**keeping 10% of the tokens unchanged**?

To reduce the computational cost of the softmax layer during training.
To prevent the network from overfitting to the [[IMAGE:a28d0ec52ebbc652_3_10]] token.

To mitigate the mismatch between pre-training (where [[IMAGE:a28d0ec52ebbc652_3_11]] appears) and
fine-tuning (where [[IMAGE:a28d0ec52ebbc652_3_12]] does not appear), ensuring the model creates meaningful
representations for non-masked words.


To act as a regularizer similar to Dropout.
A published solution is not available for this question yet.
Question 4 MCQ · 2.0 marks
Why is the standard GPT architecture (Decoder-only with causal masking) generally **unsuitable** for
the Masked Language Modeling (MLM) objective as implemented in BERT?
GPT models are too small to learn bidirectional contexts.
The causal mask in GPT prevents the model from attending to future tokens,
making it impossible to use right-side context to predict a masked token.
GPT does not have positional embeddings, which are required for MLM.
GPT uses ReLU activation, while BERT uses GELU, which is required for MLM.
A published solution is not available for this question yet.
Question 5 MCQ · 2.0 marks
In the T5 (Text-to-Text Transfer Transformer) framework, every NLP task is cast as a text
generation problem.
If you use T5 for a **Semantic Textual Similarity (STS-B)** task, where the goal is to predict a
similarity score (e.g., 3.8) between two sentences, how does the model output this score?
It outputs a single scalar value from a regression head on top of the encoder.
It generates the string "3.8" token-by-token using the decoder.
It outputs a class label corresponding to a bucketed score range (e.g., "High
Similarity").
T5 cannot be used for regression tasks like STS-B.
A published solution is not available for this question yet.
Question 6 MCQ · 2.0 marks
When using **Temperature Sampling** to mix datasets from different tasks during multi-task pre-
training, let [[IMAGE:a28d0ec52ebbc652_3_13]] be the probability of sampling a task [[IMAGE:a28d0ec52ebbc652_3_14]] . The formula involves raising the
proportion to the power of [[IMAGE:a28d0ec52ebbc652_4_15]] .
If we set the temperature [[IMAGE:a28d0ec52ebbc652_4_16]] (a very large value), what happens to the sampling distribution
across tasks?




The model trains almost exclusively on the largest dataset (highest resource
task).
The sampling distribution approaches a uniform distribution, where all tasks
(large and small) are sampled with nearly equal probability.
The model trains almost exclusively on the smallest dataset (lowest resource
task).
The sampling distribution remains proportional to the original dataset sizes.
A published solution is not available for this question yet.
Question 7 MCQ · 2.0 marks
When constructing training datasets for large language models (LLMs), which of the following best
describes the key factors that must be balanced to achieve strong and reliable performance?
Model depth, number of parameters, and learning rate
Scale, diversity, and quality of the training data
Vocabulary size, tokenization method, and batch size
Compute budget, optimizer choice, and hardware efficiency
A published solution is not available for this question yet.
Question 8 MSQ · 3.0 marks
Which of the following are components found within a standard Transformer **Encoder** layer?
Multi-Head Self-Attention mechanism
Position-wise Feed-Forward Networks
Cross-Attention mechanism (Encoder-Decoder attention)
Masked Multi-Head Self-Attention
A published solution is not available for this question yet.
Question 9 MSQ · 3.0 marks
Select all correct statements regarding the comparison between RNNs and Transformers.
In Transformers, the path length between any two positions in the sequence is
constant [[IMAGE:a28d0ec52ebbc652_5_17]] , whereas in RNNs it is [[IMAGE:a28d0ec52ebbc652_5_18]] (where [[IMAGE:a28d0ec52ebbc652_5_19]] is the sequence length).



Transformers process tokens strictly sequentially during training in the same
way as RNNs.
Transformers allow for significantly more parallelization during training
compared to RNNs.
Attention mechanisms in Transformers utilize Query, Key, and Value vectors
derived from input embeddings.
A published solution is not available for this question yet.
Question 10 MSQ · 3.0 marks
Select all correct findings from the T5 paper regarding **Unsupervised Pre-training Objectives**.
The specific corruption rate (e.g., 10%, 15%, 25%) had a minimal effect on
downstream performance.
Using a span length of around 3 tokens performed slightly better than
masking single tokens.
The "Deshuffling" objective (reordering shuffled sentences) significantly
outperformed the Span Corruption objective.
Replacing a corrupted span with a unique sentinel token worked better than
simply dropping the tokens from the input.
A published solution is not available for this question yet.
Question 11 MSQ · 3.0 marks
Consider the concept of **Zero-Shot Transfer** as popularized by GPT-2. Why might this be preferred
over Supervised Fine-Tuning?
It allows the model to handle tasks for which no labeled training data is
available.
It always achieves higher accuracy than a fine-tuned SOTA model.
It avoids the need to store a separate specialized model (checkpoint) for every
downstream task.
It mimics the human ability to perform tasks based on instructions without
needing thousands of examples.
A published solution is not available for this question yet.
Question 12 NAT · 2.0 marks
Let [[IMAGE:a28d0ec52ebbc652_6_20]] denote the activation of the [[IMAGE:a28d0ec52ebbc652_6_21]] neuron for the [[IMAGE:a28d0ec52ebbc652_6_22]] training sample.
In **Layer Normalization**, normalization is performed **across features for each individual**
**sample**.
Let [[IMAGE:a28d0ec52ebbc652_6_23]] be the number of neurons (features) in the hidden layer.
For a given sample, the layer normalization statistics are computed as:
[[IMAGE:a28d0ec52ebbc652_6_24]]
[[IMAGE:a28d0ec52ebbc652_6_25]]
[[IMAGE:a28d0ec52ebbc652_6_26]]
where [[IMAGE:a28d0ec52ebbc652_6_27]] is a small constant for numerical stability.
[[IMAGE:a28d0ec52ebbc652_6_28]]
Layer Normalization is applied **independently to each sample** in the given data above.
For **Sample 2**, after applying Layer Normalization. **What is the maximum value in the**
**normalized output vector** [[IMAGE:a28d0ec52ebbc652_6_29]] **?**
(ignore the learnable parameters [[IMAGE:a28d0ec52ebbc652_6_30]] and [[IMAGE:a28d0ec52ebbc652_6_31]] , and assume [[IMAGE:a28d0ec52ebbc652_6_32]] ):













A published solution is not available for this question yet.
Question 13 MCQ · 2.0 marks
Consider a GPT model used for **Causal Language Modelling**.
We feed the input sentence:
[[IMAGE:a28d0ec52ebbc652_7_33]]
Assume the context length is 7 (time steps t = 0 to t = 6).
The attention matrix computed in one attention layer is given below:
[[IMAGE:a28d0ec52ebbc652_7_34]]
Based on the above data, answer the given subquestions.
Is the given attention matrix valid for a causal language modelling task?


True
False
Insufficient information
A published solution is not available for this question yet.
Question 14 NAT · 2.0 marks
Consider a GPT model used for **Causal Language Modelling**.
We feed the input sentence:
[[IMAGE:a28d0ec52ebbc652_7_33]]
Assume the context length is 7 (time steps t = 0 to t = 6).
The attention matrix computed in one attention layer is given below:
[[IMAGE:a28d0ec52ebbc652_7_34]]
Based on the above data, answer the given subquestions.
At time step t = 5 (word: data), what is the attention weight assigned to the word science? (Provide
exact answer)


A published solution is not available for this question yet.
Question 15 MCQ · 2.0 marks
**Sentence Corpus:**
"the key to artificial intelligence has always been the representation"
**Pre-processing Instructions:**
• Convert all text to lowercase.
• Append the end-of-word symbol [[IMAGE:a28d0ec52ebbc652_8_35]] to every word (e.g., "the" becomes [[IMAGE:a28d0ec52ebbc652_8_36]] ).
• The **Initial Vocabulary** is defined as the set of all unique characters present in the corpus plus
the [[IMAGE:a28d0ec52ebbc652_8_37]] symbol.
Based on the above data, answer the given subquestions.
Which of the following character pairs occurs with the **highest frequency** across the entire
sentence before any BPE merges are performed?



[[IMAGE:a28d0ec52ebbc652_8_38]]

[[IMAGE:a28d0ec52ebbc652_8_39]]

[[IMAGE:a28d0ec52ebbc652_9_40]]

[[IMAGE:a28d0ec52ebbc652_9_41]]

A published solution is not available for this question yet.
Question 16 NAT · 2.0 marks
**Sentence Corpus:**
"the key to artificial intelligence has always been the representation"
**Pre-processing Instructions:**
• Convert all text to lowercase.
• Append the end-of-word symbol [[IMAGE:a28d0ec52ebbc652_8_35]] to every word (e.g., "the" becomes [[IMAGE:a28d0ec52ebbc652_8_36]] ).
• The **Initial Vocabulary** is defined as the set of all unique characters present in the corpus plus
the [[IMAGE:a28d0ec52ebbc652_8_37]] symbol.
Based on the above data, answer the given subquestions.
Calculate the **total vocabulary size** immediately after the first merge operation is completed.
**Note:** The vocabulary size includes all individual base characters/symbols plus the newly created
merge token.



A published solution is not available for this question yet.
Question 17 NAT · 3.0 marks
Consider a vocabulary [[IMAGE:a28d0ec52ebbc652_9_42]] .
At a specific time step [[IMAGE:a28d0ec52ebbc652_9_43]] , the model outputs the following probability distribution:
[[IMAGE:a28d0ec52ebbc652_9_44]]
If we use **Top-** [[IMAGE:a28d0ec52ebbc652_9_45]] **sampling with** [[IMAGE:a28d0ec52ebbc652_9_46]] , what is the probability of selecting token **B**?
(Calculate the re-normalized probability. Enter the value correct to 2 decimal places).





A published solution is not available for this question yet.
Question 18 NAT · 3.0 marks
Consider a small BERT-like model with the following configuration:
• Embedding dimension ( [[IMAGE:a28d0ec52ebbc652_10_47]] ) = 128
• Vocabulary size ( [[IMAGE:a28d0ec52ebbc652_10_48]] ) = 5000
• Maximum sequence length ( [[IMAGE:a28d0ec52ebbc652_10_49]] ) = 64
• Number of segment types = 2
Calculate the **total number of parameters** in the **Embedding Layer** (sum of Token embeddings,
Position embeddings, and Segment embeddings).
Enter the exact integer value.



A published solution is not available for this question yet.
Question 19 MCQ · 3.0 marks
What is a major disadvantage of character-level tokenization?
Character-level tokenization cannot represent punctuation marks or special
symbols
It has very large vocabulary size because each word is broken into multiple
characters
It produces much longer input sequences which increases computational cost.
It fails to capture morphological patterns such as prefixes and suffixes,
reducing the model's ability to understand word structure
A published solution is not available for this question yet.
Question 20 MCQ · 3.0 marks
When using the SentencePiece, how are subword units selected?
Based on a probabilistic model that maximizes the likelihood of the training
data.
By randomly selecting character n-grams until the vocabulary limit is reached.
By selecting only the top 50,000 most frequent words in the corpus.
By iteratively merging the most frequent pair of adjacent characters.
A published solution is not available for this question yet.
Question 21 MSQ · 2.0 marks
To obtain high-quality training text from raw web data for large language models, which of the
following mechanisms are typically included in the pre-processing pipeline?
Tokenizing text into subword units using Byte Pair Encoding (BPE)
Deduplicating content at line, paragraph, and document levels
Detecting and filtering toxic content such as hate speech and profanity
Identifying the language of web pages
Assessing the quality of content to remove low-value or spam text
Detecting and removing Personally Identifiable Information (PII)
Fine-tuning the model using reinforcement learning from human feedback
(RLHF)
A published solution is not available for this question yet.
Question 22 MSQ · 2.0 marks
Which of the following statements correctly explain the importance of deduplication during
preprocessing of large-scale datasets used for training Deep Learning or Large Language Models?
Deduplication reduces the risk of overfitting by preventing repeated samples
from dominating the gradient updates.
Deduplication guarantees that the trained model will achieve higher accuracy
on all downstream tasks.
Deduplication helps avoid data leakage between training and evaluation sets,
leading to more reliable performance metrics.
Deduplication eliminates the need for regularization techniques such as
dropout and weight decay.
A published solution is not available for this question yet.