da5004_2025T3_Q2_NA.pdf
Large Language Models · Quiz 2 · Sep 2025
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 105 NAT · 3.0 marks
Suppose for a 2-component GMM, at some iteration you have
[[IMAGE:f8bb8de26f2feee9_1_0]]
What is the responsibility [[IMAGE:f8bb8de26f2feee9_1_1]] ?


A published solution is not available for this question yet.
Question 107 NAT · 2.0 marks
[[IMAGE:f8bb8de26f2feee9_2_2]]
Based on the above data, answer the given subquestions.
[[IMAGE:f8bb8de26f2feee9_3_3]]


A published solution is not available for this question yet.
Question 108 NAT · 2.0 marks
[[IMAGE:f8bb8de26f2feee9_2_2]]
Based on the above data, answer the given subquestions.
[[IMAGE:f8bb8de26f2feee9_3_4]]


A published solution is not available for this question yet.
Question 109 MCQ · 3.0 marks
[[IMAGE:f8bb8de26f2feee9_2_2]]
Based on the above data, answer the given subquestions.
Which token will be generated in the first merge?

[[IMAGE:f8bb8de26f2feee9_3_5]]

[[IMAGE:f8bb8de26f2feee9_3_6]]

[[IMAGE:f8bb8de26f2feee9_3_7]]

[[IMAGE:f8bb8de26f2feee9_4_8]]

A published solution is not available for this question yet.
Question 110 NAT · 2.0 marks
A study is comparing three architectures, all trained with the same compute budget:
• **Model A (Encoder-Decoder):** A 12-layer encoder and a 12-layer decoder with separate
parameters. Its total size is 220 million parameters.
• **Model B (Decoder-Only):** A 24-layer decoder-only model. Its total size is also approximately
220 million parameters.
• **Model C (Shared Enc-Dec):** A 12-layer encoder and a 12-layer decoder where the parameters
are shared ( [[IMAGE:f8bb8de26f2feee9_4_9]] ).
Assuming that the 12-layer encoder and 12-layer decoder in Model A have an equal number of
parameters, what is the approximate total parameter count (in millions) of Model C?

A published solution is not available for this question yet.
Question 111 NAT · 2.0 marks
Consider a text corpus being processed using the Byte Pair Encoding (BPE) algorithm. The initial
vocabulary contains [[IMAGE:f8bb8de26f2feee9_4_10]] tokens (representing characters), and the corpus contains a total of [[IMAGE:f8bb8de26f2feee9_4_11]]
tokens.
During BPE training, the following merge operations are performed. Each merge combines all
occurrences of a particular token pair into a single new token, reducing the total token count
accordingly:
1. **Merge 1:** Pair ('t', 'h') appears 24 times → merged into new token 'th'
2. **Merge 2:** Pair ('i', 'n') appears 18 times → merged into new token 'in'
3. **Merge 3:** Pair ('th', 'e') appears 10 times → merged into new token 'the'
4. **Merge 4:** Pair ('i', 's') appears 8 times → merged into new token 'is'
What is the final total number of tokens in the corpus after all 4 merges are completed?


A published solution is not available for this question yet.
Question 112 NAT · 2.0 marks
If Byte-Pair Encoding (BPE) starts with [[IMAGE:f8bb8de26f2feee9_5_12]] unique characters and performs [[IMAGE:f8bb8de26f2feee9_5_13]] merge
operations, what will be the final vocabulary size?


A published solution is not available for this question yet.
Question 113 NAT · 2.0 marks
In a 4-token sequence, a Transformer is calculating the attention for the third Token. Its Query
vector ( [[IMAGE:f8bb8de26f2feee9_6_14]] ) must be compared against the Key vectors ( [[IMAGE:f8bb8de26f2feee9_6_15]] ) of all tokens. The Causal LM
mask is applied to the raw scores before the softmax function. The raw dot-product scores are:
• [[IMAGE:f8bb8de26f2feee9_6_16]]
• [[IMAGE:f8bb8de26f2feee9_6_17]]
• [[IMAGE:f8bb8de26f2feee9_6_18]]
• [[IMAGE:f8bb8de26f2feee9_6_19]]
After the Causal LM mask and the subsequent softmax function are applied, what will be the final
attention probability assigned to Token 4?






A published solution is not available for this question yet.
Question 114 MSQ · 3.0 marks
Which of the following is/are **primarily** multilingual datasets? Select all that apply.
Sangraha
DOLMA
BookCorpus
CommonCrawl
A published solution is not available for this question yet.
Question 115 MSQ · 2.0 marks
Which decoding method ensures the same response every time for a given prompt (assuming no
randomness in model weights)? Select all that apply.
Top-k with k = 3, temperature > 1
Top-k with k = 3, temperature = 1
Top-p with p = 0.2, Temperature > 1
Top-p with p = 0.2, Temperature = 1
Beam Search with beam-size = 5
Beam Search with beam-size = 1
A published solution is not available for this question yet.
Question 116 MCQ · 2.0 marks
A data preprocessing pipeline removed data from the Common Crawl snapshot in three stages:
50% after language identification, 24% of the remaining after quality filtering, and 12% of what
remained after deduplication. What percentage of the original Common Crawl snapshot remains
in the final dataset? (Choose the closest number)
21%
66%
33%
50%
A published solution is not available for this question yet.
Question 117 MCQ · 3.0 marks
A data preprocessing pipeline needs to apply the following steps on raw Common Crawl data:
1. Fuzzy deduplication (using MinHash)
2. Toxicity detection (using ML-based classifiers)
3. PII removal (using regex patterns)
4. Exact deduplication (using hash-based matching)
5. Quality filtering (using several rule-based heuristics)
6. URL Filtering
What is the most efficient sequence to apply these steps?
6 [[IMAGE:f8bb8de26f2feee9_8_20]] 4 [[IMAGE:f8bb8de26f2feee9_8_21]] 1 [[IMAGE:f8bb8de26f2feee9_8_22]] 5 [[IMAGE:f8bb8de26f2feee9_8_23]] 2 [[IMAGE:f8bb8de26f2feee9_8_24]] 3





1 [[IMAGE:f8bb8de26f2feee9_8_25]] 4 [[IMAGE:f8bb8de26f2feee9_8_26]] 2 [[IMAGE:f8bb8de26f2feee9_8_27]] 5 [[IMAGE:f8bb8de26f2feee9_8_28]] 3 [[IMAGE:f8bb8de26f2feee9_8_29]] 6





5 [[IMAGE:f8bb8de26f2feee9_8_30]] 6 [[IMAGE:f8bb8de26f2feee9_8_31]] 4 [[IMAGE:f8bb8de26f2feee9_8_32]] 1 [[IMAGE:f8bb8de26f2feee9_8_33]] 2 [[IMAGE:f8bb8de26f2feee9_8_34]] 3





4 [[IMAGE:f8bb8de26f2feee9_8_35]] 5 [[IMAGE:f8bb8de26f2feee9_8_36]] 6 [[IMAGE:f8bb8de26f2feee9_8_37]] 1 [[IMAGE:f8bb8de26f2feee9_8_38]] 3 [[IMAGE:f8bb8de26f2feee9_8_39]] 2





A published solution is not available for this question yet.
Question 118 MCQ · 4.0 marks
[[IMAGE:f8bb8de26f2feee9_8_40]]

A = BART, B = T5, C = BERT, D = GPT
A = T5, B = BERT, C = GPT, D = BART
A = BERT, B = GPT, C = BART, D = T5
A = BERT, B = BART, C = GPT, D = T5
A published solution is not available for this question yet.
Question 119 MCQ · 2.0 marks
Consider the scaling law:
[[IMAGE:f8bb8de26f2feee9_9_41]]
where [[IMAGE:f8bb8de26f2feee9_9_42]] = model parameters, [[IMAGE:f8bb8de26f2feee9_9_43]] = training tokens, [[IMAGE:f8bb8de26f2feee9_9_44]] = scale constants, [[IMAGE:f8bb8de26f2feee9_9_45]] = scaling
exponents.
A team trains an LLM and observes the following losses for their model:
1. Training Loss
2. Test Loss (same distribution as training)
3. Test Loss (different distribution from training)
Which loss does [[IMAGE:f8bb8de26f2feee9_9_46]] represent?






Both 2 and 3
Both 1 and 2
1 only
2 only
A published solution is not available for this question yet.
Question 120 MCQ · 3.0 marks
A company is building a sentiment analysis system for customer reviews. They have two
BART is expected to provide significantly better accuracy and is worth the
extra computational cost since it's pre-trained on more diverse corruption tasks.
BART is better because its decoder can verify the encoder's sentiment
understanding, providing a built-in validation mechanism that improves robustness.
They should use BART's encoder only and discard the decoder, which will give
them the same efficiency as BERT with better pre-training.
BERT is more suitable because sentiment classification only requires
understanding (encoding), not generation (decoding), making it faster and more memory-efficient
without sacrificing accuracy.
A published solution is not available for this question yet.
Question 121 MCQ · 2.0 marks
Which of the following statements best describes the fundamental difference between the
unsupervised pre-training objective used in T5 (span corruption) and its supervised fine-tuning
objectives such as translation or summarization?
The pre-training objective is a denoising task, while the fine-tuning objective is
a Causal Language Model (CLM) task.
There is no fundamental difference; both are treated as "text-to-text" tasks,
where the model is trained to generate a target sequence given an input sequence.
The pre-training objective trains the model to fill in [MASK] tokens, while the
fine-tuning objective trains the model to generate text from a [CLS] token.
The pre-training objective trains the entire encoder-decoder, while the fine-
tuning objective freezes the encoder and only trains the decoder.
A published solution is not available for this question yet.
Question 122 MCQ · 2.0 marks
[[IMAGE:f8bb8de26f2feee9_11_47]]

[[IMAGE:f8bb8de26f2feee9_11_48]]

[[IMAGE:f8bb8de26f2feee9_11_49]]

[[IMAGE:f8bb8de26f2feee9_11_50]]

[[IMAGE:f8bb8de26f2feee9_11_51]]

A published solution is not available for this question yet.
Question 123 MCQ · 2.0 marks
A Transformer model is designed to handle both understanding and generation tasks. When given
the input sequence [[IMAGE:f8bb8de26f2feee9_11_52]] , its attention mask allows the following:
• When processing [[IMAGE:f8bb8de26f2feee9_11_53]] or [[IMAGE:f8bb8de26f2feee9_11_54]] , the model can see both [[IMAGE:f8bb8de26f2feee9_11_55]] and [[IMAGE:f8bb8de26f2feee9_11_56]] .
• When processing [[IMAGE:f8bb8de26f2feee9_11_57]] , the model can see [[IMAGE:f8bb8de26f2feee9_11_58]] and [[IMAGE:f8bb8de26f2feee9_11_59]] but not [[IMAGE:f8bb8de26f2feee9_11_60]] .
This attention scheme is the defining characteristic of which training objective?









Causal Language Modeling (CLM)
Masked Language Modeling (MLM)
Denoising Objective (like BART)
Prefix Language Modeling (Prefix-LM)
A published solution is not available for this question yet.
Question 124 MCQ · 2.0 marks
BART is an encoder-decoder model. During its denoising pre-training, what is fed as input to the
decoder?
The corrupted text from the encoder.
The original, uncorrupted text (shifted right by one token).
A sequence of [MASK] tokens, one for each original word.
Nothing; the decoder starts with only the < s > token.
A published solution is not available for this question yet.
Question 125 MCQ · 2.0 marks
[[IMAGE:f8bb8de26f2feee9_12_61]]

[[IMAGE:f8bb8de26f2feee9_12_62]]

[[IMAGE:f8bb8de26f2feee9_12_63]]

[[IMAGE:f8bb8de26f2feee9_12_64]]

[[IMAGE:f8bb8de26f2feee9_12_65]]

A published solution is not available for this question yet.
Question 126 MCQ · 2.0 marks
How does SentencePiece handle whitespace characters (like spaces) during tokenization?
It only keeps whitespace that appears next to punctuation marks.
It discards all whitespace as a pre-processing step.
It replaces all whitespace with a special [SPACE] token.
It treats whitespace as part of the text stream and encodes it, often using a
symbol like [[IMAGE:f8bb8de26f2feee9_13_66]] (underscore).

A published solution is not available for this question yet.
Question 127 MCQ · 2.0 marks
For a layer with [[IMAGE:f8bb8de26f2feee9_13_67]] features, how many learnable parameters do BatchNorm and LayerNorm
have?

BatchNorm: 512, LayerNorm: 512
BatchNorm: 1024, LayerNorm: 512
BatchNorm: 512, LayerNorm: 1024
BatchNorm: 1024, LayerNorm: 1024
A published solution is not available for this question yet.
Question 128 MCQ · 2.0 marks
In a standard Encoder-Decoder Transformer, where do the Query [[IMAGE:f8bb8de26f2feee9_13_68]] , Key [[IMAGE:f8bb8de26f2feee9_13_69]] , and Value [[IMAGE:f8bb8de26f2feee9_13_70]]
inputs for the Decoder's Cross-Attention sub-layer come from?



[[IMAGE:f8bb8de26f2feee9_13_71]] comes from the decoder's previous sub-layer; [[IMAGE:f8bb8de26f2feee9_13_72]] and [[IMAGE:f8bb8de26f2feee9_13_73]] come from the final
hidden states of the encoder.



[[IMAGE:f8bb8de26f2feee9_13_74]] comes from the final hidden states of the encoder; [[IMAGE:f8bb8de26f2feee9_13_75]] and [[IMAGE:f8bb8de26f2feee9_13_76]] come from the
decoder's previous sub-layer.



A published solution is not available for this question yet.