da5013_2025T2_Q1_NA.pdf
Deep Learning Practice · Quiz 1 · May 2025
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 213 MCQ · 2.0 marks
A start-up is building a new language model for a low-resource language with many compound
words and complex morphology. They are debating tokenization strategies. Which of the following
approaches is most likely to offer the best balance between vocabulary size, handling OOV words,
and capturing morphological variants effectively for this scenario?
Character-level tokenization
Word-level tokenization with a fixed vocabulary of 50,000 common words.
Subword tokenization (e.g., BPE or SentencePiece) trained on the available
corpus.
Using only pre-defined special tokens and treating all other text as raw byte
sequences.
A published solution is not available for this question yet.
Question 214 MCQ · 2.0 marks
When fully fine-tuning a large pre-trained Transformer model (e.g., >1 Billion parameters), which
of the following contributes LEAST significantly to the GPU memory bottleneck compared to the
others?
Storing the model parameters themselves.
Storing the gradients for each parameter.
Storing the optimizer states (e.g., momentum and variance for Adam).
Storing the input batch data (token IDs).
A published solution is not available for this question yet.
Question 215 MCQ · 2.0 marks
A research team wants their pre-trained language model to generate more helpful and harmless
responses without extensive task-specific dataset collection. They have a collection of prompts and
human-preferred responses. Which of the following techniques directly aligns with this goal and
data?
Pre-training the model on a larger, more diverse text corpus.
Full fine-tuning on multiple downstream classification tasks.
Instruction Tuning using prompt-completion pairs or reformatting existing
datasets into an instructional format.
Implementing Parameter-Efficient Fine-Tuning (PEFT) techniques like LoRa.
A published solution is not available for this question yet.
Question 216 MCQ · 3.0 marks
[[IMAGE:2b90cb5d8dd5ff90_4_0]]

[[IMAGE:2b90cb5d8dd5ff90_4_1]]

[[IMAGE:2b90cb5d8dd5ff90_4_2]]

[[IMAGE:2b90cb5d8dd5ff90_4_3]]

[[IMAGE:2b90cb5d8dd5ff90_4_4]]

A published solution is not available for this question yet.
Question 217 MCQ · 3.0 marks
[[IMAGE:2b90cb5d8dd5ff90_5_5]]

[[IMAGE:2b90cb5d8dd5ff90_5_6]]

[[IMAGE:2b90cb5d8dd5ff90_5_7]]

[[IMAGE:2b90cb5d8dd5ff90_5_8]]

[[IMAGE:2b90cb5d8dd5ff90_5_9]]

A published solution is not available for this question yet.
Question 218 MSQ · 4.0 marks
Which of the following statements accurately describe common characteristics or goals of
subword tokenization algorithms like BPE, WordPiece, or SentencePiece? (Select ALL that apply)
They aim to significantly reduce the vocabulary size compared to character-
level tokenization.
They can handle out-of-vocabulary (OOV) words by breaking them into known
subword units.
They primarily rely on merging the least frequent character or subword pairs
to build the vocabulary.
They can represent common words as single tokens and rare words as
sequences of subword tokens.
SentencePiece is designed to be language-agnostic, not requiring pre-
segmentation based on spaces.
A published solution is not available for this question yet.
Question 219 MSQ · 4.0 marks
A team has a powerful pre-trained language model (e.g., a GPT-3 class model). They want to adapt
it for a new summarization task but have very limited labeled summarization data and limited
compute for full fine-tuning. Which of the following strategies could be viable and effective? (Select
ALL that apply)
Zero-shot prompting by providing the text and an instruction like "Summarize
this:".
Few-shot prompting (in-context learning) by providing a few examples of text
and their summaries in the prompt before the target text.
Full supervised fine-tuning of all model parameters on the small labeled
dataset.
Using a Parameter-Efficient Fine-Tuning (PEFT) method like LoRA on the small
labeled dataset.
Collecting a much larger unlabeled corpus related to the summarization
domain and continuing pre-training.
A published solution is not available for this question yet.
Question 220 MSQ · 4.0 marks
[[IMAGE:2b90cb5d8dd5ff90_7_10]]

[[IMAGE:2b90cb5d8dd5ff90_7_11]]

[[IMAGE:2b90cb5d8dd5ff90_7_12]]

[[IMAGE:2b90cb5d8dd5ff90_7_13]]

[[IMAGE:2b90cb5d8dd5ff90_7_14]]

[[IMAGE:2b90cb5d8dd5ff90_7_15]]

A published solution is not available for this question yet.
Question 221 NAT · 3.0 marks
[[IMAGE:2b90cb5d8dd5ff90_8_16]]

A published solution is not available for this question yet.
Question 222 NAT · 3.0 marks
[[IMAGE:2b90cb5d8dd5ff90_8_17]]

A published solution is not available for this question yet.
Question 223 NAT · 4.0 marks
[[IMAGE:2b90cb5d8dd5ff90_9_18]]

A published solution is not available for this question yet.
Question 224 MCQ · 3.0 marks
[[IMAGE:2b90cb5d8dd5ff90_10_19]]
Based on the above data, answer the given subquestions.
[[IMAGE:2b90cb5d8dd5ff90_10_20]]


[[IMAGE:2b90cb5d8dd5ff90_10_21]]

[[IMAGE:2b90cb5d8dd5ff90_10_22]]

[[IMAGE:2b90cb5d8dd5ff90_10_23]]

[[IMAGE:2b90cb5d8dd5ff90_10_24]]

A published solution is not available for this question yet.
Question 225 NAT · 4.0 marks
[[IMAGE:2b90cb5d8dd5ff90_10_19]]
Based on the above data, answer the given subquestions.
Based on the provided configuration, calculate the total number of parameters in the model’s
embedding layer (token embeddings) in millions. Enter your answer rounded to one decimal
place.

A published solution is not available for this question yet.
Question 226 MCQ · 3.0 marks
[[IMAGE:2b90cb5d8dd5ff90_10_19]]
Based on the above data, answer the given subquestions.
Based on the provided configuration, what is a primary characteristic of this language model’s
architecture and training paradigm?

It’s an encoder-decoder model designed for sequence-to-sequence tasks like
translation.
It’s an encoder-only model, likely using Masked Language Modeling for pre-
training.
It’s a decoder-only model, pre-trained using a Causal Language Modeling
objective.
It’s a small model primarily intended for edge devices due to its limited
context length.
A published solution is not available for this question yet.
Question 227 NAT · 4.0 marks
[[IMAGE:2b90cb5d8dd5ff90_10_19]]
Based on the above data, answer the given subquestions.
Considering the Adam optimizer stores 2 floating-point values per model parameter and
parameters are 32-bit floats (4 bytes), if the total number of trainable parameters in the model is
exactly 350 Million, how much GPU memory (in Gigabytes, GB) would be required just for the
optimizer states? (Assume 1 GB = 10\(^{9}\) bytes). Enter your answer rounded to one decimal place.

A published solution is not available for this question yet.
Question 228 MCQ · 2.0 marks
[[IMAGE:2b90cb5d8dd5ff90_10_19]]
Based on the above data, answer the given subquestions.
The configuration states the model uses Byte Pair Encoding (BPE). What is a key implication of this
choice for handling text from diverse sources during inference?

The model will be unable to process any words not seen during BPE
vocabulary training, leading to frequent errors.
All input words will be tokenized into individual characters, increasing
sequence length significantly.
Out-of-vocabulary words can be represented as sequences of known subword
units, allowing the model to process them.
BPE ensures that every language will have roughly the same number of tokens
for a text of similar semantic content.
**Statistical Computing**
**Section Id :** 64065391705
**Section Number :** 12
**Section type :** Online
**Mandatory or Optional :** Mandatory
**Number of Questions :** 10
**Number of Questions to be attempted :** 10
**Section Marks :** 30
**Display Number Panel :** Yes
**Section Negative Marks :** 0
**Group All Questions :** No
**Enable Mark as Answered Mark for Review and**
No
**Clear Response :**
**Section Maximum Duration :** 0
**Section Minimum Duration :** 0
**Section Time In :** Minutes
**Maximum Instruction Time :** 0
A published solution is not available for this question yet.