da5013_2026T2_Q1_NA.pdf
Deep Learning Practice · Quiz 1 · May 2026
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 MCQ · 2.0 marks
A startup is building an LLM for a low resource agglutinative language where a single word can
contain information equivalent to an entire English sentence.
After training a Word Level tokenizer with a vocabulary of 30,000, they observe that nearly 18% of
tokens become [UNK].
Which action is most likely to improve the situation while keeping vocabulary growth manageable?
Use SentencePiece or BPE
Increase batch size
Increase the minimum token frequency threshold during tokenizer training
Disable normalization so every word form is stored separately
A published solution is not available for this question yet.
Question 3 MCQ · 2.0 marks
Consider the following code:
[[IMAGE:be50b0cf89de8deb_3_2]]
Assume that the variable [[IMAGE:be50b0cf89de8deb_3_3]] contains a piece of text. When tokenized without applying
truncation or padding, it produces 220 tokens.
Which of the following correctly describes the [[IMAGE:be50b0cf89de8deb_3_4]] returned in [[IMAGE:be50b0cf89de8deb_3_5]] ?




The output contains all 220 tokens because tokenization is completed before
truncation.
The output contains the first 128 tokens followed by additional padding
tokens.
The output contains exactly 128 tokens from the document.
The output length depends on the tokenizer model being used.
A published solution is not available for this question yet.
Question 4 MCQ · 2.0 marks
A team fine tunes a 13B parameter model using LoRA.
After training they discover that only 0.15% of parameters have been updated.
Which statement best explains this observation?
LoRA updates low rank matrices while freezing most pretrained weights
LoRA trains only the embedding layer
LoRA automatically quantizes the model
LoRA removes attention layers during training
A published solution is not available for this question yet.
Question 5 MCQ · 2.0 marks
A company wants to build a medical chatbot.
They possess:
• 8 million unlabeled medical documents
• 300 doctor reviewed question answer pairs
Which adaptation strategy would most likely provide the largest performance gain?
Train a new tokenizer only
Continue pretraining on the medical corpus and then instruction tune
Increase context length
Increase batch size
A published solution is not available for this question yet.
Question 6 MCQ · 2.0 marks
A 7 billion parameter model is loaded in FP32 precision.
Ignoring activations and optimizer states, approximately how much GPU memory is required for
storing model weights?
7 GB
14 GB
28 GB
56 GB
A published solution is not available for this question yet.
Question 7 MCQ · 2.0 marks
A tokenizer vocabulary is increased from 16K to 128K while keeping the training corpus
unchanged.
What tradeoff is most likely?
Shorter sequences and a larger embedding layer
Longer sequences and a smaller embedding layer
Shorter sequences and a smaller embedding layer
No significant effect on the model
A published solution is not available for this question yet.
Question 8 MCQ · 2.0 marks
Consider the following code:
[[IMAGE:be50b0cf89de8deb_5_6]]
Which method fills the blank to return **only** the token IDs as a plain Python list?

[[IMAGE:be50b0cf89de8deb_5_7]]

[[IMAGE:be50b0cf89de8deb_5_8]]

[[IMAGE:be50b0cf89de8deb_5_9]]

[[IMAGE:be50b0cf89de8deb_5_10]]

A published solution is not available for this question yet.
Question 9 MCQ · 3.0 marks
The following code is intended to implement causal masking for self attention, where each token
can attend only to itself and previous tokens.
[[IMAGE:be50b0cf89de8deb_6_11]]
What is the effect of running this code as written?
(**Note:** Assume necessary imports are already done and [[IMAGE:be50b0cf89de8deb_6_12]] is a valid tensor of attention
scores with shape [[IMAGE:be50b0cf89de8deb_6_13]] .)



No effect. This correctly implements causal masking.
It raises a [[IMAGE:be50b0cf89de8deb_6_14]] because [[IMAGE:be50b0cf89de8deb_6_15]] cannot be used with
[[IMAGE:be50b0cf89de8deb_6_16]] .



The attention weights in every row will sum to more than 1.
The current token is also masked, so each token can attend only to earlier
tokens. The first row becomes entirely [[IMAGE:be50b0cf89de8deb_6_17]] , which can lead to invalid ( [[IMAGE:be50b0cf89de8deb_6_18]] ) attention weights.


A published solution is not available for this question yet.
Question 10 MSQ · 3.0 marks
Consider the following training configuration:
[[IMAGE:be50b0cf89de8deb_6_19]]
Which statements are correct?

Effective batch size is larger than 2
Memory usage is lower than FP32 training
Optimizer updates occur after every batch
Gradient accumulation can simulate larger batches
FP16 automatically reduces optimizer state size by half
A published solution is not available for this question yet.
Question 11 MSQ · 3.0 marks
A student trains GPT 2 using Hugging Face [[IMAGE:be50b0cf89de8deb_7_20]] with:
[[IMAGE:be50b0cf89de8deb_7_21]]
The training dataset contains sequences of different lengths, but the student forgets to execute:
[[IMAGE:be50b0cf89de8deb_7_22]]
Which of the following are true?



The data collator will fail when it attempts to pad the batch.
Decoder only models never need padding during training.
Using [[IMAGE:be50b0cf89de8deb_7_23]] instead may cause padding to use a
real vocabulary token.

Reusing the EOS token as the padding token is a common practice for GPT 2.
GPT 2 already defines a [[IMAGE:be50b0cf89de8deb_7_24]] token by default.

A published solution is not available for this question yet.
Question 12 MSQ · 3.0 marks
A team experiences GPU out of memory errors while fine tuning a language model.
Which actions may help?
Mixed precision training
Quantization
Gradient checkpointing
LoRA
Increasing vocabulary size
A published solution is not available for this question yet.
Question 13 MSQ · 3.0 marks
A model is trained for sentiment classification using the following prompt format:
[[IMAGE:be50b0cf89de8deb_8_25]]
Compared with traditional classification heads, which statements are true?

The model learns to follow instructions
The output vocabulary remains open ended
Predictions are restricted to predefined class IDs
Training becomes a next token prediction task
The objective is identical to logistic regression
A published solution is not available for this question yet.
Question 14 MSQ · 2.0 marks
Which outputs can typically be produced by a Hugging Face tokenizer?
input_ids
attention_mask
token_type_ids
hidden_states
offset_mapping
A published solution is not available for this question yet.
Question 15 NAT · 4.0 marks
A GPT style model has the following configuration:
• Vocabulary size = 65,000
• Context length = 4096
• Embedding dimension = 1536
The model uses learned positional embeddings.
Calculate the total number of embedding parameters (token embeddings + positional
embeddings) in millions.
Round your answer to one decimal place.
A published solution is not available for this question yet.
Question 16 NAT · 4.0 marks
A transformer block has:
• d_model = 1536
Assume standard self attention consisting of:
• Query projection
• Key projection
• Value projection
• Output projection
Ignore biases.
Calculate the total number of parameters in millions.
Round your answer to one decimal place.
A published solution is not available for this question yet.
Question 17 NAT · 4.0 marks
A model contains 5 billion parameters.
Training uses AdamW in FP32 precision.
Assume:
• Parameters = 4 bytes
• Gradients = 4 bytes
• Momentum = 4 bytes
• Variance = 4 bytes
Ignoring activations and buffers, calculate the total memory requirement in GB.
Use:
1 GB = 10\(^{9}\) bytes
Round your answer to one decimal place.
A published solution is not available for this question yet.
Question 18 NAT · 4.0 marks
A batch of sentences is tokenized using a pretrained BERT tokenizer:
[[IMAGE:be50b0cf89de8deb_11_26]]
What is the value printed by the code?
(Assume default tokenizer behavior is used.)

A published solution is not available for this question yet.
Question 19 NAT · 3.0 marks
A training job uses:
• Batch size = 8
• Context length = 4096
• Gradient accumulation = 4
• Training steps = 10,000 optimizer updates
Calculate the total number of tokens processed after 10,000 optimizer updates.
Express your answer in millions.
Round to one decimal place.
A published solution is not available for this question yet.