da5013_2026T1_Q1_NA.pdf
Deep Learning Practice · Quiz 1 · Jan 2026
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 MCQ · 2.0 marks
Consider subword tokenization methods such as Byte Pair Encoding (BPE) and WordPiece as
employed in transformer-based language modeling pipelines. Which of the following statements
correctly describes their fundamental properties?
They always split words at true morpheme boundaries (the smallest units of
meaning, e.g., “un-” + “break” + “able”).
They learn merge rules from corpus statistics (frequency/likelihood), not token
semantics
They eliminate (Out Of Vocabulary) OOV by ensuring every word is a single
token
They generalize perfectly to out-of-domain text without increasing sequence
length.
A published solution is not available for this question yet.
Question 3 MCQ · 2.0 marks
Consider two subword tokenizers trained on the same text corpus. Tokenizer A uses a vocabulary
of 8,000 tokens, whereas Tokenizer B uses a vocabulary of 64,000 tokens. Which of the following
statements most accurately characterizes the implications of these vocabulary sizes for training
efficiency in transformer-based language models?
Larger vocabulary always reduces sequence length and total compute.
Smaller vocabulary always improves semantic alignment.
Larger vocabulary reduces sequence length but increases embedding
parameters.
Vocabulary size has no effect once model is pretrained.
A published solution is not available for this question yet.
Question 4 MCQ · 2.0 marks
A subword tokenizer decomposes rare chemical entity names into a large number of fragments.
During downstream fine-tuning, the model exhibits degraded performance on a chemical named
entity recognition (NER) task. Which of the following represents the most principled corrective
action?
Increase the dropout rate during fine-tuning.
Retrain the tokenizer using an in-domain (chemical) corpus.
Reduce the batch size during training.
Apply label smoothing to the loss function.
A published solution is not available for this question yet.
Question 5 MCQ · 2.0 marks
[[IMAGE:5098e0c90adb7cd8_4_2]]

The output sequence will always contain exactly 12 non-padding tokens.
Special tokens may consume a portion of the 12-token length budget.
Word boundaries are preserved, since truncation occurs at whitespace
boundaries.
Padding will never be applied when truncation is enabled.
A published solution is not available for this question yet.
Question 6 MCQ · 2.0 marks
[[IMAGE:5098e0c90adb7cd8_5_3]]

This is supervised fine-tuning with noisy labels.
This is instruction tuning because labels are present.
This performs continual pretraining with MLM objective.
This trains a classifier head implicitly.
A published solution is not available for this question yet.
Question 7 MCQ · 2.0 marks
In the standard Transformer architecture (e.g., GPT or BERT), which component exhibits a
computational complexity that scales quadratically ( [[IMAGE:5098e0c90adb7cd8_5_4]] ) with respect to the input sequence
length [[IMAGE:5098e0c90adb7cd8_5_5]] ?


The initial token and positional embedding lookup.
The element-wise addition in the Residual (Skip) connections.
The computation of the attention score matrix (
[[IMAGE:5098e0c90adb7cd8_6_6]] ).

The linear projections in the Position-wise Feed-Forward Network (FFN).
A published solution is not available for this question yet.
Question 8 MCQ · 2.0 marks
A model with 4B parameters (fp32) is being trained using the AdamW optimizer on a single 32GB
GPU. Why is full fine-tuning impossible in this configuration?
The 4B parameters alone require 32GB of VRAM, leaving no room for the OS
or CUDA kernels.
The optimizer states (32GB) plus the model weights (16GB) and gradients
(16GB) exceed the 32GB VRAM limit.
The tokenizer is unable to map a context length of 2048 to a 32-bit integer
space.
The fp32 precision requires 8 bytes per parameter, meaning the 4B model
needs 64GB just to load.
A published solution is not available for this question yet.
Question 9 MCQ · 2.0 marks
During the full fine-tuning of a Large Language Model, you aim to regularize the training process
by penalizing the growth of weight magnitudes. This ensures the model does not overfit the fine-
tuning data by making excessively large updates to the pretrained parameters. Which parameter
is specifically designed to apply this penalty?
Gradient Accumulation.
Batch Size.
Checkpointing.
Weight Decay.
A published solution is not available for this question yet.
Question 10 MCQ · 2.0 marks
Why does full fine-tuning of large language models often require more memory than inference
using the same model?
Inference uses larger batch sizes.
Inference stores optimizer states.
Training requires storing gradients and optimizer states.
Training uses longer input sequences.
A published solution is not available for this question yet.
Question 11 MCQ · 2.0 marks
Which technique is most effective at reducing GPU memory usage during training without
modifying the model architecture?
Gradient clipping.
Mixed-precision training.
Increasing learning rate warmup.
Saving checkpoints less frequently.
A published solution is not available for this question yet.
Question 12 MSQ · 3.0 marks
Consider the output of a Hugging Face tokenizer when applied to a batch of sequences with
padding= "max_length" and return_tensors= "pt". Which of the following statements are correct?
(Select ALL that apply)
The input_ids tensor contains the numerical indices mapped from the
vocabulary.
The attention_mask contains 0s for padding tokens and 1s for real tokens to
prevent the model from attending to padding.
The token_type_ids (if present) are used primarily to distinguish between
uppercase and lowercase letters.
In models like BERT, the input_ids will typically begin with a special token index
(e.g., [CLS]).
A published solution is not available for this question yet.
Question 13 MSQ · 3.0 marks
[[IMAGE:5098e0c90adb7cd8_8_7]]

Effective batch size is larger than 1.
Reduced memory usage compared to FP32 training.
Gradient clipping limits large parameter updates.
Gradient accumulation reduces the number of optimizer states.
A published solution is not available for this question yet.
Question 14 MSQ · 3.0 marks
Which of the following statements correctly distinguish different Transformer architectures and
their typical training objectives? (Select ALL that apply)
Encoder-only models are commonly trained with Masked Language Modeling
objectives.
Decoder-only models rely on causal masking and autoregressive inference.
Encoder-decoder models are well suited for sequence-to-sequence tasks.
All Transformer architectures use identical attention masks.
A published solution is not available for this question yet.
Question 15 MSQ · 3.0 marks
[[IMAGE:5098e0c90adb7cd8_9_8]]

Task Conditioning: The model learns to follow natural language formatting
(Instruction/Input/Answer) rather than just mapping a sequence to a single integer ID.
Open Vocabulary: The model's prediction head remains the size of its full
vocabulary (e.g., 50k+ tokens), allowing it to generate any string as an answer instead of being
restricted to fixed logits for [[IMAGE:5098e0c90adb7cd8_9_9]] classes.

Guaranteed Zero-Shot: This training format ensures the model will generalize
perfectly to any unseen task prompt without further data.
Objective Shift: The training objective moves from minimizing cross-entropy
loss over a discrete class index to minimizing next-token prediction loss over the sequence tokens.
A published solution is not available for this question yet.
Question 16 MSQ · 3.0 marks
A model is trained using the Causal Language Modeling (CLM) objective with a standard cross-
entropy loss. Which of the following statements correctly describe the training dynamics? (Select
ALL that apply)
The loss at position [[IMAGE:5098e0c90adb7cd8_9_10]] depends only on tokens [[IMAGE:5098e0c90adb7cd8_9_11]] .


The attention mask is strictly upper triangular.
Future tokens contribute gradients to earlier positions.
The joint probability of the sequence is factorized autoregressively.
A published solution is not available for this question yet.
Question 17 MSQ · 3.0 marks
When configuring TrainingArguments for a Transformer model, we include a learning rate
warmup phase (e.g., warmup_steps=500). Which of the following statements correctly describe the
purpose and behavior of this strategy? (Select ALL that apply)
It involves linearly increasing the learning rate from 0 (or a small value) to the
target maximum during the initial phase.
It helps prevent "divergence'' or instability caused by large gradients when the
model weights are far from their optimal state.
It significantly reduces the VRAM (memory) required to store optimizer states.
After the warmup phase, the learning rate typically follows a decay schedule
(like linear or cosine) to ensure convergence.
A published solution is not available for this question yet.
Question 18 NAT · 4.0 marks
A transformer-based language model has 1.2 billion parameters and is fine-tuned using AdamW.
Assume parameters, gradients, and optimizer states (first and second moments) are all stored in
32-bit precision (4 bytes each). Ignoring activations and buffers, calculate the total GPU memory
required (in GB) to store parameters, gradients, and optimizer states. (Assume 1 GB = [[IMAGE:5098e0c90adb7cd8_10_12]] bytes.)

A published solution is not available for this question yet.
Question 19 NAT · 4.0 marks
A GPT-style transformer block has an embedding dimension of 768 and uses 12 attention heads.
Consider a specific processing task with a very short sequence of only 3 tokens ( [[IMAGE:5098e0c90adb7cd8_11_13]] ).Compute
the total number of weight parameters (ignore biases) in the self-attention module, specifically for
the Query, Key, Value, and Output projection matrices. Report your answer in millions, rounded to
one decimal place.

A published solution is not available for this question yet.
Question 20 NAT · 4.0 marks
You are training a decoder-only language model with a vocabulary size of 32,000, maximum
context length of 2048, and embedding dimension of 1024. The model uses learned positional
embeddings. Calculate the total number of parameters in the embedding layer (token + positional
embeddings). Report your answer in millions, rounded to one decimal place.
A published solution is not available for this question yet.