da5004_2025T3_ET_FN.pdf
Large Language Models · End Term · Sep 2025 FN
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 MCQ · 1.0 marks
What happens when you use Batch Normalization with a batch size of 1?
It works perfectly fine
The statistics become meaningless (mean=0, std=0)
It automatically switches to Layer Normalization
It uses the running statistics from training
A published solution is not available for this question yet.
Question 3 MCQ · 2.0 marks
A language model outputs the following logits for the next token:
{cat: 3.2, dog: 2.9, bird: 1.1, fish: 0.5, snake: -0.4}
Before sampling, the decoding pipeline performs:
• Temperature scaling with T = 2.0
• Top-K filtering with K = 3
After applying both steps, which tokens remain eligible for sampling?
Only ''cat''
'cat'' or ''dog''
'cat'', ''dog'' or ''bird''
'dog'', ''bird'' or ''fish''
A published solution is not available for this question yet.
Question 4 MCQ · 2.0 marks
If a model's next-token probabilities are [[IMAGE:b55b6f7161c9d349_2_2]] , [[IMAGE:b55b6f7161c9d349_2_3]] , [[IMAGE:b55b6f7161c9d349_2_4]] , [[IMAGE:b55b6f7161c9d349_2_5]] ,
and [[IMAGE:b55b6f7161c9d349_2_6]] , and Top-P is set to [[IMAGE:b55b6f7161c9d349_2_7]] . Which set of tokens will be included in the nucleus
for sampling?






Tokens A, B, C, D, E
Tokens A, B
Tokens A, B, C
Tokens A, B, C, D
Tokens A, B, D
A published solution is not available for this question yet.
Question 5 MCQ · 2.0 marks
In a WordPiece vocabulary-building step, the corpus contains the following tokenized sequences:
• [[IMAGE:b55b6f7161c9d349_3_8]]
• [[IMAGE:b55b6f7161c9d349_3_9]]
The current vocabulary includes the tokens:
[[IMAGE:b55b6f7161c9d349_3_10]] , [[IMAGE:b55b6f7161c9d349_3_11]] , [[IMAGE:b55b6f7161c9d349_3_12]] , [[IMAGE:b55b6f7161c9d349_3_13]] , [[IMAGE:b55b6f7161c9d349_3_14]]
Two candidate merges are being evaluated:
• Merge A: [[IMAGE:b55b6f7161c9d349_3_15]] [[IMAGE:b55b6f7161c9d349_3_16]] [[IMAGE:b55b6f7161c9d349_3_17]]
• Merge B: [[IMAGE:b55b6f7161c9d349_3_18]] [[IMAGE:b55b6f7161c9d349_3_19]] [[IMAGE:b55b6f7161c9d349_3_20]]
Based purely on WordPiece's likelihood-based merge selection, which merge is more likely to be
chosen?













Merge [[IMAGE:b55b6f7161c9d349_3_21]]

Merge [[IMAGE:b55b6f7161c9d349_3_22]]

Insufficient information to determine
A published solution is not available for this question yet.
Question 6 MCQ · 2.0 marks
A research team wants to classify scientific abstracts into multiple topics simultaneously (e.g.,
''ML'', ''biology'', ''statistics''), where each abstract may belong to more than one topic. They fine-
tune BERT for this task. Which modification is MOST appropriate?
Replace the [CLS] vector with an average of the top-4 attention heads
Feed the [CLS] embedding into a linear layer with a sigmoid activation per
label
Use token embeddings individually and classify each token
Use the [SEP] token embedding for multi-label prediction
A published solution is not available for this question yet.
Question 7 MCQ · 1.0 marks
A company wants to automatically correct noisy OCR text extracted from scanned documents. The
text contains spelling mistakes, missing words, and scrambled phrases. Which model should they
fine-tune?
BERT, because it masks tokens and predicts them independently
BART, because it is trained with text corruption and autoregressive
reconstruction
GPT, because it is optimal for bidirectional correction
BERT, because [CLS] captures global structure
A published solution is not available for this question yet.
Question 8 MCQ · 1.0 marks
Computing the full self-attention matrix has a time complexity of
[[IMAGE:b55b6f7161c9d349_4_23]]
where [[IMAGE:b55b6f7161c9d349_4_24]] is the sequence length and [[IMAGE:b55b6f7161c9d349_4_25]] is the embedding dimension.
Which step in the self-attention mechanism is primarily responsible for this quadratic scaling?



Computing the softmax for each row
Computing the matrix product [[IMAGE:b55b6f7161c9d349_4_26]]

Multiplying the attention matrix [[IMAGE:b55b6f7161c9d349_4_27]] with the value matrix [[IMAGE:b55b6f7161c9d349_4_28]]


Performing layer normalization
A published solution is not available for this question yet.
Question 9 MCQ · 1.0 marks
During autoregressive inference, key-value (KV) caching is used to avoid recomputing keys and
values for previously generated tokens. If the KV cache is allowed to grow without any limit, which
of the following failure modes can occur?
GPU memory exhaustion
Latency becoming quadratic in the output length
Inability to perform parallel token generation
Beam search collapsing to a single token
A published solution is not available for this question yet.
Question 10 MCQ · 2.0 marks
Consider a sequence of length [[IMAGE:b55b6f7161c9d349_5_29]] processed by a local [[IMAGE:b55b6f7161c9d349_5_30]] mechanism with block
size [[IMAGE:b55b6f7161c9d349_5_31]] , where [[IMAGE:b55b6f7161c9d349_5_32]] is a multiple of [[IMAGE:b55b6f7161c9d349_5_33]] . The sequence is partitioned into [[IMAGE:b55b6f7161c9d349_5_34]] non-overlapping
contiguous blocks, and self-attention is computed [[IMAGE:b55b6f7161c9d349_5_35]] (i.e., attention is
restricted to the main diagonal blocks of the [[IMAGE:b55b6f7161c9d349_5_36]] attention matrix).
Ignoring the cost of linear projections for [[IMAGE:b55b6f7161c9d349_5_37]] , what is the time complexity of the self-
attention operation for one layer in terms of [[IMAGE:b55b6f7161c9d349_5_38]] , [[IMAGE:b55b6f7161c9d349_5_39]] , and the head dimension [[IMAGE:b55b6f7161c9d349_5_40]] ?












[[IMAGE:b55b6f7161c9d349_5_41]]

[[IMAGE:b55b6f7161c9d349_5_42]]

[[IMAGE:b55b6f7161c9d349_5_43]]

[[IMAGE:b55b6f7161c9d349_5_44]]

A published solution is not available for this question yet.
Question 11 MCQ · 2.0 marks
A 2D vector [[IMAGE:b55b6f7161c9d349_5_45]] is rotated by [[IMAGE:b55b6f7161c9d349_5_46]] using RoPE. What is the resulting vector?


[[IMAGE:b55b6f7161c9d349_5_47]]

[[IMAGE:b55b6f7161c9d349_5_48]]

[[IMAGE:b55b6f7161c9d349_5_49]]

[[IMAGE:b55b6f7161c9d349_5_50]]

A published solution is not available for this question yet.
Question 12 MCQ · 2.0 marks
If a model uses Absolute Positional Embeddings (APE), increasing the sequence length from 100 to
1000 requires:
No new parameters
900 additional positional embedding vectors
A new rotation matrix
A distance-based bias matrix
A published solution is not available for this question yet.
Question 13 MCQ · 3.0 marks
In a sequence of length [[IMAGE:b55b6f7161c9d349_6_51]] , how many different attention pairs produce a distance of [[IMAGE:b55b6f7161c9d349_6_52]] ?


3
5
6
7
A published solution is not available for this question yet.
Question 14 MCQ · 3.0 marks
For a sequence of length [[IMAGE:b55b6f7161c9d349_6_53]] , how many learned positional vectors are required by the
following methods?
• Absolute Positional Encoding (APE)
• Relative Positional Encoding (RPE)
• ALiBi
• NoPE (No Positional Encoding)

APE = 10, RPE = 19, ALiBi = 0, NoPE = 0
APE = 10, RPE = 10, ALiBi = 19, NoPE = 0
APE = 20, RPE = 19, ALiBi = 1, NoPE = 1
APE = 10, RPE = 2T = 20, ALiBi = 10, NoPE = 0
A published solution is not available for this question yet.
Question 15 MCQ · 2.0 marks
Which of the following methods uses/use the concept: "The farther apart two tokens are, the less
they should attend to each other''?
RoPE
APE
ALiBi
RPE
A published solution is not available for this question yet.
Question 16 MCQ · 2.0 marks
You trained a model with max sequence length 512. At inference, you need to process sequence
length 4096 without retraining.
Which positional encoding will perform best out-of-the-box?
APE
RPE
RoPE
ALiBi
A published solution is not available for this question yet.
Question 17 MCQ · 2.0 marks
A model using ALiBi is extended from context length 1K [[IMAGE:b55b6f7161c9d349_7_54]] 8K.
How many new ALiBi parameters are added?

7000
8000
Depends on the number of distances
0
A published solution is not available for this question yet.
Question 18 NAT · 3.0 marks
A mini-batch has [[IMAGE:b55b6f7161c9d349_8_55]] sentences, each padded to a length [[IMAGE:b55b6f7161c9d349_8_56]] . Self-attention (single head)
computes one score matrix per sentence. How many scalar scores are computed in total for a
single head across the batch?


A published solution is not available for this question yet.
Question 19 NAT · 2.0 marks
Suppose you are working on prefix language modeling. The sequence length is [[IMAGE:b55b6f7161c9d349_8_57]] and the first
two tokens represent the task-specific prefix. How many non-infinity elements are there in the
mask for computing attention scores?

A published solution is not available for this question yet.
Question 20 NAT · 2.0 marks
A model uses a hybrid attention pattern in which each token attends to:
• A local window of size [[IMAGE:b55b6f7161c9d349_9_58]]
• A set of [[IMAGE:b55b6f7161c9d349_9_59]] random tokens,
• A set of [[IMAGE:b55b6f7161c9d349_9_60]] global tokens.
Compute the total number of attention computations per token (integer expression).



A published solution is not available for this question yet.
Question 21 NAT · 2.0 marks
During deployment of a transformer-based Large Language Model (LLM), you want to estimate
the memory usage of the KV cache for efficient batching. Compute the KV-cache memory required
per token (in Bytes) for the following model configuration:
• Number of Transformer Blocks ( [[IMAGE:b55b6f7161c9d349_9_61]] ): 24
• Attention Heads per Layer ( [[IMAGE:b55b6f7161c9d349_9_62]] ): 12
• Head Dimension ( [[IMAGE:b55b6f7161c9d349_9_63]] ): 96
• Precision used for KV tensors: 2 Bytes (FP16)



A published solution is not available for this question yet.
Question 22 NAT · 3.0 marks
You are implementing a custom Block Sparse Attention kernel. Given a sequence length
[[IMAGE:b55b6f7161c9d349_10_64]] and block size [[IMAGE:b55b6f7161c9d349_10_65]] , you decide to compute the main diagonal blocks plus one
random off-diagonal block for each block-row. How many total [[IMAGE:b55b6f7161c9d349_10_66]] block computations are
performed?



A published solution is not available for this question yet.
Question 23 MSQ · 3.0 marks
A GPT-style causal language model is trained using the next-token prediction objective:
[[IMAGE:b55b6f7161c9d349_10_67]]
During inference, the model must generate tokens autoregressively from left to right using only
past context.
Consider the following statements about GPT-style causal models.

GPT cannot condition on future tokens during training because the causal
mask zeros out all attention to positions [[IMAGE:b55b6f7161c9d349_10_68]]

If two prefixes [[IMAGE:b55b6f7161c9d349_10_69]] and [[IMAGE:b55b6f7161c9d349_10_70]] have identical embeddings at all positions, GPT
must assign identical next-token distributions for both contexts


Removing positional encodings would force GPT to treat all permutations of
the same set of tokens as equivalent contexts
Causal masking ensures that the computational cost of training scales linearly
with sequence length
A published solution is not available for this question yet.
Question 24 MSQ · 3.0 marks
Under which of the following decoding settings can the model produce different outputs across
multiple runs on the same prompt?
Top-K sampling with [[IMAGE:b55b6f7161c9d349_11_71]] and temperature [[IMAGE:b55b6f7161c9d349_11_72]]


Top-K sampling with [[IMAGE:b55b6f7161c9d349_11_73]] and temperature [[IMAGE:b55b6f7161c9d349_11_74]]


Nucleus (Top-p) sampling with [[IMAGE:b55b6f7161c9d349_11_75]] and temperature [[IMAGE:b55b6f7161c9d349_11_76]]


Greedy decoding after applying a temperature [[IMAGE:b55b6f7161c9d349_11_77]] to the logits

Beam search with beam width [[IMAGE:b55b6f7161c9d349_11_78]]

A published solution is not available for this question yet.
Question 25 MSQ · 2.0 marks
Consider a BART encoder with sequence length [[IMAGE:b55b6f7161c9d349_11_79]] and hidden size [[IMAGE:b55b6f7161c9d349_11_80]] (fixed).
Now suppose the sequence length is doubled from [[IMAGE:b55b6f7161c9d349_11_81]] to [[IMAGE:b55b6f7161c9d349_11_82]] , keeping all other parameters
unchanged.
Which of the following statements are true?




Self-attention FLOPs increase by a factor of [[IMAGE:b55b6f7161c9d349_11_83]] .

Self-attention FLOPs increase by a factor of [[IMAGE:b55b6f7161c9d349_11_84]] .

FFN FLOPs increase by a factor of [[IMAGE:b55b6f7161c9d349_11_85]] .

FFN FLOPs increase by a factor of [[IMAGE:b55b6f7161c9d349_11_86]] .

A published solution is not available for this question yet.