da5004_2025T2_Q1_NA.pdf
Large Language Models · Quiz 1 · May 2025
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 133 MCQ · 2.0 marks
We continue the notation from the previous question. Suppose the randomly assigned weights
lead to the good event, that is, there is an unique perfect matching in G with minimum weight, and
say this minimum weight is r. What can you say about the determinant of Z in this case?
[[IMAGE:a8f64888d57debe8_1_3]]

[[IMAGE:a8f64888d57debe8_1_4]]

[[IMAGE:a8f64888d57debe8_1_5]]

[[IMAGE:a8f64888d57debe8_1_6]]
**LLM**
**Section Id :** 64065391701
**Section Number :** 8
**Section type :** Online
**Mandatory or Optional :** Mandatory
**Number of Questions :** 16
**Number of Questions to be attempted :** 16
**Section Marks :** 50
**Display Number Panel :** Yes
**Section Negative Marks :** 0
**Group All Questions :** No
**Enable Mark as Answered Mark for Review and**
No
**Clear Response :**
**Section Maximum Duration :** 0
**Section Minimum Duration :** 0
**Section Time In :** Minutes
**Maximum Instruction Time :** 0

A published solution is not available for this question yet.
Question 135 MCQ · 2.0 marks
Transformers process input tokens:
One at a time (sequentially)
In reverse order
All at once (in parallel)
Only after seeing the full input
A published solution is not available for this question yet.
Question 136 MCQ · 2.0 marks
What is the purpose of the softmax function in the attention mechanism?
Normalize attention scores to a probability distribution
Add non-linearity to the model
To predict the correct class for the loss function
Remove redundant features from the input
A published solution is not available for this question yet.
Question 137 MCQ · 2.0 marks
What is the main difference between GPT and BERT pre-training objectives?
GPT uses Masked Language Modeling, BERT uses Causal Language Modeling
GPT uses Causal Language Modeling, BERT uses Masked Language Modeling
Both use Masked Language Modeling
Both use Causal Language Modeling
A published solution is not available for this question yet.
Question 138 MCQ · 2.0 marks
In Top-K sampling for language generation, increasing the value of K typically has which of the
following effects?
It makes the output more deterministic and repetitive.
It reduces the probability of selecting high-frequency words.
It increases the diversity of the generated text but may reduce coherence if K
is too large.
It guarantees grammatical correctness by focusing on top-ranked tokens only.
A published solution is not available for this question yet.
Question 139 MCQ · 2.0 marks
Why are residual connections important in transformer architectures?
They reduce memory consumption.
They add extra cost of compute by adding batch normalization.
They help in training deep networks by enabling gradient flow.
They remove the need for layer normalization.
A published solution is not available for this question yet.
Question 140 MSQ · 3.0 marks
Which of the following statements are true regarding causal language modeling (CLM)?
The model only attends to past and current tokens during training.
The model is trained by predicting the next token in a sequence.
The model uses bidirectional context.
The CLM objective is commonly used for encoder-only models.
A published solution is not available for this question yet.
Question 141 MSQ · 3.0 marks
Which of the following are valid reasons why transformer-based large language models are widely
used in natural language processing?
Transformers process input sequences in parallel, enabling faster training.
They use recurrence to remember long-term dependencies more effectively
than LSTMs.
Transformers are limited to short text inputs due to their architecture.
Large transformer models can be fine-tuned for various NLP tasks using a
single pre-trained model.
A published solution is not available for this question yet.
Question 142 MSQ · 3.0 marks
Which elements are included in BERT’s input representation for Next Sentence Prediction?
[[IMAGE:a8f64888d57debe8_5_7]]

[[IMAGE:a8f64888d57debe8_5_8]]

[[IMAGE:a8f64888d57debe8_5_9]]

[[IMAGE:a8f64888d57debe8_5_10]]

A published solution is not available for this question yet.
Question 143 MSQ · 3.0 marks
How does Top-p (nucleus) sampling differ from Top-K sampling in language generation?
It samples only from a fixed number of tokens at each step.
It samples from the smallest set of tokens whose cumulative probability
exceeds p.
It always selects the top-p tokens with equal probability.
It guarantees diversity by selecting all low-probability tokens.
A published solution is not available for this question yet.
Question 144 NAT · 3.0 marks
[[IMAGE:a8f64888d57debe8_6_11]]

A published solution is not available for this question yet.
Question 145 NAT · 3.0 marks
[[IMAGE:a8f64888d57debe8_7_12]]

A published solution is not available for this question yet.
Question 146 NAT · 3.0 marks
[[IMAGE:a8f64888d57debe8_7_13]]

A published solution is not available for this question yet.
Question 147 MCQ · 3.0 marks
[[IMAGE:a8f64888d57debe8_9_14]]
Based on the above data, answer the given subquestions.
Select the scaled dot-product attention for the first head:

[[IMAGE:a8f64888d57debe8_10_15]]

[[IMAGE:a8f64888d57debe8_10_16]]

[[IMAGE:a8f64888d57debe8_10_17]]

[[IMAGE:a8f64888d57debe8_10_18]]

A published solution is not available for this question yet.
Question 148 MCQ · 3.0 marks
[[IMAGE:a8f64888d57debe8_9_14]]
Based on the above data, answer the given subquestions.
Select the scaled dot-product attention for the second head:

[[IMAGE:a8f64888d57debe8_10_19]]

[[IMAGE:a8f64888d57debe8_10_20]]

[[IMAGE:a8f64888d57debe8_10_21]]

[[IMAGE:a8f64888d57debe8_10_22]]

A published solution is not available for this question yet.
Question 149 MCQ · 2.0 marks
[[IMAGE:a8f64888d57debe8_9_14]]
Based on the above data, answer the given subquestions.
Concatenate the outputs from both the attention heads, then apply the output projection matrix
Wo to produce the final output of the multi-head attention mechanism. Select the correct result of
this operation.

[[IMAGE:a8f64888d57debe8_11_23]]

[[IMAGE:a8f64888d57debe8_11_24]]

[[IMAGE:a8f64888d57debe8_11_25]]

[[IMAGE:a8f64888d57debe8_11_26]]

A published solution is not available for this question yet.
Question 150 NAT · 1.0 marks
[[IMAGE:a8f64888d57debe8_12_27]]
Based on the above data, answer the given subquestions.
For the given input matrix X and multihead attention output MHA(X), apply a residual connection
and store the result in matrix R, and finally compute the sum of all elements in matrix R.

A published solution is not available for this question yet.
Question 151 NAT · 2.0 marks
[[IMAGE:a8f64888d57debe8_12_27]]
Based on the above data, answer the given subquestions.
[[IMAGE:a8f64888d57debe8_13_28]]


A published solution is not available for this question yet.
Question 152 NAT · 3.0 marks
[[IMAGE:a8f64888d57debe8_12_27]]
Based on the above data, answer the given subquestions.
[[IMAGE:a8f64888d57debe8_13_29]]


A published solution is not available for this question yet.
Question 153 NAT · 2.0 marks
The table presents the **conditional probability distribution** over vocabulary tokens at each
timestep during sequence generation. Each **column** corresponds to a timestep (from 1 to 5), and
the values represent the probability of selecting each token given the tokens chosen in all previous
timesteps. For timestep t, the values in the column represent:
[[IMAGE:a8f64888d57debe8_14_30]]
Based on the above data, answer the given subquestions.
In exhaustive search, at timestep t=1, we run the decoder once to obtain probability distributions
over all tokens in the vocabulary. Given the 6 tokens available in the table in the main question,
how many times must we run the decoder at timestep t=4?

A published solution is not available for this question yet.
Question 154 NAT · 1.0 marks
The table presents the **conditional probability distribution** over vocabulary tokens at each
timestep during sequence generation. Each **column** corresponds to a timestep (from 1 to 5), and
the values represent the probability of selecting each token given the tokens chosen in all previous
timesteps. For timestep t, the values in the column represent:
[[IMAGE:a8f64888d57debe8_14_30]]
Based on the above data, answer the given subquestions.
How many total sequences of exactly length 5 are possible according to exhaustive search?

A published solution is not available for this question yet.
Question 155 NAT · 2.0 marks
The table presents the **conditional probability distribution** over vocabulary tokens at each
timestep during sequence generation. Each **column** corresponds to a timestep (from 1 to 5), and
the values represent the probability of selecting each token given the tokens chosen in all previous
timesteps. For timestep t, the values in the column represent:
[[IMAGE:a8f64888d57debe8_14_30]]
Based on the above data, answer the given subquestions.
If we use top-k sampling with k=2 at timestep 1, what is the normalized probability of selecting
token “sky” at the timestep=1?

A published solution is not available for this question yet.