MauryaHub PYQ Practice

da5004_2026T1_Q1_NA.pdf

Large Language Models · Quiz 1 · Jan 2026

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 2 MCQ · 2.0 marks

In a self-attention mechanism, let the input be a matrix [[IMAGE:36b97be09a59b1a1_3_2]] , where [[IMAGE:36b97be09a59b1a1_3_3]] is the sequence length. If the dimension of the query projection matrix is [[IMAGE:36b97be09a59b1a1_3_4]] , what is the dimension of the Query matrix [[IMAGE:36b97be09a59b1a1_3_5]] ?
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
  1. [[IMAGE:36b97be09a59b1a1_3_6]]
    Source diagram or notation
  2. [[IMAGE:36b97be09a59b1a1_3_7]]
    Source diagram or notation
  3. [[IMAGE:36b97be09a59b1a1_3_8]]
    Source diagram or notation
  4. [[IMAGE:36b97be09a59b1a1_3_9]]
    Source diagram or notation

A published solution is not available for this question yet.

Question 3 MCQ · 2.0 marks

[[IMAGE:36b97be09a59b1a1_3_10]] In the scaled dot-product attention equation , what is the primary reason for dividing the dot product by [[IMAGE:36b97be09a59b1a1_4_11]] ?
Source diagram or notationSource diagram or notation
  1. To ensure the matrix dimensions match for multiplication with [[IMAGE:36b97be09a59b1a1_4_12]]
    Source diagram or notation
  2. To reduce the number of trainable parameters in the model
  3. To prevent the dot product values from growing too large, which would push the softmax function into regions with extremely small gradients
  4. To normalize the embedding vectors to have a unit length

A published solution is not available for this question yet.

Question 4 MCQ · 2.0 marks

Consider a Transformer Encoder stack consisting of [[IMAGE:36b97be09a59b1a1_4_13]] identical layers. If the input to the first layer is a sequence of word embeddings of shape [[IMAGE:36b97be09a59b1a1_4_14]] (where [[IMAGE:36b97be09a59b1a1_4_15]] and [[IMAGE:36b97be09a59b1a1_4_16]] ), what is the shape of the output of the final (6th) layer?
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
  1. [[IMAGE:36b97be09a59b1a1_4_17]] (A single context vector)
    Source diagram or notation
  2. [[IMAGE:36b97be09a59b1a1_4_18]] (One vector per layer)
    Source diagram or notation
  3. [[IMAGE:36b97be09a59b1a1_4_19]] (One vector per token, preserving sequence length)
    Source diagram or notation
  4. [[IMAGE:36b97be09a59b1a1_4_20]] (Concatenation of all token vectors across layers)
    Source diagram or notation

A published solution is not available for this question yet.

Question 5 MCQ · 2.0 marks

In a standard Transformer encoder-decoder architecture, is masking typically applied in the cross- attention layer?
  1. Yes, both causal masking and padding masking are applied to prevent the decoder from attending to future encoder positions
  2. Yes, only causal masking is applied to maintain the autoregressive property
  3. No causal masking is needed, but padding masking may be applied to ignore padded positions in the source sequence
  4. No masking is ever applied in cross-attention since the encoder processes the complete input sequence

A published solution is not available for this question yet.

Question 6 MCQ · 2.0 marks

Consider the BERT architecture. During pre-training, it uses segment embeddings ( [[IMAGE:36b97be09a59b1a1_5_21]] ) to distinguish between sentences. However, during fine-tuning for a **single-sentence classification** **task** (e.g., Sentiment Analysis), how are these segment embeddings typically utilized?
Source diagram or notation
  1. Segment embeddings are not used; the embedding layer is disabled for single-sentence tasks.
  2. All tokens in the input sequence are assigned the same segment embedding (e.g., [[IMAGE:36b97be09a59b1a1_5_22]] ).
    Source diagram or notation
  3. The first half of the sentence gets [[IMAGE:36b97be09a59b1a1_5_23]] and the second half gets [[IMAGE:36b97be09a59b1a1_5_24]] to maintain symmetry.
    Source diagram or notationSource diagram or notation
  4. Segment embeddings are randomly initialized for every new single-sentence input.

A published solution is not available for this question yet.

Question 7 MCQ · 2.0 marks

Consider a decoding step in a language model where the logits for the next token are [[IMAGE:36b97be09a59b1a1_5_25]] . We apply a temperature scaling [[IMAGE:36b97be09a59b1a1_5_26]] before the softmax. As [[IMAGE:36b97be09a59b1a1_5_27]] (approaches zero), what does the resulting probability distribution [[IMAGE:36b97be09a59b1a1_5_28]] look like?
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
  1. It approaches a uniform distribution where all tokens have equal probability.
  2. It approaches a "one-hot" distribution where the token with the highest logit has probability 1.0 (Greedy Search).
  3. It remains unchanged from the original softmax distribution.
  4. It causes numerical instability and results in NaNs.

A published solution is not available for this question yet.

Question 8 MCQ · 2.0 marks

You are fine-tuning a pre-trained BERT model for a spam classification task. The input sequence is: [[IMAGE:36b97be09a59b1a1_5_29]] . Which vector from the final layer output is typically passed to the classification head ( [[IMAGE:36b97be09a59b1a1_5_30]] )?
Source diagram or notationSource diagram or notation
  1. The average pooling of all token vectors.
  2. The final hidden state vector corresponding to the [[IMAGE:36b97be09a59b1a1_6_31]] token ( [[IMAGE:36b97be09a59b1a1_6_32]] ).
    Source diagram or notationSource diagram or notation
  3. The final hidden state vector corresponding to the [[IMAGE:36b97be09a59b1a1_6_33]] token ( [[IMAGE:36b97be09a59b1a1_6_34]] ).
    Source diagram or notationSource diagram or notation
  4. The concatenation of all hidden state vectors.

A published solution is not available for this question yet.

Question 9 MSQ · 3.0 marks

Select all statements that correctly describe the motivation and behavior of **Multi-Head Attention** as compared to single-head attention.
  1. It allows the model to jointly attend to information from different representation subspaces at different positions.
  2. It is mathematically similar to having multiple filters/kernels in a CNN to capture different features.
  3. It reduces the total number of parameters required compared to a single head with the same total dimension.
  4. Each head can theoretically learn to capture different linguistic relationships (e.g., one head links "it" to "animal", another links "it" to "tired").

A published solution is not available for this question yet.

Question 10 MSQ · 3.0 marks

[[IMAGE:36b97be09a59b1a1_6_35]] Given the vectorized self-attention calculation , select all true statements regarding the matrix dimensions and operations. Assume [[IMAGE:36b97be09a59b1a1_6_36]] .
Source diagram or notationSource diagram or notation
  1. The matrix resulting from [[IMAGE:36b97be09a59b1a1_6_37]] has dimensions [[IMAGE:36b97be09a59b1a1_6_38]] .
    Source diagram or notationSource diagram or notation
  2. The softmax operation is applied to the entire matrix at once (globally), not row-wise.
  3. The final output [[IMAGE:36b97be09a59b1a1_6_39]] has dimensions [[IMAGE:36b97be09a59b1a1_6_40]] .
    Source diagram or notationSource diagram or notation
  4. The matrix [[IMAGE:36b97be09a59b1a1_7_41]] represents the raw affinity/score between every pair of tokens in the sequence.
    Source diagram or notation

A published solution is not available for this question yet.

Question 11 MSQ · 2.0 marks

In the context of the Transformer architecture, which of the following statements correctly describe the source of the Query ( [[IMAGE:36b97be09a59b1a1_7_42]] ), Key ( [[IMAGE:36b97be09a59b1a1_7_43]] ), and Value ( [[IMAGE:36b97be09a59b1a1_7_44]] ) vectors during the Cross-Attention (Encoder-Decoder Attention) mechanism?
Source diagram or notationSource diagram or notationSource diagram or notation
  1. The Keys ( [[IMAGE:36b97be09a59b1a1_7_45]] ) and Values ( [[IMAGE:36b97be09a59b1a1_7_46]] ) are derived from the output of the Decoder.
    Source diagram or notationSource diagram or notation
  2. The Queries ( [[IMAGE:36b97be09a59b1a1_7_47]] ) are generated from the previous hidden states of the Decoder.
    Source diagram or notation
  3. The Queries ( [[IMAGE:36b97be09a59b1a1_7_48]] ), Keys ( [[IMAGE:36b97be09a59b1a1_7_49]] ), and Values ( [[IMAGE:36b97be09a59b1a1_7_50]] ) all originate from the same input sequence.
    Source diagram or notationSource diagram or notationSource diagram or notation
  4. This mechanism allows the Decoder to focus on relevant parts of the input sequence processed by the Encoder.

A published solution is not available for this question yet.

Question 12 NAT · 2.0 marks

Consider the positional encoding for a given position [[IMAGE:36b97be09a59b1a1_7_51]] and dimension [[IMAGE:36b97be09a59b1a1_7_52]] is defined by: [[IMAGE:36b97be09a59b1a1_7_53]] [[IMAGE:36b97be09a59b1a1_7_54]] Consider a Transformer model with a hidden state dimension [[IMAGE:36b97be09a59b1a1_8_55]] . You are calculating the positional embedding for the 4th word in a sequence (index [[IMAGE:36b97be09a59b1a1_8_56]] , using 0-based indexing). Calculate the value of the 6th element (index [[IMAGE:36b97be09a59b1a1_8_57]] ) of the positional embedding vector for this word. (use radians for angle)
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

    A published solution is not available for this question yet.

    Question 13 NAT · 1.0 marks

    Suppose you are given three encoder hidden states at time [[IMAGE:36b97be09a59b1a1_8_58]] : [[IMAGE:36b97be09a59b1a1_8_59]] The previous decoder hidden state is: [[IMAGE:36b97be09a59b1a1_8_60]] Given the attention score function: [[IMAGE:36b97be09a59b1a1_8_61]] where the hyperbolic tangent function is defined as: [[IMAGE:36b97be09a59b1a1_8_62]] and [[IMAGE:36b97be09a59b1a1_9_63]] Based on the above data, answer the given subquestions.
    Compute the attention score for hidden state [[IMAGE:36b97be09a59b1a1_9_64]] using the given function. Enter the final answer correct to two decimal places.
    Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

      A published solution is not available for this question yet.

      Question 14 NAT · 3.0 marks

      Suppose you are given three encoder hidden states at time [[IMAGE:36b97be09a59b1a1_8_58]] : [[IMAGE:36b97be09a59b1a1_8_59]] The previous decoder hidden state is: [[IMAGE:36b97be09a59b1a1_8_60]] Given the attention score function: [[IMAGE:36b97be09a59b1a1_8_61]] where the hyperbolic tangent function is defined as: [[IMAGE:36b97be09a59b1a1_8_62]] and [[IMAGE:36b97be09a59b1a1_9_63]] Based on the above data, answer the given subquestions.
      Normalize the attention scores using the softmax function to obtain the attention weights [[IMAGE:36b97be09a59b1a1_9_65]] . Submit [[IMAGE:36b97be09a59b1a1_9_66]] (i.e., the first element of the [[IMAGE:36b97be09a59b1a1_9_67]] vector). Enter the final answer correct to two decimal places. [[IMAGE:36b97be09a59b1a1_9_68]]
      Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

        A published solution is not available for this question yet.

        Question 15 NAT · 2.0 marks

        Suppose you are given three encoder hidden states at time [[IMAGE:36b97be09a59b1a1_8_58]] : [[IMAGE:36b97be09a59b1a1_8_59]] The previous decoder hidden state is: [[IMAGE:36b97be09a59b1a1_8_60]] Given the attention score function: [[IMAGE:36b97be09a59b1a1_8_61]] where the hyperbolic tangent function is defined as: [[IMAGE:36b97be09a59b1a1_8_62]] and [[IMAGE:36b97be09a59b1a1_9_63]] Based on the above data, answer the given subquestions.
        Calculate the context vector [[IMAGE:36b97be09a59b1a1_10_69]] : [[IMAGE:36b97be09a59b1a1_10_70]] Provide the **sum of all elements** of [[IMAGE:36b97be09a59b1a1_10_71]] . Enter the final answer correct to two decimal places.
        Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

          A published solution is not available for this question yet.

          Question 16 MCQ · 2.0 marks

          Consider the following Configuration for a GPT model: • Vocabulary size: 40,000 tokens • Embedding dimension (d_model): 768 • Maximum sequence length: 512 • Number of transformer blocks: 12 • Number of attention heads per block: 12 • Feed-forward network hidden dimension: 3,072 • Activation function: GELU Note: In multi-head attention, the model dimension is split equally among all heads. Based on the above data, answer the given subquestions.
          What is the number of parameters in the token embedding matrix?
          1. 30,720,000
          2. 40,768
          3. 393,216
          4. 15,728,640,000

          A published solution is not available for this question yet.

          Question 17 MCQ · 2.0 marks

          Consider the following Configuration for a GPT model: • Vocabulary size: 40,000 tokens • Embedding dimension (d_model): 768 • Maximum sequence length: 512 • Number of transformer blocks: 12 • Number of attention heads per block: 12 • Feed-forward network hidden dimension: 3,072 • Activation function: GELU Note: In multi-head attention, the model dimension is split equally among all heads. Based on the above data, answer the given subquestions.
          What is the total number of parameters in the positional embedding matrix?
          1. 512,768
          2. 393,216
          3. 3,932,160
          4. 39,321

          A published solution is not available for this question yet.

          Question 18 MCQ · 2.0 marks

          Consider the following Configuration for a GPT model: • Vocabulary size: 40,000 tokens • Embedding dimension (d_model): 768 • Maximum sequence length: 512 • Number of transformer blocks: 12 • Number of attention heads per block: 12 • Feed-forward network hidden dimension: 3,072 • Activation function: GELU Note: In multi-head attention, the model dimension is split equally among all heads. Based on the above data, answer the given subquestions.
          For ONE complete attention head, what is the total number of parameters (Q + K + V projections combined)?
          1. 49,152
          2. 196,608
          3. 589,824
          4. 147,456

          A published solution is not available for this question yet.

          Question 19 MCQ · 2.0 marks

          Consider the following Configuration for a GPT model: • Vocabulary size: 40,000 tokens • Embedding dimension (d_model): 768 • Maximum sequence length: 512 • Number of transformer blocks: 12 • Number of attention heads per block: 12 • Feed-forward network hidden dimension: 3,072 • Activation function: GELU Note: In multi-head attention, the model dimension is split equally among all heads. Based on the above data, answer the given subquestions.
          After the multi-head attention computation, all head outputs are concatenated and projected back to the model dimension. How many parameters are in the output projection matrix (weights only, excluding bias)?
          1. 768
          2. 49,152
          3. 589,824
          4. 294,912

          A published solution is not available for this question yet.

          Question 20 MCQ · 2.0 marks

          Consider the following Configuration for a GPT model: • Vocabulary size: 40,000 tokens • Embedding dimension (d_model): 768 • Maximum sequence length: 512 • Number of transformer blocks: 12 • Number of attention heads per block: 12 • Feed-forward network hidden dimension: 3,072 • Activation function: GELU Note: In multi-head attention, the model dimension is split equally among all heads. Based on the above data, answer the given subquestions.
          What is the total number of parameters in the complete FFN (including both weight matrices and bias vectors)?
          1. 2,359,296
          2. 4,718,592
          3. 4,722,432
          4. 9,437,184

          A published solution is not available for this question yet.