MauryaHub PYQ Practice

da5004_2026T2_ET_FN.pdf

Large Language Models · End Term · May 2026 FN

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 2 NAT · 2.0 marks

Consider a Multi-Head Attention sub-layer with the following hyper-parameters: • Embedding dimension ( [[IMAGE:2ef6735b833c260e_2_2]] ) = 256 • Number of heads ( [[IMAGE:2ef6735b833c260e_2_3]] ) = 4 • key/value dimension per head ( [[IMAGE:2ef6735b833c260e_2_4]] ) = [[IMAGE:2ef6735b833c260e_2_5]] . Calculate the total number of trainable parameters contained **only** in the projection matrices [[IMAGE:2ef6735b833c260e_2_6]] and [[IMAGE:2ef6735b833c260e_2_7]] . (Ignore biases). Enter the integer value.
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

    A published solution is not available for this question yet.

    Question 3 NAT · 2.0 marks

    For T5-Base, the model dimension is 768 and FFN dimension is 3072. Ignore bias. Find the number of parameters in the FFN of one encoder layer.

      A published solution is not available for this question yet.

      Question 4 NAT · 3.0 marks

      A batch of 5 sentences is passed through a Transformer Encoder. The sequences are padded to a maximum length of 20 tokens. The embedding dimension ( [[IMAGE:2ef6735b833c260e_3_8]] ) is 64. What is the total number of elements (volume) in the output tensor produced by the final layer of the encoder for this batch?
      Source diagram or notation

        A published solution is not available for this question yet.

        Question 5 NAT · 3.0 marks

        Calculate the KV Cache memory size in Megabytes for a single request (batch size = 1) with the following parameters: • Number of layers: 24 • Embedding dimension (d_model): 1024 • Sequence length (n): 2048 tokens • Precision: 16-bit (2 bytes per element) • Assume standard Multi-Head Attention (MHA) where total KV dimension per layer is 2 x d_model (2 for K+V)

          A published solution is not available for this question yet.

          Question 6 MCQ · 2.0 marks

          In BERT's Next Sentence Prediction task, how are the negative (NotNext) examples created?
          1. By taking the sentence that immediately precedes the current sentence.
          2. By reversing the word order of the actual next sentence.
          3. By picking a random sentence from the corpus (50% of the time).
          4. By masking all verbs in the actual next sentence.

          A published solution is not available for this question yet.

          Question 7 MCQ · 2.0 marks

          What happens when a very large model is trained for many epochs on a relatively small dataset in the T5 scaling-law setting?
          1. The performance continues to improve linearly with more training steps.
          2. The performance degrades due to overfitting (memorization) of the training data.
          3. The model automatically learns to augment the data, preventing overfitting.
          4. The performance plateaus but does not degrade.

          A published solution is not available for this question yet.

          Question 8 MCQ · 2.0 marks

          **PagedAttention** is designed to address which specific inefficiency in Large Language Model (LLM) serving?
          1. High computational latency of matrix multiplication.
          2. Memory fragmentation and waste in the KV cache due to pre-allocating contiguous memory blocks for variable-length sequences.
          3. Low bandwidth of the internet connection during API calls.
          4. The difficulty of training models on long sequences.

          A published solution is not available for this question yet.

          Question 9 MCQ · 2.0 marks

          The standard Self-Attention mechanism has a time complexity of [[IMAGE:2ef6735b833c260e_5_9]] , where [[IMAGE:2ef6735b833c260e_5_10]] is the sequence length and [[IMAGE:2ef6735b833c260e_5_11]] is the embedding dimension. In Kernel-based Linear Attention methods (like Linear Transformers), the associativity of matrix multiplication is exploited to compute [[IMAGE:2ef6735b833c260e_5_12]] instead of [[IMAGE:2ef6735b833c260e_5_13]] . What is the resulting time complexity of this Linear Attention mechanism?
          Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
          1. [[IMAGE:2ef6735b833c260e_5_14]]
            Source diagram or notation
          2. [[IMAGE:2ef6735b833c260e_5_15]]
            Source diagram or notation
          3. [[IMAGE:2ef6735b833c260e_5_16]]
            Source diagram or notation
          4. [[IMAGE:2ef6735b833c260e_5_17]]
            Source diagram or notation

          A published solution is not available for this question yet.

          Question 10 MCQ · 2.0 marks

          In a standard Transformer layer utilizing Rotary Positional Embeddings (RoPE), at which specific stage of the forward pass is the rotary transformation applied?
          1. To the input embeddings [[IMAGE:2ef6735b833c260e_5_18]] immediately before the Linear projections for [[IMAGE:2ef6735b833c260e_5_19]] , [[IMAGE:2ef6735b833c260e_5_20]] , and [[IMAGE:2ef6735b833c260e_5_21]] .
            Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
          2. To the [[IMAGE:2ef6735b833c260e_5_22]] and [[IMAGE:2ef6735b833c260e_5_23]] vectors simultaneously using a coupled weight matrix within the Attention head.
            Source diagram or notationSource diagram or notation
          3. To the [[IMAGE:2ef6735b833c260e_5_24]] and [[IMAGE:2ef6735b833c260e_5_25]] vectors independently, after their linear projections but before the scaled dot-product calculation.
            Source diagram or notationSource diagram or notation
          4. To the [[IMAGE:2ef6735b833c260e_5_26]] and [[IMAGE:2ef6735b833c260e_5_27]] vectors after the linear projections to ensure the entire latent space is rotationally invariant.
            Source diagram or notationSource diagram or notation

          A published solution is not available for this question yet.

          Question 11 MCQ · 2.0 marks

          In a Transformer using ALiBi, the attention score is modified as: [[IMAGE:2ef6735b833c260e_6_28]] Suppose a researcher proposes modifying ALiBi to: [[IMAGE:2ef6735b833c260e_6_29]] and claims that: “The model will still learn to focus on nearby tokens through training.” Which of the following is the most accurate critique of this proposal?
          Source diagram or notationSource diagram or notation
          1. The modification will not affect attention behavior significantly because softmax normalizes all scores, making additive biases irrelevant.
          2. The modification will cause instability because the attention scores will become negative for nearby tokens and positive for distant tokens.
          3. The modification is equivalent to scaling (Q) and (K) vectors differently and therefore does not fundamentally change the positional bias.
          4. The modification will bias attention toward distant tokens, and due to the exponential nature of softmax, this effect can dominate learned [[IMAGE:2ef6735b833c260e_6_30]] similarities, making it difficult for training to recover locality.
            Source diagram or notation

          A published solution is not available for this question yet.

          Question 12 MSQ · 3.0 marks

          Comparing the architectures of BERT and GPT, select all the **structural differences** that are true.
          1. BERT uses bidirectional self-attention, allowing tokens to attend to both left and right contexts.
          2. GPT uses masked (causal) self-attention, allowing tokens to attend only to previous tokens.
          3. BERT is an autoregressive model, while GPT is an autoencoding model.
          4. During pre-training, BERT predicts masked tokens, whereas GPT predicts the next token in the sequence.

          A published solution is not available for this question yet.

          Question 13 MSQ · 3.0 marks

          Based on the architectural comparison experiments (Encoder-Decoder vs. Decoder-only vs. Encoder-only) discussed in the lectures:
          1. Encoder-Decoder architectures generally perform best for sequence-to- sequence tasks like Translation and Summarization.
          2. A Decoder-only model with a **Prefix-LM** objective can match the performance of an Encoder-Decoder model.
          3. Sharing parameters between the encoder and decoder always improves performance regardless of task.
          4. Encoder-only models (like BERT) are typically less suitable for generation tasks compared to Encoder-Decoder models.

          A published solution is not available for this question yet.

          Question 14 MSQ · 3.0 marks

          Select all true statements regarding **KV Caching** during the autoregressive inference of LLMs.
          1. It trades off increased memory usage for reduced computational latency.
          2. It avoids recomputing the Key and Value vectors for tokens that have already been processed in previous steps.
          3. It is essential during the training phase to speed up backpropagation.
          4. The memory required for the KV cache grows linearly with the sequence length and batch size.

          A published solution is not available for this question yet.

          Question 15 MSQ · 2.0 marks

          Select all correct statements regarding Decoding Strategies in language generation.
          1. Beam Search with a beam width [[IMAGE:2ef6735b833c260e_8_31]] is mathematically equivalent to Greedy Search.
            Source diagram or notation
          2. Greedy search is deterministic and often leads to repetitive or degenerative text loops.
          3. Top- [[IMAGE:2ef6735b833c260e_8_32]] sampling guarantees that the generated text will always be grammatically correct.
            Source diagram or notation
          4. Top- [[IMAGE:2ef6735b833c260e_8_33]] (Nucleus) sampling allows for a dynamic vocabulary size at each step, whereas Top- [[IMAGE:2ef6735b833c260e_8_34]] uses a fixed number of candidates.
            Source diagram or notationSource diagram or notation

          A published solution is not available for this question yet.

          Question 16 MSQ · 2.0 marks

          In a feature-based approach using a pretrained transformer (like BERT) for a Multiple-Choice QA task, which of the following statements regarding the methodology are correct? (Select all that apply)
          1. The weights of the pretrained transformer are updated via backpropagation during the training phase.
          2. The input must be formatted as multiple pairs, such as [CLS] Question [SEP] Choice_n, creating a distinct representation for each candidate answer.
          3. The final prediction is determined by passing the fixed output representations through a trainable linear layer and applying a Softmax function.
          4. The model generates the correct answer string token-by-token using an autoregressive decoding strategy like Greedy Search.

          A published solution is not available for this question yet.

          Question 17 MCQ · 3.0 marks

          A Transformer processes a sequence of length 5 using 6 identical self-attention layers. Count the total number of allowed attention links across all layers and Choose the option representing the correct pair (A) BERT (B) GPT
          1. BERT: 150, GPT: 90
          2. BERT: 90, GPT: 150
          3. BERT: 150, GPT: 150
          4. BERT: 90, GPT: 90
          5. BERT: 30, GPT: 36
          6. BERT: 36, GPT: 30

          A published solution is not available for this question yet.

          Question 18 MCQ · 3.0 marks

          Consider the architecture of **BART**. Which of the following descriptions best match its structural design?
          1. A Decoder-only model similar to GPT, but with bidirectional attention in the first layer.
          2. An Encoder-only model similar to BERT, but trained with a causal masking objective.
          3. An encoder-decoder model where the encoder processes a corrupted input bidirectionally and the decoder autoregressively reconstructs the original sequence.
          4. A dual-encoder model where one encoder processes the context and another processes the query.

          A published solution is not available for this question yet.

          Question 19 MCQ · 3.0 marks

          Why is **Flash Attention** considered an "IO-aware" algorithm?
          1. It reduces the number of parameters in the model to fit in GPU memory.
          2. It compresses the input data using JPEG-like encoding before processing.
          3. It optimizes the movement of data between the high-bandwidth memory (HBM) and the faster on-chip SRAM to avoid memory bandwidth bottlenecks.
          4. It writes all intermediate attention matrices to the hard disk to save RAM.

          A published solution is not available for this question yet.

          Question 20 MCQ · 3.0 marks

          In a transformer with relative positional encoding (RPE), attention logits between a query at position 4 and keys at positions 2 and 6 depend on relative distances [[IMAGE:2ef6735b833c260e_9_35]] where [[IMAGE:2ef6735b833c260e_9_36]] is the position of query and [[IMAGE:2ef6735b833c260e_9_37]] is the position of key. Suppose the learned scalar biases as per RPE are: • [[IMAGE:2ef6735b833c260e_10_38]] • [[IMAGE:2ef6735b833c260e_10_39]] • [[IMAGE:2ef6735b833c260e_10_40]] • [[IMAGE:2ef6735b833c260e_10_41]] • [[IMAGE:2ef6735b833c260e_10_42]] The base (content-only) attention logits are: • [[IMAGE:2ef6735b833c260e_10_43]] • [[IMAGE:2ef6735b833c260e_10_44]] What are the final attention logits after incorporating RPE?
          Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
          1. [[IMAGE:2ef6735b833c260e_10_45]]
            Source diagram or notation
          2. [[IMAGE:2ef6735b833c260e_10_46]]
            Source diagram or notation
          3. [[IMAGE:2ef6735b833c260e_10_47]]
            Source diagram or notation
          4. [[IMAGE:2ef6735b833c260e_10_48]]
            Source diagram or notation

          A published solution is not available for this question yet.

          Question 21 MCQ · 3.0 marks

          An engineer is designing a Transformer model capable of length extrapolation (handling sequences longer than those seen during training) while minimizing the number of trainable weights. They are evaluating Absolute Positional Encodings (APE), Relative Positional Encodings (RPE), Rotary Positional Encodings (RoPE), and Attention with Linear Biases (ALiBi).Which of the following statements correctly describes the parameterization and behavior of these techniques?
          1. RoPE is considered a learned encoding because the model must optimize the rotation angles [[IMAGE:2ef6735b833c260e_10_49]] as trainable parameters during the backpropagation phase.
            Source diagram or notation
          2. The engineer should choose APE or RPE to minimize weights, as both rely exclusively on fixed sinusoidal functions without requiring a lookup table.
          3. RoPE and ALiBi are preferable for this use case because they utilize fixed mathematical transformations, whereas standard APE and RPE typically require learning position- specific embedding weights.
          4. ALiBi is the only technique among the four that requires the model to learn a unique "slope" parameter for each attention head to determine the rate of penalty decay.

          A published solution is not available for this question yet.