MauryaHub PYQ Practice

da5004_2026T2_Q2_NA.pdf

Large Language Models · Quiz 2 · May 2026

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 2 SHORT_TEXT · 2.0 marks

Answer the subquestions for the given transformer architecture for one (N = 1) **Encoder Decoder** **Block** [[IMAGE:f854547ec11886a8_2_2]]
In the given Transformer architecture for one Encoder Decoder block, identify every **Add & Norm** layer. Submit the corresponding component alphabets in **alphabetical (ascending) order** as a single uppercase string without spaces, commas, or quotation marks. Example: If the correct components are Q, G, and S, your answer should be GQS
Source diagram or notation

    A published solution is not available for this question yet.

    Question 3 MCQ · 1.0 marks

    Answer the subquestions for the given transformer architecture for one (N = 1) **Encoder Decoder** **Block** [[IMAGE:f854547ec11886a8_2_2]]
    Identify the function of component **"L"** in the given Transformer architecture.
    Source diagram or notation
    1. Positional Encoding
    2. Feed-Forward Network
    3. Masked Multi-Head Self-Attention
    4. Multi-Head Cross-Attention
    5. Linear Output Projection

    A published solution is not available for this question yet.

    Question 4 MSQ · 2.0 marks

    Answer the subquestions for the given transformer architecture for one (N = 1) **Encoder Decoder** **Block** [[IMAGE:f854547ec11886a8_2_2]]
    Based on the architecture shown in the given image, identify all components that appear more than once. Assume the architecture contains exactly one Encoder block and one Decoder block.
    Source diagram or notation
    1. Positional Encoding
    2. Feed-Forward Network
    3. Masked Multi-Head Self-Attention
    4. Multi-Head Cross-Attention
    5. tokenizer
    6. Add & Norm

    A published solution is not available for this question yet.

    Question 5 MSQ · 3.0 marks

    [[IMAGE:f854547ec11886a8_4_3]]
    Source diagram or notation
    1. The approach violates the autoregressive property
    2. The approach ensures autoregressive property is not violated by ensuring that the final attention weights of future tokens are 0
    3. The final attention weights in the new approach will be greater than or equal to that of the original approach
    4. The attention weights for the final token of the sequence will always be the same for both the approaches

    A published solution is not available for this question yet.

    Question 6 MSQ · 3.0 marks

    Which of the following options correctly reflect the internal state updates performed during a Byte Pair Encoding (BPE) training iteration?
    1. After a pair is chosen for merging, it is permanently added to the active vocabulary list.
    2. The algorithm recalculates the log probability of the entire corpus before the next iteration.
    3. The independent frequencies of the two constituent tokens that were part of the merge are reduced by the frequency of the merged token.
    4. The original constituent tokens of the merged pair are removed from the vocabulary to optimize dictionary size.

    A published solution is not available for this question yet.

    Question 7 MSQ · 3.0 marks

    A baseline T5 model with denoising objective on a span of corrupted tokens is used for unsupervised pre-training. Consider the phrase **"life is like a box of chocolates"**. If the word "life" and the span "box of chocolates" are selected for corruption, which of the following statements correctly describe the training process?
    1. The input sequence replaces the corrupted spans with unique sentinal tokens as: "<X> is like a <Y>".
    2. The output sequence generated takes the form "<X> life <Y> box of chocolates <Z>", with the final sentinal token to indicate sequence completion.
    3. The loss is calculated exclusively over the generated sentinal tokens and the missing text.
    4. The decoder reconstructs the entire original sequence autoregressively to compute the total loss over all positions.

    A published solution is not available for this question yet.

    Question 8 MCQ · 3.0 marks

    During the fine-tuning of GPT-1 for Textual Entailment tasks, what is the purpose of the delimiter token ($)?
    1. To trigger the softmax function.
    2. To separate premise and hypothesis.
    3. To mark the end of the sentence.
    4. To initialize the linear head.

    A published solution is not available for this question yet.

    Question 9 MCQ · 3.0 marks

    During the fine-tuning of a decoder-only Transformer for sequence classification, why is the hidden representation of the final input token used as the input to the classification layer?
    1. The first token's representation relies exclusively on absolute positional encodings and lacks semantic context for classification.
    2. The final token's representation is computed independently of self-attention, making its gradients more stable during the fine-tuning process.
    3. Selecting the final token aligns with the autoregressive objective, eliminating the need for teacher forcing during the classification forward pass.
    4. Causal self-attention ensures that only the final token's representation aggregates information from all preceding tokens in the sequence.

    A published solution is not available for this question yet.

    Question 10 MCQ · 3.0 marks

    In the decoding phase of a Language model, temperature scaling with [[IMAGE:f854547ec11886a8_6_4]] is applied to the output logits before performing top- [[IMAGE:f854547ec11886a8_6_5]] (nucleus) sampling. Which of the following options best describes the impact of the temperature scaling on the subsequent sampling step?
    Source diagram or notationSource diagram or notation
    1. Increases the required number of candidate tokens thus expanding the nucleus.
    2. Decreases the required number of candidate tokens thus shrinking the nucleus.
    3. Truncates the tail of the distribution causing the sampling to behave similar to top- [[IMAGE:f854547ec11886a8_6_6]] sampling.
      Source diagram or notation
    4. Size of the nucleus remains unaffected by the temperature scaling as it only depends on order of the tokens.

    A published solution is not available for this question yet.

    Question 11 MCQ · 3.0 marks

    In the WordPiece tokenization algorithm, how does the scoring formula prioritize token merges?
    1. It evaluates the log-probability of the pair and selects the highest scoring transition.
    2. It merges the pair with the highest absolute frequency in the corpus.
    3. It prioritizes pairs that occur together frequently but rarely appear independently in the corpus.
    4. None of these

    A published solution is not available for this question yet.

    Question 12 MCQ · 3.0 marks

    Consider the T5 "text-to-text" paradigm. How does the model process fundamentally different tasks such as Semantic Similarity (continuous regression) and Sentiment Analysis (discrete classification) during the fine-tuning phase?
    1. The model uses the encoder to output the regression score and decoder for the classification tokens switching based on the input prompt.
    2. The decoder part of the model is appended with separate prediction heads for regression and classification which are updated based on Mean Squared Error and Cross-Entropy respectively.
    3. The model treats both tasks as sequence generation objectives with task- specific prefix added in the input and all output is generated as strings.
    4. None of these.

    A published solution is not available for this question yet.

    Question 13 MCQ · 3.0 marks

    In the context of the BART (Bidirectional AutoRegressive Transformer) architecture, how does the model handle corrupted input sequences compared to its output?
    1. The decoder ignores the encoder output and performs standard causal language modeling.
    2. The encoder processes the original sequence and the decoder predicts the corrupted tokens.
    3. The encoder processes the corrupted sequence and the decoder predicts the entire original sequence.
    4. Both the encoder and decoder process the corrupted sequence to calculate masked loss.

    A published solution is not available for this question yet.

    Question 14 MCQ · 3.0 marks

    When comparing the C4 dataset to an 'Unfiltered-C4' dataset that is 8 times larger, what was the observed effect on downstream task performance?
    1. Performance remained identical, suggesting a saturation point in data utility.
    2. Performance improved only for high-resource tasks like translation.
    3. Performance improved significantly due to the increased diversity of data.
    4. Performance degraded across all tasks despite the larger scale.

    A published solution is not available for this question yet.

    Question 15 MCQ · 3.0 marks

    Which of the following describes the 'Adapter Layers' fine-tuning strategy?
    1. Randomly freezing 50% of the attention heads in each layer during training.
    2. Adding small dense-ReLU-dense blocks after FFN layers and updating only those and Layer Norm parameters.
    3. Updating all parameters in the model but using a much smaller learning rate.
    4. Increasing the dimension [[IMAGE:f854547ec11886a8_8_7]] of the existing Feed-Forward networks.
      Source diagram or notation

    A published solution is not available for this question yet.

    Question 16 MCQ · 3.0 marks

    Which of the following is a reported negative effect of including significantly duplicated content in LLM training data?
    1. It decreases the model's overall inference speed during deployment.
    2. It causes the model to generate repeated sequences much more frequently.
    3. It permanently reduces the maximum supported context window length.
    4. It causes the model to lose the ability to process multilingual input.

    A published solution is not available for this question yet.

    Question 17 MCQ · 3.0 marks

    If you have a fixed compute budget ( [[IMAGE:f854547ec11886a8_8_8]] ) and choose to train an extraordinarily large model (massive [[IMAGE:f854547ec11886a8_8_9]] ) on a relatively small dataset (small [[IMAGE:f854547ec11886a8_8_10]] ), what is the likely outcome according to scaling laws?
    Source diagram or notationSource diagram or notationSource diagram or notation
    1. The model will be over-trained and memorize the dataset.
    2. The model will be under-trained and less capable than a smaller model trained on more data with the same compute budget.
    3. The model will perfectly memorize the small dataset and generalize flawlessly.
    4. The model's inference speed will increase dramatically.

    A published solution is not available for this question yet.

    Question 18 MSQ · 3.0 marks

    Consider the following corpus consisting of 4 words (ignore spaces and punctuation). **there is no spoon** The vocabulary is constructed using the whole words in this corpus and the individual characters in this corpus. The vocabulary should now have 13 tokens. Use this vocabulary for the given subquestions.
    Using the vocabulary constructed, two new continuous strings: "prison" and "rhino" are to be segmented. Which of the following proposed segmentations are valid?
    1. prison
    2. p, r, is, o, n
    3. p, r, is, on
    4. r, h, i, n, o
    5. r, h, i, no

    A published solution is not available for this question yet.

    Question 19 MCQ · 3.0 marks

    Consider the following corpus consisting of 4 words (ignore spaces and punctuation). **there is no spoon** The vocabulary is constructed using the whole words in this corpus and the individual characters in this corpus. The vocabulary should now have 13 tokens. Use this vocabulary for the given subquestions.
    Using a unigram language model, identify the segment with the highest probability from the given
    1. n, o, i, s, e, s
    2. no, is, e, s
    3. nois, e, s
    4. no, i, s, e, s

    A published solution is not available for this question yet.