MauryaHub PYQ Practice

da5004_2026T1_Q2_NA.pdf

Large Language Models · Quiz 2 · Jan 2026

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 2 MCQ · 2.0 marks

In the context of Multi-Head Attention, if the model dimension [[IMAGE:a28d0ec52ebbc652_2_2]] and we employ [[IMAGE:a28d0ec52ebbc652_2_3]] heads, what is the dimension of the concatenated output of the 8 heads before it is passed through the final linear output projection [[IMAGE:a28d0ec52ebbc652_2_4]] ?
Source diagram or notationSource diagram or notationSource diagram or notation
  1. [[IMAGE:a28d0ec52ebbc652_2_5]]
    Source diagram or notation
  2. [[IMAGE:a28d0ec52ebbc652_2_6]]
    Source diagram or notation
  3. [[IMAGE:a28d0ec52ebbc652_2_7]]
    Source diagram or notation
  4. [[IMAGE:a28d0ec52ebbc652_2_8]]
    Source diagram or notation

A published solution is not available for this question yet.

Question 3 MCQ · 2.0 marks

In the Masked Language Modeling (MLM) objective of BERT, 15% of tokens are chosen for prediction. Of these chosen tokens, 80% are replaced with [[IMAGE:a28d0ec52ebbc652_2_9]] , 10% are replaced with a random word, and 10% are left unchanged. What is the primary theoretical motivation for **keeping 10% of the tokens unchanged**?
Source diagram or notation
  1. To reduce the computational cost of the softmax layer during training.
  2. To prevent the network from overfitting to the [[IMAGE:a28d0ec52ebbc652_3_10]] token.
    Source diagram or notation
  3. To mitigate the mismatch between pre-training (where [[IMAGE:a28d0ec52ebbc652_3_11]] appears) and fine-tuning (where [[IMAGE:a28d0ec52ebbc652_3_12]] does not appear), ensuring the model creates meaningful representations for non-masked words.
    Source diagram or notationSource diagram or notation
  4. To act as a regularizer similar to Dropout.

A published solution is not available for this question yet.

Question 4 MCQ · 2.0 marks

Why is the standard GPT architecture (Decoder-only with causal masking) generally **unsuitable** for the Masked Language Modeling (MLM) objective as implemented in BERT?
  1. GPT models are too small to learn bidirectional contexts.
  2. The causal mask in GPT prevents the model from attending to future tokens, making it impossible to use right-side context to predict a masked token.
  3. GPT does not have positional embeddings, which are required for MLM.
  4. GPT uses ReLU activation, while BERT uses GELU, which is required for MLM.

A published solution is not available for this question yet.

Question 5 MCQ · 2.0 marks

In the T5 (Text-to-Text Transfer Transformer) framework, every NLP task is cast as a text generation problem. If you use T5 for a **Semantic Textual Similarity (STS-B)** task, where the goal is to predict a similarity score (e.g., 3.8) between two sentences, how does the model output this score?
  1. It outputs a single scalar value from a regression head on top of the encoder.
  2. It generates the string "3.8" token-by-token using the decoder.
  3. It outputs a class label corresponding to a bucketed score range (e.g., "High Similarity").
  4. T5 cannot be used for regression tasks like STS-B.

A published solution is not available for this question yet.

Question 6 MCQ · 2.0 marks

When using **Temperature Sampling** to mix datasets from different tasks during multi-task pre- training, let [[IMAGE:a28d0ec52ebbc652_3_13]] be the probability of sampling a task [[IMAGE:a28d0ec52ebbc652_3_14]] . The formula involves raising the proportion to the power of [[IMAGE:a28d0ec52ebbc652_4_15]] . If we set the temperature [[IMAGE:a28d0ec52ebbc652_4_16]] (a very large value), what happens to the sampling distribution across tasks?
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
  1. The model trains almost exclusively on the largest dataset (highest resource task).
  2. The sampling distribution approaches a uniform distribution, where all tasks (large and small) are sampled with nearly equal probability.
  3. The model trains almost exclusively on the smallest dataset (lowest resource task).
  4. The sampling distribution remains proportional to the original dataset sizes.

A published solution is not available for this question yet.

Question 7 MCQ · 2.0 marks

When constructing training datasets for large language models (LLMs), which of the following best describes the key factors that must be balanced to achieve strong and reliable performance?
  1. Model depth, number of parameters, and learning rate
  2. Scale, diversity, and quality of the training data
  3. Vocabulary size, tokenization method, and batch size
  4. Compute budget, optimizer choice, and hardware efficiency

A published solution is not available for this question yet.

Question 8 MSQ · 3.0 marks

Which of the following are components found within a standard Transformer **Encoder** layer?
  1. Multi-Head Self-Attention mechanism
  2. Position-wise Feed-Forward Networks
  3. Cross-Attention mechanism (Encoder-Decoder attention)
  4. Masked Multi-Head Self-Attention

A published solution is not available for this question yet.

Question 9 MSQ · 3.0 marks

Select all correct statements regarding the comparison between RNNs and Transformers.
  1. In Transformers, the path length between any two positions in the sequence is constant [[IMAGE:a28d0ec52ebbc652_5_17]] , whereas in RNNs it is [[IMAGE:a28d0ec52ebbc652_5_18]] (where [[IMAGE:a28d0ec52ebbc652_5_19]] is the sequence length).
    Source diagram or notationSource diagram or notationSource diagram or notation
  2. Transformers process tokens strictly sequentially during training in the same way as RNNs.
  3. Transformers allow for significantly more parallelization during training compared to RNNs.
  4. Attention mechanisms in Transformers utilize Query, Key, and Value vectors derived from input embeddings.

A published solution is not available for this question yet.

Question 10 MSQ · 3.0 marks

Select all correct findings from the T5 paper regarding **Unsupervised Pre-training Objectives**.
  1. The specific corruption rate (e.g., 10%, 15%, 25%) had a minimal effect on downstream performance.
  2. Using a span length of around 3 tokens performed slightly better than masking single tokens.
  3. The "Deshuffling" objective (reordering shuffled sentences) significantly outperformed the Span Corruption objective.
  4. Replacing a corrupted span with a unique sentinel token worked better than simply dropping the tokens from the input.

A published solution is not available for this question yet.

Question 11 MSQ · 3.0 marks

Consider the concept of **Zero-Shot Transfer** as popularized by GPT-2. Why might this be preferred over Supervised Fine-Tuning?
  1. It allows the model to handle tasks for which no labeled training data is available.
  2. It always achieves higher accuracy than a fine-tuned SOTA model.
  3. It avoids the need to store a separate specialized model (checkpoint) for every downstream task.
  4. It mimics the human ability to perform tasks based on instructions without needing thousands of examples.

A published solution is not available for this question yet.

Question 12 NAT · 2.0 marks

Let [[IMAGE:a28d0ec52ebbc652_6_20]] denote the activation of the [[IMAGE:a28d0ec52ebbc652_6_21]] neuron for the [[IMAGE:a28d0ec52ebbc652_6_22]] training sample. In **Layer Normalization**, normalization is performed **across features for each individual** **sample**. Let [[IMAGE:a28d0ec52ebbc652_6_23]] be the number of neurons (features) in the hidden layer. For a given sample, the layer normalization statistics are computed as: [[IMAGE:a28d0ec52ebbc652_6_24]] [[IMAGE:a28d0ec52ebbc652_6_25]] [[IMAGE:a28d0ec52ebbc652_6_26]] where [[IMAGE:a28d0ec52ebbc652_6_27]] is a small constant for numerical stability. [[IMAGE:a28d0ec52ebbc652_6_28]] Layer Normalization is applied **independently to each sample** in the given data above. For **Sample 2**, after applying Layer Normalization. **What is the maximum value in the** **normalized output vector** [[IMAGE:a28d0ec52ebbc652_6_29]] **?** (ignore the learnable parameters [[IMAGE:a28d0ec52ebbc652_6_30]] and [[IMAGE:a28d0ec52ebbc652_6_31]] , and assume [[IMAGE:a28d0ec52ebbc652_6_32]] ):
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

    A published solution is not available for this question yet.

    Question 13 MCQ · 2.0 marks

    Consider a GPT model used for **Causal Language Modelling**. We feed the input sentence: [[IMAGE:a28d0ec52ebbc652_7_33]] Assume the context length is 7 (time steps t = 0 to t = 6). The attention matrix computed in one attention layer is given below: [[IMAGE:a28d0ec52ebbc652_7_34]] Based on the above data, answer the given subquestions.
    Is the given attention matrix valid for a causal language modelling task?
    Source diagram or notationSource diagram or notation
    1. True
    2. False
    3. Insufficient information

    A published solution is not available for this question yet.

    Question 14 NAT · 2.0 marks

    Consider a GPT model used for **Causal Language Modelling**. We feed the input sentence: [[IMAGE:a28d0ec52ebbc652_7_33]] Assume the context length is 7 (time steps t = 0 to t = 6). The attention matrix computed in one attention layer is given below: [[IMAGE:a28d0ec52ebbc652_7_34]] Based on the above data, answer the given subquestions.
    At time step t = 5 (word: data), what is the attention weight assigned to the word science? (Provide exact answer)
    Source diagram or notationSource diagram or notation

      A published solution is not available for this question yet.

      Question 15 MCQ · 2.0 marks

      **Sentence Corpus:** "the key to artificial intelligence has always been the representation" **Pre-processing Instructions:** • Convert all text to lowercase. • Append the end-of-word symbol [[IMAGE:a28d0ec52ebbc652_8_35]] to every word (e.g., "the" becomes [[IMAGE:a28d0ec52ebbc652_8_36]] ). • The **Initial Vocabulary** is defined as the set of all unique characters present in the corpus plus the [[IMAGE:a28d0ec52ebbc652_8_37]] symbol. Based on the above data, answer the given subquestions.
      Which of the following character pairs occurs with the **highest frequency** across the entire sentence before any BPE merges are performed?
      Source diagram or notationSource diagram or notationSource diagram or notation
      1. [[IMAGE:a28d0ec52ebbc652_8_38]]
        Source diagram or notation
      2. [[IMAGE:a28d0ec52ebbc652_8_39]]
        Source diagram or notation
      3. [[IMAGE:a28d0ec52ebbc652_9_40]]
        Source diagram or notation
      4. [[IMAGE:a28d0ec52ebbc652_9_41]]
        Source diagram or notation

      A published solution is not available for this question yet.

      Question 16 NAT · 2.0 marks

      **Sentence Corpus:** "the key to artificial intelligence has always been the representation" **Pre-processing Instructions:** • Convert all text to lowercase. • Append the end-of-word symbol [[IMAGE:a28d0ec52ebbc652_8_35]] to every word (e.g., "the" becomes [[IMAGE:a28d0ec52ebbc652_8_36]] ). • The **Initial Vocabulary** is defined as the set of all unique characters present in the corpus plus the [[IMAGE:a28d0ec52ebbc652_8_37]] symbol. Based on the above data, answer the given subquestions.
      Calculate the **total vocabulary size** immediately after the first merge operation is completed. **Note:** The vocabulary size includes all individual base characters/symbols plus the newly created merge token.
      Source diagram or notationSource diagram or notationSource diagram or notation

        A published solution is not available for this question yet.

        Question 17 NAT · 3.0 marks

        Consider a vocabulary [[IMAGE:a28d0ec52ebbc652_9_42]] . At a specific time step [[IMAGE:a28d0ec52ebbc652_9_43]] , the model outputs the following probability distribution: [[IMAGE:a28d0ec52ebbc652_9_44]] If we use **Top-** [[IMAGE:a28d0ec52ebbc652_9_45]] **sampling with** [[IMAGE:a28d0ec52ebbc652_9_46]] , what is the probability of selecting token **B**? (Calculate the re-normalized probability. Enter the value correct to 2 decimal places).
        Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

          A published solution is not available for this question yet.

          Question 18 NAT · 3.0 marks

          Consider a small BERT-like model with the following configuration: • Embedding dimension ( [[IMAGE:a28d0ec52ebbc652_10_47]] ) = 128 • Vocabulary size ( [[IMAGE:a28d0ec52ebbc652_10_48]] ) = 5000 • Maximum sequence length ( [[IMAGE:a28d0ec52ebbc652_10_49]] ) = 64 • Number of segment types = 2 Calculate the **total number of parameters** in the **Embedding Layer** (sum of Token embeddings, Position embeddings, and Segment embeddings). Enter the exact integer value.
          Source diagram or notationSource diagram or notationSource diagram or notation

            A published solution is not available for this question yet.

            Question 19 MCQ · 3.0 marks

            What is a major disadvantage of character-level tokenization?
            1. Character-level tokenization cannot represent punctuation marks or special symbols
            2. It has very large vocabulary size because each word is broken into multiple characters
            3. It produces much longer input sequences which increases computational cost.
            4. It fails to capture morphological patterns such as prefixes and suffixes, reducing the model's ability to understand word structure

            A published solution is not available for this question yet.

            Question 20 MCQ · 3.0 marks

            When using the SentencePiece, how are subword units selected?
            1. Based on a probabilistic model that maximizes the likelihood of the training data.
            2. By randomly selecting character n-grams until the vocabulary limit is reached.
            3. By selecting only the top 50,000 most frequent words in the corpus.
            4. By iteratively merging the most frequent pair of adjacent characters.

            A published solution is not available for this question yet.

            Question 21 MSQ · 2.0 marks

            To obtain high-quality training text from raw web data for large language models, which of the following mechanisms are typically included in the pre-processing pipeline?
            1. Tokenizing text into subword units using Byte Pair Encoding (BPE)
            2. Deduplicating content at line, paragraph, and document levels
            3. Detecting and filtering toxic content such as hate speech and profanity
            4. Identifying the language of web pages
            5. Assessing the quality of content to remove low-value or spam text
            6. Detecting and removing Personally Identifiable Information (PII)
            7. Fine-tuning the model using reinforcement learning from human feedback (RLHF)

            A published solution is not available for this question yet.

            Question 22 MSQ · 2.0 marks

            Which of the following statements correctly explain the importance of deduplication during preprocessing of large-scale datasets used for training Deep Learning or Large Language Models?
            1. Deduplication reduces the risk of overfitting by preventing repeated samples from dominating the gradient updates.
            2. Deduplication guarantees that the trained model will achieve higher accuracy on all downstream tasks.
            3. Deduplication helps avoid data leakage between training and evaluation sets, leading to more reliable performance metrics.
            4. Deduplication eliminates the need for regularization techniques such as dropout and weight decay.

            A published solution is not available for this question yet.