MauryaHub PYQ Practice

da5004_2025T3_Q2_NA.pdf

Large Language Models · Quiz 2 · Sep 2025

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 105 NAT · 3.0 marks

Suppose for a 2-component GMM, at some iteration you have [[IMAGE:f8bb8de26f2feee9_1_0]] What is the responsibility [[IMAGE:f8bb8de26f2feee9_1_1]] ?
Source diagram or notationSource diagram or notation

    A published solution is not available for this question yet.

    Question 107 NAT · 2.0 marks

    [[IMAGE:f8bb8de26f2feee9_2_2]] Based on the above data, answer the given subquestions.
    [[IMAGE:f8bb8de26f2feee9_3_3]]
    Source diagram or notationSource diagram or notation

      A published solution is not available for this question yet.

      Question 108 NAT · 2.0 marks

      [[IMAGE:f8bb8de26f2feee9_2_2]] Based on the above data, answer the given subquestions.
      [[IMAGE:f8bb8de26f2feee9_3_4]]
      Source diagram or notationSource diagram or notation

        A published solution is not available for this question yet.

        Question 109 MCQ · 3.0 marks

        [[IMAGE:f8bb8de26f2feee9_2_2]] Based on the above data, answer the given subquestions.
        Which token will be generated in the first merge?
        Source diagram or notation
        1. [[IMAGE:f8bb8de26f2feee9_3_5]]
          Source diagram or notation
        2. [[IMAGE:f8bb8de26f2feee9_3_6]]
          Source diagram or notation
        3. [[IMAGE:f8bb8de26f2feee9_3_7]]
          Source diagram or notation
        4. [[IMAGE:f8bb8de26f2feee9_4_8]]
          Source diagram or notation

        A published solution is not available for this question yet.

        Question 110 NAT · 2.0 marks

        A study is comparing three architectures, all trained with the same compute budget: • **Model A (Encoder-Decoder):** A 12-layer encoder and a 12-layer decoder with separate parameters. Its total size is 220 million parameters. • **Model B (Decoder-Only):** A 24-layer decoder-only model. Its total size is also approximately 220 million parameters. • **Model C (Shared Enc-Dec):** A 12-layer encoder and a 12-layer decoder where the parameters are shared ( [[IMAGE:f8bb8de26f2feee9_4_9]] ). Assuming that the 12-layer encoder and 12-layer decoder in Model A have an equal number of parameters, what is the approximate total parameter count (in millions) of Model C?
        Source diagram or notation

          A published solution is not available for this question yet.

          Question 111 NAT · 2.0 marks

          Consider a text corpus being processed using the Byte Pair Encoding (BPE) algorithm. The initial vocabulary contains [[IMAGE:f8bb8de26f2feee9_4_10]] tokens (representing characters), and the corpus contains a total of [[IMAGE:f8bb8de26f2feee9_4_11]] tokens. During BPE training, the following merge operations are performed. Each merge combines all occurrences of a particular token pair into a single new token, reducing the total token count accordingly: 1. **Merge 1:** Pair ('t', 'h') appears 24 times → merged into new token 'th' 2. **Merge 2:** Pair ('i', 'n') appears 18 times → merged into new token 'in' 3. **Merge 3:** Pair ('th', 'e') appears 10 times → merged into new token 'the' 4. **Merge 4:** Pair ('i', 's') appears 8 times → merged into new token 'is' What is the final total number of tokens in the corpus after all 4 merges are completed?
          Source diagram or notationSource diagram or notation

            A published solution is not available for this question yet.

            Question 112 NAT · 2.0 marks

            If Byte-Pair Encoding (BPE) starts with [[IMAGE:f8bb8de26f2feee9_5_12]] unique characters and performs [[IMAGE:f8bb8de26f2feee9_5_13]] merge operations, what will be the final vocabulary size?
            Source diagram or notationSource diagram or notation

              A published solution is not available for this question yet.

              Question 113 NAT · 2.0 marks

              In a 4-token sequence, a Transformer is calculating the attention for the third Token. Its Query vector ( [[IMAGE:f8bb8de26f2feee9_6_14]] ) must be compared against the Key vectors ( [[IMAGE:f8bb8de26f2feee9_6_15]] ) of all tokens. The Causal LM mask is applied to the raw scores before the softmax function. The raw dot-product scores are: • [[IMAGE:f8bb8de26f2feee9_6_16]] • [[IMAGE:f8bb8de26f2feee9_6_17]] • [[IMAGE:f8bb8de26f2feee9_6_18]] • [[IMAGE:f8bb8de26f2feee9_6_19]] After the Causal LM mask and the subsequent softmax function are applied, what will be the final attention probability assigned to Token 4?
              Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

                A published solution is not available for this question yet.

                Question 114 MSQ · 3.0 marks

                Which of the following is/are **primarily** multilingual datasets? Select all that apply.
                1. Sangraha
                2. DOLMA
                3. BookCorpus
                4. CommonCrawl

                A published solution is not available for this question yet.

                Question 115 MSQ · 2.0 marks

                Which decoding method ensures the same response every time for a given prompt (assuming no randomness in model weights)? Select all that apply.
                1. Top-k with k = 3, temperature > 1
                2. Top-k with k = 3, temperature = 1
                3. Top-p with p = 0.2, Temperature > 1
                4. Top-p with p = 0.2, Temperature = 1
                5. Beam Search with beam-size = 5
                6. Beam Search with beam-size = 1

                A published solution is not available for this question yet.

                Question 116 MCQ · 2.0 marks

                A data preprocessing pipeline removed data from the Common Crawl snapshot in three stages: 50% after language identification, 24% of the remaining after quality filtering, and 12% of what remained after deduplication. What percentage of the original Common Crawl snapshot remains in the final dataset? (Choose the closest number)
                1. 21%
                2. 66%
                3. 33%
                4. 50%

                A published solution is not available for this question yet.

                Question 117 MCQ · 3.0 marks

                A data preprocessing pipeline needs to apply the following steps on raw Common Crawl data: 1. Fuzzy deduplication (using MinHash) 2. Toxicity detection (using ML-based classifiers) 3. PII removal (using regex patterns) 4. Exact deduplication (using hash-based matching) 5. Quality filtering (using several rule-based heuristics) 6. URL Filtering What is the most efficient sequence to apply these steps?
                1. 6 [[IMAGE:f8bb8de26f2feee9_8_20]] 4 [[IMAGE:f8bb8de26f2feee9_8_21]] 1 [[IMAGE:f8bb8de26f2feee9_8_22]] 5 [[IMAGE:f8bb8de26f2feee9_8_23]] 2 [[IMAGE:f8bb8de26f2feee9_8_24]] 3
                  Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
                2. 1 [[IMAGE:f8bb8de26f2feee9_8_25]] 4 [[IMAGE:f8bb8de26f2feee9_8_26]] 2 [[IMAGE:f8bb8de26f2feee9_8_27]] 5 [[IMAGE:f8bb8de26f2feee9_8_28]] 3 [[IMAGE:f8bb8de26f2feee9_8_29]] 6
                  Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
                3. 5 [[IMAGE:f8bb8de26f2feee9_8_30]] 6 [[IMAGE:f8bb8de26f2feee9_8_31]] 4 [[IMAGE:f8bb8de26f2feee9_8_32]] 1 [[IMAGE:f8bb8de26f2feee9_8_33]] 2 [[IMAGE:f8bb8de26f2feee9_8_34]] 3
                  Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
                4. 4 [[IMAGE:f8bb8de26f2feee9_8_35]] 5 [[IMAGE:f8bb8de26f2feee9_8_36]] 6 [[IMAGE:f8bb8de26f2feee9_8_37]] 1 [[IMAGE:f8bb8de26f2feee9_8_38]] 3 [[IMAGE:f8bb8de26f2feee9_8_39]] 2
                  Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

                A published solution is not available for this question yet.

                Question 118 MCQ · 4.0 marks

                [[IMAGE:f8bb8de26f2feee9_8_40]]
                Source diagram or notation
                1. A = BART, B = T5, C = BERT, D = GPT
                2. A = T5, B = BERT, C = GPT, D = BART
                3. A = BERT, B = GPT, C = BART, D = T5
                4. A = BERT, B = BART, C = GPT, D = T5

                A published solution is not available for this question yet.

                Question 119 MCQ · 2.0 marks

                Consider the scaling law: [[IMAGE:f8bb8de26f2feee9_9_41]] where [[IMAGE:f8bb8de26f2feee9_9_42]] = model parameters, [[IMAGE:f8bb8de26f2feee9_9_43]] = training tokens, [[IMAGE:f8bb8de26f2feee9_9_44]] = scale constants, [[IMAGE:f8bb8de26f2feee9_9_45]] = scaling exponents. A team trains an LLM and observes the following losses for their model: 1. Training Loss 2. Test Loss (same distribution as training) 3. Test Loss (different distribution from training) Which loss does [[IMAGE:f8bb8de26f2feee9_9_46]] represent?
                Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
                1. Both 2 and 3
                2. Both 1 and 2
                3. 1 only
                4. 2 only

                A published solution is not available for this question yet.

                Question 120 MCQ · 3.0 marks

                A company is building a sentiment analysis system for customer reviews. They have two
                1. BART is expected to provide significantly better accuracy and is worth the extra computational cost since it's pre-trained on more diverse corruption tasks.
                2. BART is better because its decoder can verify the encoder's sentiment understanding, providing a built-in validation mechanism that improves robustness.
                3. They should use BART's encoder only and discard the decoder, which will give them the same efficiency as BERT with better pre-training.
                4. BERT is more suitable because sentiment classification only requires understanding (encoding), not generation (decoding), making it faster and more memory-efficient without sacrificing accuracy.

                A published solution is not available for this question yet.

                Question 121 MCQ · 2.0 marks

                Which of the following statements best describes the fundamental difference between the unsupervised pre-training objective used in T5 (span corruption) and its supervised fine-tuning objectives such as translation or summarization?
                1. The pre-training objective is a denoising task, while the fine-tuning objective is a Causal Language Model (CLM) task.
                2. There is no fundamental difference; both are treated as "text-to-text" tasks, where the model is trained to generate a target sequence given an input sequence.
                3. The pre-training objective trains the model to fill in [MASK] tokens, while the fine-tuning objective trains the model to generate text from a [CLS] token.
                4. The pre-training objective trains the entire encoder-decoder, while the fine- tuning objective freezes the encoder and only trains the decoder.

                A published solution is not available for this question yet.

                Question 122 MCQ · 2.0 marks

                [[IMAGE:f8bb8de26f2feee9_11_47]]
                Source diagram or notation
                1. [[IMAGE:f8bb8de26f2feee9_11_48]]
                  Source diagram or notation
                2. [[IMAGE:f8bb8de26f2feee9_11_49]]
                  Source diagram or notation
                3. [[IMAGE:f8bb8de26f2feee9_11_50]]
                  Source diagram or notation
                4. [[IMAGE:f8bb8de26f2feee9_11_51]]
                  Source diagram or notation

                A published solution is not available for this question yet.

                Question 123 MCQ · 2.0 marks

                A Transformer model is designed to handle both understanding and generation tasks. When given the input sequence [[IMAGE:f8bb8de26f2feee9_11_52]] , its attention mask allows the following: • When processing [[IMAGE:f8bb8de26f2feee9_11_53]] or [[IMAGE:f8bb8de26f2feee9_11_54]] , the model can see both [[IMAGE:f8bb8de26f2feee9_11_55]] and [[IMAGE:f8bb8de26f2feee9_11_56]] . • When processing [[IMAGE:f8bb8de26f2feee9_11_57]] , the model can see [[IMAGE:f8bb8de26f2feee9_11_58]] and [[IMAGE:f8bb8de26f2feee9_11_59]] but not [[IMAGE:f8bb8de26f2feee9_11_60]] . This attention scheme is the defining characteristic of which training objective?
                Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
                1. Causal Language Modeling (CLM)
                2. Masked Language Modeling (MLM)
                3. Denoising Objective (like BART)
                4. Prefix Language Modeling (Prefix-LM)

                A published solution is not available for this question yet.

                Question 124 MCQ · 2.0 marks

                BART is an encoder-decoder model. During its denoising pre-training, what is fed as input to the decoder?
                1. The corrupted text from the encoder.
                2. The original, uncorrupted text (shifted right by one token).
                3. A sequence of [MASK] tokens, one for each original word.
                4. Nothing; the decoder starts with only the < s > token.

                A published solution is not available for this question yet.

                Question 125 MCQ · 2.0 marks

                [[IMAGE:f8bb8de26f2feee9_12_61]]
                Source diagram or notation
                1. [[IMAGE:f8bb8de26f2feee9_12_62]]
                  Source diagram or notation
                2. [[IMAGE:f8bb8de26f2feee9_12_63]]
                  Source diagram or notation
                3. [[IMAGE:f8bb8de26f2feee9_12_64]]
                  Source diagram or notation
                4. [[IMAGE:f8bb8de26f2feee9_12_65]]
                  Source diagram or notation

                A published solution is not available for this question yet.

                Question 126 MCQ · 2.0 marks

                How does SentencePiece handle whitespace characters (like spaces) during tokenization?
                1. It only keeps whitespace that appears next to punctuation marks.
                2. It discards all whitespace as a pre-processing step.
                3. It replaces all whitespace with a special [SPACE] token.
                4. It treats whitespace as part of the text stream and encodes it, often using a symbol like [[IMAGE:f8bb8de26f2feee9_13_66]] (underscore).
                  Source diagram or notation

                A published solution is not available for this question yet.

                Question 127 MCQ · 2.0 marks

                For a layer with [[IMAGE:f8bb8de26f2feee9_13_67]] features, how many learnable parameters do BatchNorm and LayerNorm have?
                Source diagram or notation
                1. BatchNorm: 512, LayerNorm: 512
                2. BatchNorm: 1024, LayerNorm: 512
                3. BatchNorm: 512, LayerNorm: 1024
                4. BatchNorm: 1024, LayerNorm: 1024

                A published solution is not available for this question yet.

                Question 128 MCQ · 2.0 marks

                In a standard Encoder-Decoder Transformer, where do the Query [[IMAGE:f8bb8de26f2feee9_13_68]] , Key [[IMAGE:f8bb8de26f2feee9_13_69]] , and Value [[IMAGE:f8bb8de26f2feee9_13_70]] inputs for the Decoder's Cross-Attention sub-layer come from?
                Source diagram or notationSource diagram or notationSource diagram or notation
                1. [[IMAGE:f8bb8de26f2feee9_13_71]] comes from the decoder's previous sub-layer; [[IMAGE:f8bb8de26f2feee9_13_72]] and [[IMAGE:f8bb8de26f2feee9_13_73]] come from the final hidden states of the encoder.
                  Source diagram or notationSource diagram or notationSource diagram or notation
                2. [[IMAGE:f8bb8de26f2feee9_13_74]] comes from the final hidden states of the encoder; [[IMAGE:f8bb8de26f2feee9_13_75]] and [[IMAGE:f8bb8de26f2feee9_13_76]] come from the decoder's previous sub-layer.
                  Source diagram or notationSource diagram or notationSource diagram or notation

                A published solution is not available for this question yet.