MauryaHub PYQ Practice

da5013_2024T3_Q1_NA.pdf

Deep Learning Practice · Quiz 1 · Sep 2024

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 293 NAT · 3.0 marks

[[IMAGE:7a74ca87f5851ef0_1_0]]
Source diagram or notation

    A published solution is not available for this question yet.

    Question 295 MCQ · 4.0 marks

    [[IMAGE:7a74ca87f5851ef0_2_1]]
    Source diagram or notation
    1. Pre-Tokenization
    2. Normalization
    3. Post Processor
    4. Tokenization Algorithm
    5. Decoder

    A published solution is not available for this question yet.

    Question 296 MCQ · 3.0 marks

    Choose the Hugging Face module that helps us train a tokenizer from scratch on a specific dataset.
    1. tokenizers
    2. transformers
    3. evaluate
    4. Autotrain
    5. Accelerate

    A published solution is not available for this question yet.

    Question 297 MCQ · 3.0 marks

    A dataset contains 10 billion words ( separated by a single white space). Suppose we use a pre- trained tokenizer that has a vocabulary of size 10,000 to tokenize the dataset, then the number of tokens in the dataset will always be greater than or equal to the number of words in the dataset. The statement is
    1. True
    2. False

    A published solution is not available for this question yet.

    Question 298 MCQ · 3.0 marks

    Which of the following tokenization algorithms can be applied to languages that do not have any word delimiters?
    1. BPE (Byte Pair Encoding)
    2. Wordpiece
    3. Sentencepiece

    A published solution is not available for this question yet.

    Question 299 MCQ · 3.0 marks

    Consider the Wikipedia dataset scraped from the web that contains 2 billion words. A team decided to use the BPE tokenization algorithm with the varying vocabulary size from 2K to 52K, then the statement that increasing the number of merges will increase the size of the vocabulary is
    1. True
    2. False
    3. Insufficient information

    A published solution is not available for this question yet.

    Question 300 MCQ · 3.0 marks

    Suppose that we pre-train a Causal Language Model. Choose the data collator function from the Hugging Face library that is suitable for this task
    1. DataCollator(tokenizer)
    2. DefaultDataCollator(tokenizer)
    3. DataCollatorForLanguageModelling(tokenizer,mlm=False)
    4. DataCollatorForCausalLanguageModelling(tokenizer)
    5. DataLoader(tokenizer)

    A published solution is not available for this question yet.

    Question 301 MSQ · 4.0 marks

    [[IMAGE:7a74ca87f5851ef0_4_2]]
    Source diagram or notation
    1. [[IMAGE:7a74ca87f5851ef0_5_3]]
      Source diagram or notation
    2. [[IMAGE:7a74ca87f5851ef0_5_4]]
      Source diagram or notation
    3. [[IMAGE:7a74ca87f5851ef0_5_5]]
      Source diagram or notation
    4. [[IMAGE:7a74ca87f5851ef0_5_6]]
      Source diagram or notation
    5. [[IMAGE:7a74ca87f5851ef0_5_7]]
      Source diagram or notation
    6. [[IMAGE:7a74ca87f5851ef0_5_8]]
      Source diagram or notation

    A published solution is not available for this question yet.

    Question 302 MSQ · 4.0 marks

    [[IMAGE:7a74ca87f5851ef0_6_9]]
    Source diagram or notation
    1. [[IMAGE:7a74ca87f5851ef0_6_10]]
      Source diagram or notation
    2. [[IMAGE:7a74ca87f5851ef0_6_11]]
      Source diagram or notation
    3. [[IMAGE:7a74ca87f5851ef0_6_12]]
      Source diagram or notation
    4. [[IMAGE:7a74ca87f5851ef0_6_13]]
      Source diagram or notation

    A published solution is not available for this question yet.

    Question 303 MSQ · 3.0 marks

    [[IMAGE:7a74ca87f5851ef0_7_14]]
    Source diagram or notation
    1. [[IMAGE:7a74ca87f5851ef0_7_15]]
      Source diagram or notation
    2. [[IMAGE:7a74ca87f5851ef0_7_16]]
      Source diagram or notation
    3. [[IMAGE:7a74ca87f5851ef0_7_17]]
      Source diagram or notation

    A published solution is not available for this question yet.

    Question 304 MSQ · 3.0 marks

    [[IMAGE:7a74ca87f5851ef0_7_18]]
    Source diagram or notation
    1. ids
    2. tokens
    3. offsets
    4. attention_mask
    5. special_token_mask
    6. type_ids
    7. vocab_size

    A published solution is not available for this question yet.

    Question 305 NAT · 3.0 marks

    [[IMAGE:7a74ca87f5851ef0_8_19]] Based on the above data, answer the given subquestions.
    Enter the number of parameters in the embedding layer of the model in millions. For example, if the answer is 1234567. Then enter it as 1.23
    Source diagram or notation

      A published solution is not available for this question yet.

      Question 306 NAT · 3.0 marks

      [[IMAGE:7a74ca87f5851ef0_8_19]] Based on the above data, answer the given subquestions.
      Enter the context length.
      Source diagram or notation

        A published solution is not available for this question yet.

        Question 307 NAT · 5.0 marks

        [[IMAGE:7a74ca87f5851ef0_9_20]] Based on the above data, answer the given subquestions.
        Enter the number of tokens (in millions) processed by the model after 1000 steps. Enter the answer to 2 decimal places. For example, if your answer is 123456789, then enter it as 123.45.
        Source diagram or notation

          A published solution is not available for this question yet.

          Question 308 NAT · 3.0 marks

          [[IMAGE:7a74ca87f5851ef0_9_20]] Based on the above data, answer the given subquestions.
          How many steps does it take to complete one epoch of training? Enter the answer in thousands (round down to an integer). For example, if your answer is 1234567.89, then enter it as 1234567.
          Source diagram or notation

            A published solution is not available for this question yet.

            Question 309 NAT · 3.0 marks

            [[IMAGE:7a74ca87f5851ef0_10_21]]
            Source diagram or notation

              A published solution is not available for this question yet.