MauryaHub PYQ Practice

da5007_2025T3_Q1_NA.pdf

Reinforcement Learning · Quiz 1 · Sep 2025

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 235 SHORT_TEXT · 1.5 marks

Enter the correct answer for Blank (b) ________________ **NOTE:** Enter the exact answer without any extra space in the beginning or at the end.

    A published solution is not available for this question yet.

    Question 237 MCQ · 3.0 marks

    [[IMAGE:8e32aab7e59b4517_2_0]]
    Source diagram or notation
    1. [[IMAGE:8e32aab7e59b4517_2_1]]
      Source diagram or notation
    2. [[IMAGE:8e32aab7e59b4517_2_2]]
      Source diagram or notation
    3. [[IMAGE:8e32aab7e59b4517_2_3]]
      Source diagram or notation
    4. [[IMAGE:8e32aab7e59b4517_2_4]]
      Source diagram or notation

    A published solution is not available for this question yet.

    Question 238 MCQ · 3.0 marks

    Consider following assertion reason pair: **Assertion:** In the UCB algorithm for multi-armed bandits, replacing the upper confidence bound with a lower confidence bound (and greedily selecting actions based on that lower bound) would still promote effective exploration or lead to optimal reward maximisation. **Reason:** Even when lower confidence bounds are used instead of upper bounds, if the algorithm greedily selects arms based on these lower bounds, it would still encourage exploration and ultimately achieve optimal reward maximisation.
    1. Both Assertion and Reason are true, and Reason is the correct explanation of Assertion.
    2. Both Assertion and Reason are true, but Reason is NOT the correct explanation of Assertion.
    3. Assertion is true, Reason is false
    4. Assertion is false, Reason is false

    A published solution is not available for this question yet.

    Question 239 MCQ · 5.0 marks

    [[IMAGE:8e32aab7e59b4517_3_5]]
    Source diagram or notation
    1. Policy Iteration
    2. Value Iteration
    3. Both require equal updates
    4. It cannot be determined

    A published solution is not available for this question yet.

    Question 240 MCQ · 2.0 marks

    [[IMAGE:8e32aab7e59b4517_4_6]]
    Source diagram or notation
    1. Arm 1
    2. Arm 2
    3. Arm 3
    4. Arm 4

    A published solution is not available for this question yet.

    Question 241 MSQ · 5.0 marks

    Which of the following statements correctly describes the differences in computational complexity and convergence behaviour between Policy Iteration and Value Iteration algorithms in solving Markov Decision Processes?
    1. Value Iteration requires fewer iterations to converge, but each iteration is computationally more expensive due to the policy evaluation step.
    2. Value Iteration performs a combined update of value estimation and policy improvement in one step and requires more iterations to converge compared to Policy Iteration.
    3. Policy Iteration can only be applied to small state spaces, while Value Iteration scales well to large state spaces.
    4. Value Iteration is simpler to implement because it only maintains a value function, while Policy Iteration maintains both policy and value function.
    5. Both Policy Iteration and Value Iteration are guaranteed to converge to the optimal policy.

    A published solution is not available for this question yet.

    Question 242 NAT · 5.0 marks

    [[IMAGE:8e32aab7e59b4517_5_7]]
    Source diagram or notation

      A published solution is not available for this question yet.

      Question 243 NAT · 4.0 marks

      [[IMAGE:8e32aab7e59b4517_6_8]]
      Source diagram or notation

        A published solution is not available for this question yet.

        Question 244 MCQ · 1.0 marks

        Suppose you face a 2-armed bandit task where, at each time step, the true action values are either (10, 20) with probability 0.7 (case A) or (90, 80) with probability 0.3 (case B). Based on the above data, answer the given subquestions.
        If you cannot observe which case you face, what is the best expected reward per step you can achieve, and what strategy should you follow?
        1. Randomly choose between Action 1 and Action 2; expected reward is 3.
        2. Choose Action 1 always; expected reward is 34.
        3. Choose Action 2 always; expected reward is 38.
        4. Alternate between Action 1 and Action 2; expected reward is 40.

        A published solution is not available for this question yet.

        Question 245 MCQ · 1.0 marks

        Suppose you face a 2-armed bandit task where, at each time step, the true action values are either (10, 20) with probability 0.7 (case A) or (90, 80) with probability 0.3 (case B). Based on the above data, answer the given subquestions.
        If, at each time step, you are told whether you face case A or case B (but not the true action values), what is the best expected reward per step you can achieve,and what strategy should you follow?
        1. Pick Action 2 in case A, Action 1 in case B; expected reward is 41.
        2. Always pick Action 2; expected reward is 38.
        3. Always pick Action 1; expected reward is 34.
        4. The expected reward stated in all the given options is incorrect.

        A published solution is not available for this question yet.

        Question 246 MCQ · 2.0 marks

        [[IMAGE:8e32aab7e59b4517_8_9]] Based on the above data, answer the given subquestions.
        Which of the following best explains why a contextual bandit approach is preferred over a standard multi-armed bandit in this scenario?
        Source diagram or notation
        1. Because the optimal waiting time is the same for all VMs, regardless of their context.
        2. Because contextual bandits can adaptively select actions based on the specific features of each VM failure event, maximising expected reward across diverse situations.
        3. Because multi-armed bandits are unable to balance exploration and exploitation.
        4. Because contextual bandits always guarantee zero regret after every round.

        A published solution is not available for this question yet.

        Question 247 MCQ · 4.0 marks

        [[IMAGE:8e32aab7e59b4517_8_9]] Based on the above data, answer the given subquestions.
        [[IMAGE:8e32aab7e59b4517_9_10]]
        Source diagram or notationSource diagram or notation
        1. It determines the penalty for rebooting or migrating the VM.
        2. It is used to predict the expected reward for each possible waiting time,allowing the algorithm to personalize decisions for each failure event.
        3. It is ignored by LinUCB, which only uses past rewards.
        4. It is used to randomly select an action to ensure exploration.

        A published solution is not available for this question yet.

        Question 248 MCQ · 2.0 marks

        [[IMAGE:8e32aab7e59b4517_8_9]] Based on the above data, answer the given subquestions.
        What is the main optimality goal of a contextual bandit algorithm over many rounds?
        Source diagram or notation
        1. Maximise the total expected reward by acting as closely as possible to the best policy mapping contexts to actions.
        2. Achieve zero regret after every individual round by always picking the empirically best action so far for each context
        3. Always exploit the highest-reward action seen so far
        4. Minimise the loss on the worst round

        A published solution is not available for this question yet.

        Question 249 NAT · 2.0 marks

        [[IMAGE:8e32aab7e59b4517_10_11]] Based on the above data, answer the given subquestions.
        What is the total reward if the robot collects both treasures and then exits,taking the shortest possible path and never hitting a wall?
        Source diagram or notation

          A published solution is not available for this question yet.

          Question 250 MSQ · 3.0 marks

          [[IMAGE:8e32aab7e59b4517_10_11]] Based on the above data, answer the given subquestions.
          [[IMAGE:8e32aab7e59b4517_11_12]]
          Source diagram or notationSource diagram or notation
          1. [[IMAGE:8e32aab7e59b4517_11_13]]
            Source diagram or notation
          2. [[IMAGE:8e32aab7e59b4517_11_14]]
            Source diagram or notation
          3. [[IMAGE:8e32aab7e59b4517_11_15]]
            Source diagram or notation
          4. [[IMAGE:8e32aab7e59b4517_11_16]]
            Source diagram or notation

          A published solution is not available for this question yet.

          Question 251 MCQ · 4.0 marks

          Consider a pole-balancing task where the goal is to apply forces to a cart moving along a track to keep a hinged pole from falling. A failure event occurs when the pole falls past a certain angle or the cart moves off the track, after which the pole resets to vertical. [[IMAGE:8e32aab7e59b4517_12_17]] The problem can be formulated either as an episodic task (with episodes ending at failure) or as a continuing task (after each failure, pole gets back into its initial position, potentially infinite time horizon with discounting).Rewards can be defined as: **• Episodic:** Reward +1 for each timestamp without failure, so the return is the number of steps until failure (possibly infinite if balanced forever). **• Continuing:** Reward 0 for each timestamp except at each failure, where the reward is −1. Returns are discounted sums of future rewards. Based on the above data, answer the given subquestions.
          [[IMAGE:8e32aab7e59b4517_12_18]]
          Source diagram or notationSource diagram or notation
          1. [[IMAGE:8e32aab7e59b4517_12_19]]
            Source diagram or notation
          2. [[IMAGE:8e32aab7e59b4517_12_20]]
            Source diagram or notation
          3. [[IMAGE:8e32aab7e59b4517_12_21]]
            Source diagram or notation

          A published solution is not available for this question yet.