MauryaHub PYQ Practice

da5007_2025T3_Q2_NA.pdf

Reinforcement Learning · Quiz 2 · Sep 2025

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 180 NAT · 0.5 marks

[[IMAGE:af9915b9b07d2db6_1_0]]
Source diagram or notation

    A published solution is not available for this question yet.

    Question 182 MCQ · 2.0 marks

    Consider the following assertion reason pair: **Assertion**: In Monte Carlo (MC), the value function converges to the certainty equivalence estimate. **Reason**: Monte Carlo methods use complete trajectories to fit the value function as closely as possible to the sampled returns.
    1. Both Assertion and Reason are correct, and Reason is the correct explanation.
    2. Assertion is correct, Reason is incorrect
    3. Assertion is incorrect, Reason is correct
    4. Both Assertion and Reason are correct, but Reason is not the correct explanation.

    A published solution is not available for this question yet.

    Question 183 MSQ · 3.0 marks

    Which of the following is the correct way to represent policy for policy search methods? Assume [[IMAGE:af9915b9b07d2db6_2_1]] represents a real-valued parameter.
    Source diagram or notation
    1. [[IMAGE:af9915b9b07d2db6_2_2]]
      Source diagram or notation
    2. [[IMAGE:af9915b9b07d2db6_3_3]]
      Source diagram or notation
    3. [[IMAGE:af9915b9b07d2db6_3_4]]
      Source diagram or notation
    4. [[IMAGE:af9915b9b07d2db6_3_5]]
      Source diagram or notation
    5. [[IMAGE:af9915b9b07d2db6_3_6]]
      Source diagram or notation

    A published solution is not available for this question yet.

    Question 184 MSQ · 3.0 marks

    Recall the incremental update rule for REINFORCE: [[IMAGE:af9915b9b07d2db6_3_7]] Consider the following binary-bandit problem: [[IMAGE:af9915b9b07d2db6_3_8]] Which of the following expressions are equivalent to [[IMAGE:af9915b9b07d2db6_3_9]] ?
    Source diagram or notationSource diagram or notationSource diagram or notation
    1. [[IMAGE:af9915b9b07d2db6_3_10]]
      Source diagram or notation
    2. [[IMAGE:af9915b9b07d2db6_3_11]]
      Source diagram or notation
    3. [[IMAGE:af9915b9b07d2db6_3_12]]
      Source diagram or notation
    4. [[IMAGE:af9915b9b07d2db6_4_13]]
      Source diagram or notation

    A published solution is not available for this question yet.

    Question 185 MSQ · 2.0 marks

    In Q-Learning, the update rule for the action-value function is based on bootstrapping from the current estimate. Which of the following correctly represents this update?
    1. [[IMAGE:af9915b9b07d2db6_4_14]]
      Source diagram or notation
    2. [[IMAGE:af9915b9b07d2db6_4_15]]
      Source diagram or notation
    3. [[IMAGE:af9915b9b07d2db6_4_16]]
      Source diagram or notation
    4. [[IMAGE:af9915b9b07d2db6_4_17]]
      Source diagram or notation

    A published solution is not available for this question yet.

    Question 186 NAT · 3.0 marks

    Consider an episodic task where the [[IMAGE:af9915b9b07d2db6_4_18]] -step returns follow a decaying exponential pattern: [[IMAGE:af9915b9b07d2db6_4_19]] Assume the episode is sufficiently long so the forward-view infinite sum is valid. Compute the [[IMAGE:af9915b9b07d2db6_4_20]] -return for [[IMAGE:af9915b9b07d2db6_4_21]] . Enter your answer correct to **two decimal places**.
    Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

      A published solution is not available for this question yet.

      Question 187 MCQ · 2.0 marks

      In Deep Q-Networks (DQN), a separate **target network** is maintained, whose weights are periodically copied from the main (online) network instead of being updated at every step. Based on the above data, answer the given subquestions.
      What is the primary reason for using a separate target network with periodically updated weights in DQN?
      1. To accelerate convergence by increasing the learning rate.
      2. To stabilize learning by keeping target values fixed for several steps.
      3. To ensure the Q-values are always up to date.
      4. To share weights between actor and critic models.

      A published solution is not available for this question yet.

      Question 188 MSQ · 2.0 marks

      In Deep Q-Networks (DQN), a separate **target network** is maintained, whose weights are periodically copied from the main (online) network instead of being updated at every step. Based on the above data, answer the given subquestions.
      What are the effects of periodically updating the weights of the target network in DQN?
      1. Provides stable targets to the main network during training.
      2. Causes the learning targets to be non-stationary.
      3. Mitigates instabilities in Q-learning updates.
      4. Avoids the need for a replay buffer.

      A published solution is not available for this question yet.

      Question 189 MCQ · 1.0 marks

      [[IMAGE:af9915b9b07d2db6_6_22]] Based on the above data, answer the given subquestions.
      Find the Q values for all three actions [[IMAGE:af9915b9b07d2db6_6_23]] for the state [[IMAGE:af9915b9b07d2db6_6_24]] . Express your answer as a vector q, where [[IMAGE:af9915b9b07d2db6_6_25]]
      Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
      1. [[IMAGE:af9915b9b07d2db6_6_26]]
        Source diagram or notation
      2. [[IMAGE:af9915b9b07d2db6_6_27]]
        Source diagram or notation
      3. [[IMAGE:af9915b9b07d2db6_6_28]]
        Source diagram or notation
      4. [[IMAGE:af9915b9b07d2db6_6_29]]
        Source diagram or notation

      A published solution is not available for this question yet.

      Question 190 NAT · 4.0 marks

      [[IMAGE:af9915b9b07d2db6_6_22]] Based on the above data, answer the given subquestions.
      Compute the TD target for this transition using SARSA with [[IMAGE:af9915b9b07d2db6_7_30]] . Use a greedy policy derived from the Q-values at [[IMAGE:af9915b9b07d2db6_7_31]] to select the next action.
      Source diagram or notationSource diagram or notationSource diagram or notation

        A published solution is not available for this question yet.

        Question 191 MCQ · 4.0 marks

        [[IMAGE:af9915b9b07d2db6_6_22]] Based on the above data, answer the given subquestions.
        Perform one Step of semi-gradient TD using this transition. What would be [[IMAGE:af9915b9b07d2db6_7_32]] if [[IMAGE:af9915b9b07d2db6_7_33]] ?
        Source diagram or notationSource diagram or notationSource diagram or notation
        1. [[IMAGE:af9915b9b07d2db6_7_34]]
          Source diagram or notation
        2. [[IMAGE:af9915b9b07d2db6_7_35]]
          Source diagram or notation
        3. [[IMAGE:af9915b9b07d2db6_7_36]]
          Source diagram or notation
        4. [[IMAGE:af9915b9b07d2db6_7_37]]
          Source diagram or notation

        A published solution is not available for this question yet.

        Question 192 MCQ · 4.0 marks

        Consider the following algorithm for learning action-value estimates in an episodic MDP: [[IMAGE:af9915b9b07d2db6_8_38]] Based on the above data, answer the given subquestions.
        Which of the following issues in standard Q-learning does the given algorithm aim to reduce?
        Source diagram or notation
        1. Overestimation of action values due to maximisation bias.
        2. Underestimation of terminal state values.
        3. Instability due to non-stationary rewards.
        4. Bias introduced by [[IMAGE:af9915b9b07d2db6_8_39]] -greedy exploration.
          Source diagram or notation

        A published solution is not available for this question yet.

        Question 193 MSQ · 4.0 marks

        Consider the following algorithm for learning action-value estimates in an episodic MDP: [[IMAGE:af9915b9b07d2db6_8_38]] Based on the above data, answer the given subquestions.
        In the given pseudocode, the algorithm maintains two value functions, [[IMAGE:af9915b9b07d2db6_8_40]] and [[IMAGE:af9915b9b07d2db6_8_41]] , and updates them alternately. Which of the following best describes the purpose and effect of this design choice?
        Source diagram or notationSource diagram or notationSource diagram or notation
        1. It ensures that the same target network is used for both selection and evaluation, improving stability.
        2. It ensures the overall estimate equals the true expected return by averaging [[IMAGE:af9915b9b07d2db6_9_42]] and [[IMAGE:af9915b9b07d2db6_9_43]] .
          Source diagram or notationSource diagram or notation
        3. It decorrelates the selection and evaluation of the greedy action, reducing overestimation bias.
        4. It doubles the effective learning rate by updating two estimators in parallel.

        A published solution is not available for this question yet.

        Question 194 MCQ · 3.0 marks

        Recall from the lectures, the forward-view [[IMAGE:af9915b9b07d2db6_9_44]] -**return** is defined as: [[IMAGE:af9915b9b07d2db6_9_45]] where [[IMAGE:af9915b9b07d2db6_9_46]] denotes the [[IMAGE:af9915b9b07d2db6_9_47]] -step return starting at time [[IMAGE:af9915b9b07d2db6_9_48]] .Consider using the TD( [[IMAGE:af9915b9b07d2db6_9_49]] ) algorithm for an **episodic** task. Based on the above data, answer the given subquestions.
        What is the effect of setting [[IMAGE:af9915b9b07d2db6_9_50]] in TD( [[IMAGE:af9915b9b07d2db6_9_51]] ) on the weighting of [[IMAGE:af9915b9b07d2db6_9_52]] -step returns?
        Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
        1. All [[IMAGE:af9915b9b07d2db6_9_53]] -step returns are ignored.
          Source diagram or notation
        2. Only the 1-step return contributes, with weight 1.
        3. All [[IMAGE:af9915b9b07d2db6_9_54]] -step returns contribute equally.
          Source diagram or notation
        4. The weighting distribution remains unchanged.

        A published solution is not available for this question yet.

        Question 195 MCQ · 3.0 marks

        Recall from the lectures, the forward-view [[IMAGE:af9915b9b07d2db6_9_44]] -**return** is defined as: [[IMAGE:af9915b9b07d2db6_9_45]] where [[IMAGE:af9915b9b07d2db6_9_46]] denotes the [[IMAGE:af9915b9b07d2db6_9_47]] -step return starting at time [[IMAGE:af9915b9b07d2db6_9_48]] .Consider using the TD( [[IMAGE:af9915b9b07d2db6_9_49]] ) algorithm for an **episodic** task. Based on the above data, answer the given subquestions.
        As [[IMAGE:af9915b9b07d2db6_10_55]] in an episodic TD( [[IMAGE:af9915b9b07d2db6_10_56]] ) task, the [[IMAGE:af9915b9b07d2db6_10_57]] -return [[IMAGE:af9915b9b07d2db6_10_58]] approaches which of the following?
        Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
        1. The Monte Carlo return.
        2. The one-step TD target.
        3. The average of all [[IMAGE:af9915b9b07d2db6_10_59]] -step returns.
          Source diagram or notation
        4. Zero, since weights vanish as [[IMAGE:af9915b9b07d2db6_10_60]] .
          Source diagram or notation

        A published solution is not available for this question yet.

        Question 196 NAT · 3.0 marks

        Consider an episodic task with four non-terminal states: A, B, C and D. The following are some episodes experienced by an agent following a fixed policy. The terminal state is not explicitly mentioned for any of the episodes. A,1, C,1, D,0, C,0, A,1, B,1 A,0, C,0, D,1, C,1, B,1 D,1, C,0, A,0, C,1, B,0 C,0, A,1, C,1, B,0 C,1, B,0 D,0, B,1 B,1 B,0 (Enter your answer correct to **three decimal places**.) Based on the above data, answer the given subquestions.
        Using every-visit Monte Carlo, estimate [[IMAGE:af9915b9b07d2db6_10_61]] and [[IMAGE:af9915b9b07d2db6_10_62]] . What is the value of [[IMAGE:af9915b9b07d2db6_10_63]] ? (use [[IMAGE:af9915b9b07d2db6_10_64]] )
        Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

          A published solution is not available for this question yet.