MauryaHub PYQ Practice

da5007_2026T1_Q2_NA.pdf

Reinforcement Learning · Quiz 2 · Jan 2026

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 2 MCQ · 1.0 marks

What is a key advantage of using n-step TD prediction over one-step TD prediction?
  1. n-step TD prediction requires less memory and computational resources.
  2. n-step TD prediction can handle delayed rewards and credit assignment over multiple time steps
  3. n-step TD prediction converges faster to the optimal policy.
  4. n-step TD prediction guarantees convergence to the optimal value function for any choice of the learning rate.

A published solution is not available for this question yet.

Question 3 MCQ · 1.0 marks

In which of the following scenarios is Expected SARSA a good fit?
  1. When function approximation is necessary due to large state spaces, requiring deep networks for learning representations
  2. When learning needs to prioritize separating state value and advantage functions for better decision-making
  3. When the environment has a high degree of stochasticity, making bootstrapping unstable in value-based methods
  4. When experience replay is essential for stable learning and sample efficiency

A published solution is not available for this question yet.

Question 4 MCQ · 2.0 marks

Consider the following target equations: **Equation 1** [[IMAGE:a489e35127f63335_3_2]] **Equation 2** [[IMAGE:a489e35127f63335_3_3]] In both equations, [[IMAGE:a489e35127f63335_3_4]] refers to the shared parameters between the state-value network and advantage network. [[IMAGE:a489e35127f63335_3_5]] and [[IMAGE:a489e35127f63335_3_6]] refer to the non-shared parameters of the state-value and advantage network, respectively. In both equations, [[IMAGE:a489e35127f63335_3_7]] is calculated using the following network. [[IMAGE:a489e35127f63335_3_8]] Note: [[IMAGE:a489e35127f63335_3_9]] are target network parameters for online network parameters [[IMAGE:a489e35127f63335_3_10]] Which among the following are correct statements?
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
  1. Target in Equation 1 is characteristic of Dueling DQN algorithm, and the target in Equation 2 is characteristic of Dueling Double DQN algorithm
  2. Target in Equation 1 is characteristic of the DQN algorithm, and the target in Equation 2 is characteristic of the Double DQN algorithm
  3. Target in Equation 1 is characteristic of the DQN algorithm, and the target in Equation 2 is characteristic of the Dueling Double DQN algorithm
  4. Target in Equation 1 is characteristic of the Double Q-learning algorithm, and the target in Equation 2 is characteristic of Double DQN

A published solution is not available for this question yet.

Question 5 MSQ · 1.0 marks

Which of the following methods are a form of Generalized Policy Iteration?
  1. Policy Improvement
  2. Value Iteration
  3. Q-Learning and SARSA
  4. Monte Carlo control methods

A published solution is not available for this question yet.

Question 6 MSQ · 2.0 marks

In value-based deep reinforcement learning, the max operator in Q-learning can lead to systematic overestimation of Q-values when neural networks are used for function approximation Which option best explains why vanilla Deep Q-Netwok (DQN) tends to produce overoptimistic Q- value estimates?
  1. The same neural approximator [[IMAGE:a489e35127f63335_4_11]] is used to select the best action [[IMAGE:a489e35127f63335_4_12]] and to evaluate that action [[IMAGE:a489e35127f63335_4_13]] , causing the max operator to preferentially propagate positive approximation errors across correlated action-value predictions
    Source diagram or notationSource diagram or notationSource diagram or notation
  2. Delayed target network polyak-averaging [[IMAGE:a489e35127f63335_4_14]] introduces temporal mismatch in Bellman targets [[IMAGE:a489e35127f63335_4_15]] , causing overestimation
    Source diagram or notationSource diagram or notation
  3. DQN assumes a learned model of the environment [[IMAGE:a489e35127f63335_4_16]] , and errors in this model cause biased targets.
    Source diagram or notation
  4. DQN uses Monte Carlo returns [[IMAGE:a489e35127f63335_5_17]] instead of bootstrapped temporal-difference targets
    Source diagram or notation

A published solution is not available for this question yet.

Question 7 NAT · 1.0 marks

Consider the following MDP: [[IMAGE:a489e35127f63335_5_18]] Evaluate the value function [[IMAGE:a489e35127f63335_5_19]] , where [[IMAGE:a489e35127f63335_5_20]] is the equiprobable random policy. What is [[IMAGE:a489e35127f63335_5_21]] ? (Set [[IMAGE:a489e35127f63335_5_22]] )
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

    A published solution is not available for this question yet.

    Question 8 NAT · 1.0 marks

    Consider a 2D navigation environment (as shown in the figure below) where an autonomous cleaning robot must locate and collect trash items (T1 and T2) scattered in the environment and then deliver them to a designated disposal bin (G). The environment is shaped by curved, contour- like walls, forming nested spaces with only one narrow passage connecting them, which requires deliberate path planning. [[IMAGE:a489e35127f63335_6_23]] In this environment, the robot takes discrete actions from its action space [[IMAGE:a489e35127f63335_6_24]] , [[IMAGE:a489e35127f63335_6_25]] { NORTH, SOUTH, EAST, WEST } Each action attempts to move the robot by 0.1 cm in the corresponding direction. If the action would cause a collision with a wall boundary, the robot remains in the same position and receives a penalty. The grey regions represent potholes. If the robot ends up in one, it cannot escape on its own and requires human help, leading to a substantial penalty and relocation to the nearest safe position. When the agent reaches a trash item, it automatically collects it and receives a reward. Likewise, upon reaching a disposal bin, it automatically unloads the collected trash. Therefore, explicit pickup and drop actions are not required. The episode terminates either when the agent reaches the Goal (G) while carrying at least one piece of trash, or when it has fallen into potholes a total of 10 times. The position of the Robot ( [[IMAGE:a489e35127f63335_6_26]] ) can be represented as a tuple [[IMAGE:a489e35127f63335_6_27]] , where [[IMAGE:a489e35127f63335_6_28]] and [[IMAGE:a489e35127f63335_6_29]] are the distances of the robot from the bottom-left corner of the environment. For state representation, we would consider state discretisation by performing tile as well as coarse coding, which converts [[IMAGE:a489e35127f63335_6_30]] into [[IMAGE:a489e35127f63335_6_31]] { [[IMAGE:a489e35127f63335_6_32]] } [[IMAGE:a489e35127f63335_6_33]] . [[IMAGE:a489e35127f63335_6_34]] [[IMAGE:a489e35127f63335_7_35]] Observe that there are 3 different grid-tilings here: [[IMAGE:a489e35127f63335_7_36]] Based on the above data, answer the given subquestions.
    Using the coded state representation ( [[IMAGE:a489e35127f63335_7_37]] ) as input to a DQN with two hidden layers and a fully connected output layer, what should be the dimensionality of the network’s input layer?
    Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

      A published solution is not available for this question yet.

      Question 9 NAT · 1.0 marks

      Consider a 2D navigation environment (as shown in the figure below) where an autonomous cleaning robot must locate and collect trash items (T1 and T2) scattered in the environment and then deliver them to a designated disposal bin (G). The environment is shaped by curved, contour- like walls, forming nested spaces with only one narrow passage connecting them, which requires deliberate path planning. [[IMAGE:a489e35127f63335_6_23]] In this environment, the robot takes discrete actions from its action space [[IMAGE:a489e35127f63335_6_24]] , [[IMAGE:a489e35127f63335_6_25]] { NORTH, SOUTH, EAST, WEST } Each action attempts to move the robot by 0.1 cm in the corresponding direction. If the action would cause a collision with a wall boundary, the robot remains in the same position and receives a penalty. The grey regions represent potholes. If the robot ends up in one, it cannot escape on its own and requires human help, leading to a substantial penalty and relocation to the nearest safe position. When the agent reaches a trash item, it automatically collects it and receives a reward. Likewise, upon reaching a disposal bin, it automatically unloads the collected trash. Therefore, explicit pickup and drop actions are not required. The episode terminates either when the agent reaches the Goal (G) while carrying at least one piece of trash, or when it has fallen into potholes a total of 10 times. The position of the Robot ( [[IMAGE:a489e35127f63335_6_26]] ) can be represented as a tuple [[IMAGE:a489e35127f63335_6_27]] , where [[IMAGE:a489e35127f63335_6_28]] and [[IMAGE:a489e35127f63335_6_29]] are the distances of the robot from the bottom-left corner of the environment. For state representation, we would consider state discretisation by performing tile as well as coarse coding, which converts [[IMAGE:a489e35127f63335_6_30]] into [[IMAGE:a489e35127f63335_6_31]] { [[IMAGE:a489e35127f63335_6_32]] } [[IMAGE:a489e35127f63335_6_33]] . [[IMAGE:a489e35127f63335_6_34]] [[IMAGE:a489e35127f63335_7_35]] Observe that there are 3 different grid-tilings here: [[IMAGE:a489e35127f63335_7_36]] Based on the above data, answer the given subquestions.
      Using the coded state representation ( [[IMAGE:a489e35127f63335_8_38]] ) as input to a DQN with two hidden layers and a fully connected output layer, what should be the size of the output layer?
      Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

        A published solution is not available for this question yet.

        Question 10 MSQ · 1.0 marks

        Consider a 2D navigation environment (as shown in the figure below) where an autonomous cleaning robot must locate and collect trash items (T1 and T2) scattered in the environment and then deliver them to a designated disposal bin (G). The environment is shaped by curved, contour- like walls, forming nested spaces with only one narrow passage connecting them, which requires deliberate path planning. [[IMAGE:a489e35127f63335_6_23]] In this environment, the robot takes discrete actions from its action space [[IMAGE:a489e35127f63335_6_24]] , [[IMAGE:a489e35127f63335_6_25]] { NORTH, SOUTH, EAST, WEST } Each action attempts to move the robot by 0.1 cm in the corresponding direction. If the action would cause a collision with a wall boundary, the robot remains in the same position and receives a penalty. The grey regions represent potholes. If the robot ends up in one, it cannot escape on its own and requires human help, leading to a substantial penalty and relocation to the nearest safe position. When the agent reaches a trash item, it automatically collects it and receives a reward. Likewise, upon reaching a disposal bin, it automatically unloads the collected trash. Therefore, explicit pickup and drop actions are not required. The episode terminates either when the agent reaches the Goal (G) while carrying at least one piece of trash, or when it has fallen into potholes a total of 10 times. The position of the Robot ( [[IMAGE:a489e35127f63335_6_26]] ) can be represented as a tuple [[IMAGE:a489e35127f63335_6_27]] , where [[IMAGE:a489e35127f63335_6_28]] and [[IMAGE:a489e35127f63335_6_29]] are the distances of the robot from the bottom-left corner of the environment. For state representation, we would consider state discretisation by performing tile as well as coarse coding, which converts [[IMAGE:a489e35127f63335_6_30]] into [[IMAGE:a489e35127f63335_6_31]] { [[IMAGE:a489e35127f63335_6_32]] } [[IMAGE:a489e35127f63335_6_33]] . [[IMAGE:a489e35127f63335_6_34]] [[IMAGE:a489e35127f63335_7_35]] Observe that there are 3 different grid-tilings here: [[IMAGE:a489e35127f63335_7_36]] Based on the above data, answer the given subquestions.
        You have decided to use the Dueling DQN architecture with the aggregation. [[IMAGE:a489e35127f63335_8_39]] What is the primary purpose of subtracting the term [[IMAGE:a489e35127f63335_8_40]] in this formulation?
        Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
        1. It encourages the network to always choose the action with the highest advantage, thereby directly implementing a greedy policy inside the Q-network.
        2. It normalises the advantage values for a given state so that at least one action has zero advantage, reducing identifiability issues between [[IMAGE:a489e35127f63335_8_41]] and [[IMAGE:a489e35127f63335_8_42]] .
          Source diagram or notationSource diagram or notation
        3. It rescales the advantages to lie in the range [[IMAGE:a489e35127f63335_8_43]] , making training more numerically stable.
          Source diagram or notation
        4. It removes the need for a separate value stream [[IMAGE:a489e35127f63335_8_44]] , since the max operator can recover the state value from the advantages alone.
          Source diagram or notation

        A published solution is not available for this question yet.

        Question 11 MSQ · 1.0 marks

        In Deep Q-Networks (DQN), a separate target network is maintained whose weights are periodically copied from the main (online) network rather than updated at every step. Based on the above data, answer the given subquestions.
        What is the primary reason for using a separate target network with periodically updated weights in DQN?
        1. To accelerate convergence by increasing the learning rate.
        2. To stabilize learning by keeping target values fixed for several steps.
        3. To ensure the Q-values are always up to date.
        4. To share weights between the actor and the critic.

        A published solution is not available for this question yet.

        Question 12 MSQ · 1.0 marks

        In Deep Q-Networks (DQN), a separate target network is maintained whose weights are periodically copied from the main (online) network rather than updated at every step. Based on the above data, answer the given subquestions.
        What are the effects of periodically updating the weights of the target network in DQN?
        1. Provides stable targets to the main network during training.
        2. Causes the learning targets to be non-stationary.
        3. Mitigates instabilities in Q-learning updates.
        4. Avoids the need for a replay buffer.

        A published solution is not available for this question yet.

        Question 13 NAT · 0.5 marks

        Consider a discounted ( [[IMAGE:a489e35127f63335_9_45]] ), episodic task that has two non-terminal states, A and B. The rewards belongs to { 0,1,2,3 } set. The following are some episodes experienced by an agent following a fixed policy. The terminal state is not explicitly mentioned for any of the episodes. [[IMAGE:a489e35127f63335_10_46]] Based on the above data, answer the given subquestions.
        What is the estimate of [[IMAGE:a489e35127f63335_10_47]] returned by first-visit MC?
        Source diagram or notationSource diagram or notationSource diagram or notation

          A published solution is not available for this question yet.

          Question 14 NAT · 0.5 marks

          Consider a discounted ( [[IMAGE:a489e35127f63335_9_45]] ), episodic task that has two non-terminal states, A and B. The rewards belongs to { 0,1,2,3 } set. The following are some episodes experienced by an agent following a fixed policy. The terminal state is not explicitly mentioned for any of the episodes. [[IMAGE:a489e35127f63335_10_46]] Based on the above data, answer the given subquestions.
          What is the estimate of [[IMAGE:a489e35127f63335_10_48]] returned by first-visit MC?
          Source diagram or notationSource diagram or notationSource diagram or notation

            A published solution is not available for this question yet.

            Question 15 NAT · 0.5 marks

            Consider a discounted ( [[IMAGE:a489e35127f63335_9_45]] ), episodic task that has two non-terminal states, A and B. The rewards belongs to { 0,1,2,3 } set. The following are some episodes experienced by an agent following a fixed policy. The terminal state is not explicitly mentioned for any of the episodes. [[IMAGE:a489e35127f63335_10_46]] Based on the above data, answer the given subquestions.
            What is the estimate of [[IMAGE:a489e35127f63335_11_49]] returned by batch [[IMAGE:a489e35127f63335_11_50]] ?
            Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

              A published solution is not available for this question yet.

              Question 16 NAT · 0.5 marks

              Consider a discounted ( [[IMAGE:a489e35127f63335_9_45]] ), episodic task that has two non-terminal states, A and B. The rewards belongs to { 0,1,2,3 } set. The following are some episodes experienced by an agent following a fixed policy. The terminal state is not explicitly mentioned for any of the episodes. [[IMAGE:a489e35127f63335_10_46]] Based on the above data, answer the given subquestions.
              What is the estimate of [[IMAGE:a489e35127f63335_11_51]] returned by batch [[IMAGE:a489e35127f63335_11_52]] ?
              Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

                A published solution is not available for this question yet.

                Question 17 NAT · 1.0 marks

                Consider the following Deep Q-Network given below that approximates the Q-values for actions [[IMAGE:a489e35127f63335_11_53]] and [[IMAGE:a489e35127f63335_11_54]] for states [[IMAGE:a489e35127f63335_11_55]] . The hidden layer nodes (colored green) employ the [[IMAGE:a489e35127f63335_11_56]] activation function. The weights of the connections are given directly above the edges connecting the nodes. For instance, 3 is the weight for edge [[IMAGE:a489e35127f63335_12_57]] and -3 is the weight for edge [[IMAGE:a489e35127f63335_12_58]] . [[IMAGE:a489e35127f63335_12_59]] We are interested in the Q-values calculated by the network for the following states: [[IMAGE:a489e35127f63335_12_60]] Based on the above data, answer the given subquestions.
                Recall the linear-approximator view for DQNs ( [[IMAGE:a489e35127f63335_12_61]] ). Evaluate [[IMAGE:a489e35127f63335_12_62]] i.e. the state-representation for [[IMAGE:a489e35127f63335_12_63]] . Enter [[IMAGE:a489e35127f63335_12_64]] (the L-1 norm of the state-representation) as your answer.
                Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

                  A published solution is not available for this question yet.

                  Question 18 MCQ · 1.0 marks

                  Consider the following Deep Q-Network given below that approximates the Q-values for actions [[IMAGE:a489e35127f63335_11_53]] and [[IMAGE:a489e35127f63335_11_54]] for states [[IMAGE:a489e35127f63335_11_55]] . The hidden layer nodes (colored green) employ the [[IMAGE:a489e35127f63335_11_56]] activation function. The weights of the connections are given directly above the edges connecting the nodes. For instance, 3 is the weight for edge [[IMAGE:a489e35127f63335_12_57]] and -3 is the weight for edge [[IMAGE:a489e35127f63335_12_58]] . [[IMAGE:a489e35127f63335_12_59]] We are interested in the Q-values calculated by the network for the following states: [[IMAGE:a489e35127f63335_12_60]] Based on the above data, answer the given subquestions.
                  Obtain the state-representations for [[IMAGE:a489e35127f63335_12_65]] and [[IMAGE:a489e35127f63335_12_66]] as well. Which of the following pairs of states are the closest (euclidean distance) in their respective representations?
                  Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
                  1. [[IMAGE:a489e35127f63335_13_67]]
                    Source diagram or notation
                  2. [[IMAGE:a489e35127f63335_13_68]]
                    Source diagram or notation
                  3. [[IMAGE:a489e35127f63335_13_69]]
                    Source diagram or notation

                  A published solution is not available for this question yet.

                  Question 19 NAT · 1.0 marks

                  Consider the following Deep Q-Network given below that approximates the Q-values for actions [[IMAGE:a489e35127f63335_11_53]] and [[IMAGE:a489e35127f63335_11_54]] for states [[IMAGE:a489e35127f63335_11_55]] . The hidden layer nodes (colored green) employ the [[IMAGE:a489e35127f63335_11_56]] activation function. The weights of the connections are given directly above the edges connecting the nodes. For instance, 3 is the weight for edge [[IMAGE:a489e35127f63335_12_57]] and -3 is the weight for edge [[IMAGE:a489e35127f63335_12_58]] . [[IMAGE:a489e35127f63335_12_59]] We are interested in the Q-values calculated by the network for the following states: [[IMAGE:a489e35127f63335_12_60]] Based on the above data, answer the given subquestions.
                  What is [[IMAGE:a489e35127f63335_13_70]] ?
                  Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

                    A published solution is not available for this question yet.

                    Question 20 NAT · 1.0 marks

                    Consider the following Deep Q-Network given below that approximates the Q-values for actions [[IMAGE:a489e35127f63335_11_53]] and [[IMAGE:a489e35127f63335_11_54]] for states [[IMAGE:a489e35127f63335_11_55]] . The hidden layer nodes (colored green) employ the [[IMAGE:a489e35127f63335_11_56]] activation function. The weights of the connections are given directly above the edges connecting the nodes. For instance, 3 is the weight for edge [[IMAGE:a489e35127f63335_12_57]] and -3 is the weight for edge [[IMAGE:a489e35127f63335_12_58]] . [[IMAGE:a489e35127f63335_12_59]] We are interested in the Q-values calculated by the network for the following states: [[IMAGE:a489e35127f63335_12_60]] Based on the above data, answer the given subquestions.
                    What is [[IMAGE:a489e35127f63335_13_71]] ?
                    Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

                      A published solution is not available for this question yet.

                      Question 21 NAT · 1.0 marks

                      Consider the following Deep Q-Network given below that approximates the Q-values for actions [[IMAGE:a489e35127f63335_11_53]] and [[IMAGE:a489e35127f63335_11_54]] for states [[IMAGE:a489e35127f63335_11_55]] . The hidden layer nodes (colored green) employ the [[IMAGE:a489e35127f63335_11_56]] activation function. The weights of the connections are given directly above the edges connecting the nodes. For instance, 3 is the weight for edge [[IMAGE:a489e35127f63335_12_57]] and -3 is the weight for edge [[IMAGE:a489e35127f63335_12_58]] . [[IMAGE:a489e35127f63335_12_59]] We are interested in the Q-values calculated by the network for the following states: [[IMAGE:a489e35127f63335_12_60]] Based on the above data, answer the given subquestions.
                      Suppose you use [[IMAGE:a489e35127f63335_13_72]] -greedy with [[IMAGE:a489e35127f63335_13_73]] selection for balancing exploration with exploitation during training. What is the probability that you'll pick action O2 given that the observed state is [[IMAGE:a489e35127f63335_14_74]] ?
                      Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

                        A published solution is not available for this question yet.