MauryaHub PYQ Practice

da5007_2025T3_ET_FN.pdf

Reinforcement Learning · End Term · Sep 2025 FN

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 2 MCQ · 2.0 marks

Consider a discounted return [[IMAGE:cfe1f3d865067824_2_2]] in an infinite-horizon MDP with bounded rewards [[IMAGE:cfe1f3d865067824_2_3]] and discount factor [[IMAGE:cfe1f3d865067824_2_4]] . Which of the following statements is/are correct? 1. For fixed rewards, the contribution of [[IMAGE:cfe1f3d865067824_2_5]] to [[IMAGE:cfe1f3d865067824_2_6]] decays geometrically as [[IMAGE:cfe1f3d865067824_2_7]] . 2. If [[IMAGE:cfe1f3d865067824_2_8]] and rewards are uniformly bounded, the infinite sum [[IMAGE:cfe1f3d865067824_2_9]] is always finite. 3. For [[IMAGE:cfe1f3d865067824_2_10]] , the discounted return [[IMAGE:cfe1f3d865067824_2_11]] is bounded above in magnitude by [[IMAGE:cfe1f3d865067824_2_12]] .
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
  1. 1, 2, 3
  2. 1, 3
  3. 2, 3
  4. only 1
  5. only 3

A published solution is not available for this question yet.

Question 3 MCQ · 1.0 marks

In the UCT (Upper Confidence bounds applied to Trees) variant of Monte Carlo Tree Search, the tree policy at a state [[IMAGE:cfe1f3d865067824_2_13]] typically selects an action [[IMAGE:cfe1f3d865067824_2_14]] by maximising
Source diagram or notationSource diagram or notation
  1. A pure exploitation term based only on the estimated value [[IMAGE:cfe1f3d865067824_2_15]]
    Source diagram or notation
  2. A pure exploration term based only on visit counts [[IMAGE:cfe1f3d865067824_2_16]]
    Source diagram or notation
  3. A combination of estimated value and an exploration bonus based on visit counts
  4. A random choice among all child actions

A published solution is not available for this question yet.

Question 4 MCQ · 2.0 marks

Consider the following target equations: **Equation 1** [[IMAGE:cfe1f3d865067824_3_17]] **Equation 2** [[IMAGE:cfe1f3d865067824_3_18]] In both equations, [[IMAGE:cfe1f3d865067824_3_19]] refers to the shared parameters between the state-value network and advantage network. [[IMAGE:cfe1f3d865067824_3_20]] and [[IMAGE:cfe1f3d865067824_3_21]] refer to the non-shared parameters of the state-value and advantage network, respectively. In both equations, [[IMAGE:cfe1f3d865067824_3_22]] is calculated as follows: [[IMAGE:cfe1f3d865067824_3_23]] Note: [[IMAGE:cfe1f3d865067824_3_24]] are target network parameters for online network parameters [[IMAGE:cfe1f3d865067824_3_25]] Which among the following are correct statements?
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
  1. Target in Equation 1 is characteristic of Dueling DQN algorithm and target in equation 2 is characteristic of Dueling DDQN alogrithm
  2. Target in Equation 2 is characteristic of DQN algorithm and target in equation 2 is characteristic of Double DQN alogrithm
  3. Target in Equation 2 is characteristic of the DQN algorithm, and the target in Equation 2 is characteristic of the Dueling DDQN alogrithm
  4. Target in Equation 2 is characteristic of the Double Q-learning algorithm and the target in Equation 2 is characteristic of DDQN

A published solution is not available for this question yet.

Question 5 MSQ · 2.0 marks

Consider the REINFORCE update with baseline for a stochastic policy [[IMAGE:cfe1f3d865067824_4_26]] : [[IMAGE:cfe1f3d865067824_4_27]] Which of the following statements are true?
Source diagram or notationSource diagram or notation
  1. Using a state-dependent baseline [[IMAGE:cfe1f3d865067824_4_28]] can reduce the variance of the gradient estimate without introducing bias.
    Source diagram or notation
  2. Choosing [[IMAGE:cfe1f3d865067824_4_29]] makes the update equivalent to using an advantage function.
    Source diagram or notation
  3. If [[IMAGE:cfe1f3d865067824_4_30]] depends on the action [[IMAGE:cfe1f3d865067824_4_31]] , the gradient estimate can become biased.
    Source diagram or notationSource diagram or notation
  4. Any baseline [[IMAGE:cfe1f3d865067824_4_32]] that is independent of [[IMAGE:cfe1f3d865067824_4_33]] will necessarily make the gradient estimate unbiased.
    Source diagram or notationSource diagram or notation

A published solution is not available for this question yet.

Question 6 NAT · 2.0 marks

Consider following MDP, assume [[IMAGE:cfe1f3d865067824_5_34]] : [[IMAGE:cfe1f3d865067824_5_35]] The edges have the value [[IMAGE:cfe1f3d865067824_5_36]] , where [[IMAGE:cfe1f3d865067824_5_37]] denotes the transition probability and [[IMAGE:cfe1f3d865067824_5_38]] is the immediate expected reward. [[IMAGE:cfe1f3d865067824_5_39]] are non-terminal states and [[IMAGE:cfe1f3d865067824_5_40]] is a terminal state. Based on the above data, answer the given
What is the probability that the accumulating eligibility trace for state [[IMAGE:cfe1f3d865067824_5_41]] is more than 3 at the end of a trajectory? **(Write your answer up to 4 decimal places)**
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

    A published solution is not available for this question yet.

    Question 7 NAT · 2.0 marks

    Consider following MDP, assume [[IMAGE:cfe1f3d865067824_5_34]] : [[IMAGE:cfe1f3d865067824_5_35]] The edges have the value [[IMAGE:cfe1f3d865067824_5_36]] , where [[IMAGE:cfe1f3d865067824_5_37]] denotes the transition probability and [[IMAGE:cfe1f3d865067824_5_38]] is the immediate expected reward. [[IMAGE:cfe1f3d865067824_5_39]] are non-terminal states and [[IMAGE:cfe1f3d865067824_5_40]] is a terminal state. Based on the above data, answer the given
    Consider the following trajectory that the agent follows: [[IMAGE:cfe1f3d865067824_5_42]] What will the eligibility trace of state [[IMAGE:cfe1f3d865067824_5_43]] just before the episode terminates?**(Write your answer** **up to 4 decimal places)**
    Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

      A published solution is not available for this question yet.

      Question 8 NAT · 2.0 marks

      Consider a 2D navigation environment (as shown in the figure below) where an autonomous cleaning robot must locate and collect trash items (T1 and T2) scattered in the environment and then deliver them to a designated disposal bin (G). The environment is shaped by curved, contour- like walls, forming nested spaces with only one narrow passage connecting them, which requires deliberate path planning. [[IMAGE:cfe1f3d865067824_7_44]] In this environment, the robot takes discrete actions from its action space [[IMAGE:cfe1f3d865067824_7_45]] , [[IMAGE:cfe1f3d865067824_7_46]] { NORTH, SOUTH, EAST, WEST } Each action attempts to move the robot by [[IMAGE:cfe1f3d865067824_7_47]] cm in the corresponding direction. If the action would cause a collision with a wall boundary, the robot remains in the same position and receives a penalty. The grey regions represent potholes. If the robot ends up in one, it cannot escape on its own and requires human help, leading to a substantial penalty and relocation to the nearest safe position. When the agent reaches a trash item, it automatically collects it and receives a reward. Likewise, upon reaching a disposal bin, it automatically unloads the collected trash. Therefore, explicit pickup and drop actions are not required. The episode terminates either when the agent reaches the Goal (G) while carrying at least one piece of trash, or when it has fallen into potholes a total of 10 times. The position of the Robot ( [[IMAGE:cfe1f3d865067824_7_48]] ) can be represented as a tuple [[IMAGE:cfe1f3d865067824_7_49]] , where [[IMAGE:cfe1f3d865067824_7_50]] and [[IMAGE:cfe1f3d865067824_7_51]] are the distances of the robot from the bottom-left corner of the environment. For state representation, we would consider state discretisation by performing tile as well as coarse coding, which converts [[IMAGE:cfe1f3d865067824_8_52]] into [[IMAGE:cfe1f3d865067824_8_53]] . [[IMAGE:cfe1f3d865067824_8_54]] Observe that there is 1 grid-tiling and 1 coarse-tiling here: [[IMAGE:cfe1f3d865067824_8_55]] [[IMAGE:cfe1f3d865067824_8_56]] Based on the above data, answer the given
      Using the coded state representation ( [[IMAGE:cfe1f3d865067824_8_57]] ) as input to a DQN with two hidden layers and a fully connected output layer, what should be the dimensionality of the network’s input layer?
      Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

        A published solution is not available for this question yet.

        Question 9 NAT · 2.0 marks

        Consider a 2D navigation environment (as shown in the figure below) where an autonomous cleaning robot must locate and collect trash items (T1 and T2) scattered in the environment and then deliver them to a designated disposal bin (G). The environment is shaped by curved, contour- like walls, forming nested spaces with only one narrow passage connecting them, which requires deliberate path planning. [[IMAGE:cfe1f3d865067824_7_44]] In this environment, the robot takes discrete actions from its action space [[IMAGE:cfe1f3d865067824_7_45]] , [[IMAGE:cfe1f3d865067824_7_46]] { NORTH, SOUTH, EAST, WEST } Each action attempts to move the robot by [[IMAGE:cfe1f3d865067824_7_47]] cm in the corresponding direction. If the action would cause a collision with a wall boundary, the robot remains in the same position and receives a penalty. The grey regions represent potholes. If the robot ends up in one, it cannot escape on its own and requires human help, leading to a substantial penalty and relocation to the nearest safe position. When the agent reaches a trash item, it automatically collects it and receives a reward. Likewise, upon reaching a disposal bin, it automatically unloads the collected trash. Therefore, explicit pickup and drop actions are not required. The episode terminates either when the agent reaches the Goal (G) while carrying at least one piece of trash, or when it has fallen into potholes a total of 10 times. The position of the Robot ( [[IMAGE:cfe1f3d865067824_7_48]] ) can be represented as a tuple [[IMAGE:cfe1f3d865067824_7_49]] , where [[IMAGE:cfe1f3d865067824_7_50]] and [[IMAGE:cfe1f3d865067824_7_51]] are the distances of the robot from the bottom-left corner of the environment. For state representation, we would consider state discretisation by performing tile as well as coarse coding, which converts [[IMAGE:cfe1f3d865067824_8_52]] into [[IMAGE:cfe1f3d865067824_8_53]] . [[IMAGE:cfe1f3d865067824_8_54]] Observe that there is 1 grid-tiling and 1 coarse-tiling here: [[IMAGE:cfe1f3d865067824_8_55]] [[IMAGE:cfe1f3d865067824_8_56]] Based on the above data, answer the given
        Using the coded state representation ( [[IMAGE:cfe1f3d865067824_9_58]] ) as input to a DQN with two hidden layers and a fully connected output layer, what should be the size of the output layer?
        Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

          A published solution is not available for this question yet.

          Question 10 MSQ · 2.0 marks

          Consider a 2D navigation environment (as shown in the figure below) where an autonomous cleaning robot must locate and collect trash items (T1 and T2) scattered in the environment and then deliver them to a designated disposal bin (G). The environment is shaped by curved, contour- like walls, forming nested spaces with only one narrow passage connecting them, which requires deliberate path planning. [[IMAGE:cfe1f3d865067824_7_44]] In this environment, the robot takes discrete actions from its action space [[IMAGE:cfe1f3d865067824_7_45]] , [[IMAGE:cfe1f3d865067824_7_46]] { NORTH, SOUTH, EAST, WEST } Each action attempts to move the robot by [[IMAGE:cfe1f3d865067824_7_47]] cm in the corresponding direction. If the action would cause a collision with a wall boundary, the robot remains in the same position and receives a penalty. The grey regions represent potholes. If the robot ends up in one, it cannot escape on its own and requires human help, leading to a substantial penalty and relocation to the nearest safe position. When the agent reaches a trash item, it automatically collects it and receives a reward. Likewise, upon reaching a disposal bin, it automatically unloads the collected trash. Therefore, explicit pickup and drop actions are not required. The episode terminates either when the agent reaches the Goal (G) while carrying at least one piece of trash, or when it has fallen into potholes a total of 10 times. The position of the Robot ( [[IMAGE:cfe1f3d865067824_7_48]] ) can be represented as a tuple [[IMAGE:cfe1f3d865067824_7_49]] , where [[IMAGE:cfe1f3d865067824_7_50]] and [[IMAGE:cfe1f3d865067824_7_51]] are the distances of the robot from the bottom-left corner of the environment. For state representation, we would consider state discretisation by performing tile as well as coarse coding, which converts [[IMAGE:cfe1f3d865067824_8_52]] into [[IMAGE:cfe1f3d865067824_8_53]] . [[IMAGE:cfe1f3d865067824_8_54]] Observe that there is 1 grid-tiling and 1 coarse-tiling here: [[IMAGE:cfe1f3d865067824_8_55]] [[IMAGE:cfe1f3d865067824_8_56]] Based on the above data, answer the given
          In the given environment, the state space is continuous while the action space is discrete. Now consider a similar environment where the action space is also continuous, with the action defined as the robot’s rotation angle relative to the positive horizontal axis. You decide to use a **Deterministic Policy Gradient (DPG)** method with a deterministic policy [[IMAGE:cfe1f3d865067824_9_59]] . Which of the following statements are valid for ensuring adequate exploration and obtaining unbiased policy-gradient estimates?
          Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
          1. In a deterministic but stochastic environment, the environment’s own randomness can provide sufficient exploration, so an explicit stochastic behaviour policy is not strictly necessary for exploration.
          2. In a deterministic and non-stochastic environment, to ensure exploration we can use an off-policy actor–critic setup, where the behaviour policy [[IMAGE:cfe1f3d865067824_9_60]] is more exploratory (e.g., [[IMAGE:cfe1f3d865067824_9_61]] , where [[IMAGE:cfe1f3d865067824_9_62]] is small noise) than the deterministic target policy [[IMAGE:cfe1f3d865067824_9_63]] .
            Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
          3. In off-policy DPG, the policy gradient is taken with respect to the behaviour policy [[IMAGE:cfe1f3d865067824_9_64]] , since it is the one generating the data, rather than with respect to the deterministic target policy [[IMAGE:cfe1f3d865067824_9_65]] .
            Source diagram or notationSource diagram or notation
          4. For DPG with a function-approximated critic [[IMAGE:cfe1f3d865067824_9_66]] , using a compatible parametrisation between [[IMAGE:cfe1f3d865067824_9_67]] and [[IMAGE:cfe1f3d865067824_9_68]] can help ensure that the deterministic policy gradient estimate remains unbiased, even when [[IMAGE:cfe1f3d865067824_10_69]] is approximate.
            Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

          A published solution is not available for this question yet.

          Question 11 MCQ · 2.0 marks

          Consider a 2D navigation environment (as shown in the figure below) where an autonomous cleaning robot must locate and collect trash items (T1 and T2) scattered in the environment and then deliver them to a designated disposal bin (G). The environment is shaped by curved, contour- like walls, forming nested spaces with only one narrow passage connecting them, which requires deliberate path planning. [[IMAGE:cfe1f3d865067824_7_44]] In this environment, the robot takes discrete actions from its action space [[IMAGE:cfe1f3d865067824_7_45]] , [[IMAGE:cfe1f3d865067824_7_46]] { NORTH, SOUTH, EAST, WEST } Each action attempts to move the robot by [[IMAGE:cfe1f3d865067824_7_47]] cm in the corresponding direction. If the action would cause a collision with a wall boundary, the robot remains in the same position and receives a penalty. The grey regions represent potholes. If the robot ends up in one, it cannot escape on its own and requires human help, leading to a substantial penalty and relocation to the nearest safe position. When the agent reaches a trash item, it automatically collects it and receives a reward. Likewise, upon reaching a disposal bin, it automatically unloads the collected trash. Therefore, explicit pickup and drop actions are not required. The episode terminates either when the agent reaches the Goal (G) while carrying at least one piece of trash, or when it has fallen into potholes a total of 10 times. The position of the Robot ( [[IMAGE:cfe1f3d865067824_7_48]] ) can be represented as a tuple [[IMAGE:cfe1f3d865067824_7_49]] , where [[IMAGE:cfe1f3d865067824_7_50]] and [[IMAGE:cfe1f3d865067824_7_51]] are the distances of the robot from the bottom-left corner of the environment. For state representation, we would consider state discretisation by performing tile as well as coarse coding, which converts [[IMAGE:cfe1f3d865067824_8_52]] into [[IMAGE:cfe1f3d865067824_8_53]] . [[IMAGE:cfe1f3d865067824_8_54]] Observe that there is 1 grid-tiling and 1 coarse-tiling here: [[IMAGE:cfe1f3d865067824_8_55]] [[IMAGE:cfe1f3d865067824_8_56]] Based on the above data, answer the given
          You have decided to use the Dueling DQN architecture with the aggregation. [[IMAGE:cfe1f3d865067824_10_70]] What is the primary purpose of subtracting the term [[IMAGE:cfe1f3d865067824_10_71]] in this formulation?
          Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
          1. To clip the magnitude of the advantage values and thereby prevent exploding Q-values.
          2. To guarantee that the action with the largest advantage always has a strictly positive Q-value.
          3. To force the average advantage over all actions in a state to be zero, making the decomposition into [[IMAGE:cfe1f3d865067824_10_72]] and [[IMAGE:cfe1f3d865067824_10_73]] well-defined.
            Source diagram or notationSource diagram or notation
          4. To allow the network to represent any Q-function using only the advantage stream, making the value stream redundant.

          A published solution is not available for this question yet.

          Question 12 NAT · 2.0 marks

          Consider the following Dueling DQN architecture given below that approximates the state-value function ( [[IMAGE:cfe1f3d865067824_11_74]] ) and advantage for actions A and B for states [[IMAGE:cfe1f3d865067824_11_75]] . The hidden layer nodes (grey colored nodes) employ the RELU activation function. [[IMAGE:cfe1f3d865067824_11_76]] The weight matrices of the network are given by: [[IMAGE:cfe1f3d865067824_11_77]] Here, [[IMAGE:cfe1f3d865067824_11_78]] contains the weights connecting the **input layer** to the **first hidden layer**. Similarly, [[IMAGE:cfe1f3d865067824_11_79]] contains the weights connecting the **first hidden layer** [[IMAGE:cfe1f3d865067824_11_80]] to the **second hidden** **layer** [[IMAGE:cfe1f3d865067824_11_81]] . [[IMAGE:cfe1f3d865067824_12_82]] [[IMAGE:cfe1f3d865067824_12_83]] contains the weights connecting the [[IMAGE:cfe1f3d865067824_12_84]] to the [[IMAGE:cfe1f3d865067824_12_85]] . [[IMAGE:cfe1f3d865067824_12_86]] contains the weights connecting the layer [[IMAGE:cfe1f3d865067824_12_87]] to the layer [[IMAGE:cfe1f3d865067824_12_88]] . Lastly, [[IMAGE:cfe1f3d865067824_12_89]] contains the weights connecting the layer [[IMAGE:cfe1f3d865067824_12_90]] to the layer [[IMAGE:cfe1f3d865067824_12_91]] . Each entry in these matrices corresponds to a single scalar weight. For example, the entry [[IMAGE:cfe1f3d865067824_12_92]] in [[IMAGE:cfe1f3d865067824_12_93]] denotes the weight from neuron [[IMAGE:cfe1f3d865067824_12_94]] in the first hidden layer to neuron [[IMAGE:cfe1f3d865067824_12_95]] in the second hidden layer. At some timestep [[IMAGE:cfe1f3d865067824_12_96]] , the weight matrices are as follows [[IMAGE:cfe1f3d865067824_12_97]] [[IMAGE:cfe1f3d865067824_12_98]] Let [[IMAGE:cfe1f3d865067824_12_99]] denote the set containing the weight matrices [[IMAGE:cfe1f3d865067824_12_100]] and [[IMAGE:cfe1f3d865067824_12_101]] . Let [[IMAGE:cfe1f3d865067824_12_102]] represent the set of weight matrices [[IMAGE:cfe1f3d865067824_12_103]] and [[IMAGE:cfe1f3d865067824_12_104]] , and let [[IMAGE:cfe1f3d865067824_12_105]] denote the set consisting of the weight matrix [[IMAGE:cfe1f3d865067824_12_106]] . Recall the dueling network formulation used in the programming assignment: [[IMAGE:cfe1f3d865067824_12_107]] We are interested in computing Q-values for the following state: [[IMAGE:cfe1f3d865067824_12_108]] Based on the above data, answer the given
          Compute the **state-value** [[IMAGE:cfe1f3d865067824_13_109]] for state [[IMAGE:cfe1f3d865067824_13_110]] .
          Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

            A published solution is not available for this question yet.

            Question 13 NAT · 3.0 marks

            Consider the following Dueling DQN architecture given below that approximates the state-value function ( [[IMAGE:cfe1f3d865067824_11_74]] ) and advantage for actions A and B for states [[IMAGE:cfe1f3d865067824_11_75]] . The hidden layer nodes (grey colored nodes) employ the RELU activation function. [[IMAGE:cfe1f3d865067824_11_76]] The weight matrices of the network are given by: [[IMAGE:cfe1f3d865067824_11_77]] Here, [[IMAGE:cfe1f3d865067824_11_78]] contains the weights connecting the **input layer** to the **first hidden layer**. Similarly, [[IMAGE:cfe1f3d865067824_11_79]] contains the weights connecting the **first hidden layer** [[IMAGE:cfe1f3d865067824_11_80]] to the **second hidden** **layer** [[IMAGE:cfe1f3d865067824_11_81]] . [[IMAGE:cfe1f3d865067824_12_82]] [[IMAGE:cfe1f3d865067824_12_83]] contains the weights connecting the [[IMAGE:cfe1f3d865067824_12_84]] to the [[IMAGE:cfe1f3d865067824_12_85]] . [[IMAGE:cfe1f3d865067824_12_86]] contains the weights connecting the layer [[IMAGE:cfe1f3d865067824_12_87]] to the layer [[IMAGE:cfe1f3d865067824_12_88]] . Lastly, [[IMAGE:cfe1f3d865067824_12_89]] contains the weights connecting the layer [[IMAGE:cfe1f3d865067824_12_90]] to the layer [[IMAGE:cfe1f3d865067824_12_91]] . Each entry in these matrices corresponds to a single scalar weight. For example, the entry [[IMAGE:cfe1f3d865067824_12_92]] in [[IMAGE:cfe1f3d865067824_12_93]] denotes the weight from neuron [[IMAGE:cfe1f3d865067824_12_94]] in the first hidden layer to neuron [[IMAGE:cfe1f3d865067824_12_95]] in the second hidden layer. At some timestep [[IMAGE:cfe1f3d865067824_12_96]] , the weight matrices are as follows [[IMAGE:cfe1f3d865067824_12_97]] [[IMAGE:cfe1f3d865067824_12_98]] Let [[IMAGE:cfe1f3d865067824_12_99]] denote the set containing the weight matrices [[IMAGE:cfe1f3d865067824_12_100]] and [[IMAGE:cfe1f3d865067824_12_101]] . Let [[IMAGE:cfe1f3d865067824_12_102]] represent the set of weight matrices [[IMAGE:cfe1f3d865067824_12_103]] and [[IMAGE:cfe1f3d865067824_12_104]] , and let [[IMAGE:cfe1f3d865067824_12_105]] denote the set consisting of the weight matrix [[IMAGE:cfe1f3d865067824_12_106]] . Recall the dueling network formulation used in the programming assignment: [[IMAGE:cfe1f3d865067824_12_107]] We are interested in computing Q-values for the following state: [[IMAGE:cfe1f3d865067824_12_108]] Based on the above data, answer the given
            In order to compute Q-values, we need the advantage of action A and action B for state [[IMAGE:cfe1f3d865067824_13_111]] ? Compute the following advantage vector [[IMAGE:cfe1f3d865067824_13_112]] What is the value of [[IMAGE:cfe1f3d865067824_13_113]] + [[IMAGE:cfe1f3d865067824_13_114]] ?
            Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

              A published solution is not available for this question yet.

              Question 14 NAT · 3.0 marks

              Consider the following Dueling DQN architecture given below that approximates the state-value function ( [[IMAGE:cfe1f3d865067824_11_74]] ) and advantage for actions A and B for states [[IMAGE:cfe1f3d865067824_11_75]] . The hidden layer nodes (grey colored nodes) employ the RELU activation function. [[IMAGE:cfe1f3d865067824_11_76]] The weight matrices of the network are given by: [[IMAGE:cfe1f3d865067824_11_77]] Here, [[IMAGE:cfe1f3d865067824_11_78]] contains the weights connecting the **input layer** to the **first hidden layer**. Similarly, [[IMAGE:cfe1f3d865067824_11_79]] contains the weights connecting the **first hidden layer** [[IMAGE:cfe1f3d865067824_11_80]] to the **second hidden** **layer** [[IMAGE:cfe1f3d865067824_11_81]] . [[IMAGE:cfe1f3d865067824_12_82]] [[IMAGE:cfe1f3d865067824_12_83]] contains the weights connecting the [[IMAGE:cfe1f3d865067824_12_84]] to the [[IMAGE:cfe1f3d865067824_12_85]] . [[IMAGE:cfe1f3d865067824_12_86]] contains the weights connecting the layer [[IMAGE:cfe1f3d865067824_12_87]] to the layer [[IMAGE:cfe1f3d865067824_12_88]] . Lastly, [[IMAGE:cfe1f3d865067824_12_89]] contains the weights connecting the layer [[IMAGE:cfe1f3d865067824_12_90]] to the layer [[IMAGE:cfe1f3d865067824_12_91]] . Each entry in these matrices corresponds to a single scalar weight. For example, the entry [[IMAGE:cfe1f3d865067824_12_92]] in [[IMAGE:cfe1f3d865067824_12_93]] denotes the weight from neuron [[IMAGE:cfe1f3d865067824_12_94]] in the first hidden layer to neuron [[IMAGE:cfe1f3d865067824_12_95]] in the second hidden layer. At some timestep [[IMAGE:cfe1f3d865067824_12_96]] , the weight matrices are as follows [[IMAGE:cfe1f3d865067824_12_97]] [[IMAGE:cfe1f3d865067824_12_98]] Let [[IMAGE:cfe1f3d865067824_12_99]] denote the set containing the weight matrices [[IMAGE:cfe1f3d865067824_12_100]] and [[IMAGE:cfe1f3d865067824_12_101]] . Let [[IMAGE:cfe1f3d865067824_12_102]] represent the set of weight matrices [[IMAGE:cfe1f3d865067824_12_103]] and [[IMAGE:cfe1f3d865067824_12_104]] , and let [[IMAGE:cfe1f3d865067824_12_105]] denote the set consisting of the weight matrix [[IMAGE:cfe1f3d865067824_12_106]] . Recall the dueling network formulation used in the programming assignment: [[IMAGE:cfe1f3d865067824_12_107]] We are interested in computing Q-values for the following state: [[IMAGE:cfe1f3d865067824_12_108]] Based on the above data, answer the given
              Compute the Q-value of action A and action B for state [[IMAGE:cfe1f3d865067824_13_115]] . What is the value of [[IMAGE:cfe1f3d865067824_13_116]] + [[IMAGE:cfe1f3d865067824_14_117]] ?
              Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

                A published solution is not available for this question yet.

                Question 15 MCQ · 1.0 marks

                Consider the following navigation problem: [[IMAGE:cfe1f3d865067824_15_118]] The agent is spawned at the Start state, and has 4 actions on each state (Up, Down, Left, Right). Teleport to X is a special state; if the agent lands on this state, it is **instantly** teleported to the state marked X. Every action from every state, excluding the End state, incurs a reward of -1 (Episode terminates once the agent reaches the End state) A hierarchical learning approach is proposed, with the following hierarchy: [[IMAGE:cfe1f3d865067824_15_119]] where zone A and B correspond to rows 2-4 and 5-7 on the grid, respectively. Based on the above data, answer the given
                Which of the routes illustrated on the grid is taken when the Hierarchically Optimal policy is executed?
                Source diagram or notationSource diagram or notation
                1. Route 1
                2. Route 2
                3. Route 3

                A published solution is not available for this question yet.

                Question 16 MCQ · 1.0 marks

                Consider the following navigation problem: [[IMAGE:cfe1f3d865067824_15_118]] The agent is spawned at the Start state, and has 4 actions on each state (Up, Down, Left, Right). Teleport to X is a special state; if the agent lands on this state, it is **instantly** teleported to the state marked X. Every action from every state, excluding the End state, incurs a reward of -1 (Episode terminates once the agent reaches the End state) A hierarchical learning approach is proposed, with the following hierarchy: [[IMAGE:cfe1f3d865067824_15_119]] where zone A and B correspond to rows 2-4 and 5-7 on the grid, respectively. Based on the above data, answer the given
                Which of the routes illustrated on the grid is taken when the Recursively Optimal policy is executed?
                Source diagram or notationSource diagram or notation
                1. Route 1
                2. Route 2
                3. Route 3

                A published solution is not available for this question yet.

                Question 17 NAT · 1.0 marks

                Consider the following navigation problem: [[IMAGE:cfe1f3d865067824_15_118]] The agent is spawned at the Start state, and has 4 actions on each state (Up, Down, Left, Right). Teleport to X is a special state; if the agent lands on this state, it is **instantly** teleported to the state marked X. Every action from every state, excluding the End state, incurs a reward of -1 (Episode terminates once the agent reaches the End state) A hierarchical learning approach is proposed, with the following hierarchy: [[IMAGE:cfe1f3d865067824_15_119]] where zone A and B correspond to rows 2-4 and 5-7 on the grid, respectively. Based on the above data, answer the given
                Calculate the number of time-steps taken to finish the episode when the flat optimal policy is executed.
                Source diagram or notationSource diagram or notation

                  A published solution is not available for this question yet.