da5007_2025T3_ET_FN.pdf
Reinforcement Learning · End Term · Sep 2025 FN
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 MCQ · 2.0 marks
Consider a discounted return
[[IMAGE:cfe1f3d865067824_2_2]]
in an infinite-horizon MDP with bounded rewards [[IMAGE:cfe1f3d865067824_2_3]] and discount factor [[IMAGE:cfe1f3d865067824_2_4]] .
Which of the following statements is/are correct?
1. For fixed rewards, the contribution of [[IMAGE:cfe1f3d865067824_2_5]] to [[IMAGE:cfe1f3d865067824_2_6]] decays geometrically as [[IMAGE:cfe1f3d865067824_2_7]] .
2. If [[IMAGE:cfe1f3d865067824_2_8]] and rewards are uniformly bounded, the infinite sum [[IMAGE:cfe1f3d865067824_2_9]] is always finite.
3. For [[IMAGE:cfe1f3d865067824_2_10]] , the discounted return [[IMAGE:cfe1f3d865067824_2_11]] is bounded above in magnitude by [[IMAGE:cfe1f3d865067824_2_12]] .











1, 2, 3
1, 3
2, 3
only 1
only 3
A published solution is not available for this question yet.
Question 3 MCQ · 1.0 marks
In the UCT (Upper Confidence bounds applied to Trees) variant of Monte Carlo Tree Search, the
tree policy at a state [[IMAGE:cfe1f3d865067824_2_13]] typically selects an action [[IMAGE:cfe1f3d865067824_2_14]] by maximising


A pure exploitation term based only on the estimated value [[IMAGE:cfe1f3d865067824_2_15]]

A pure exploration term based only on visit counts [[IMAGE:cfe1f3d865067824_2_16]]

A combination of estimated value and an exploration bonus based on visit
counts
A random choice among all child actions
A published solution is not available for this question yet.
Question 4 MCQ · 2.0 marks
Consider the following target equations:
**Equation 1**
[[IMAGE:cfe1f3d865067824_3_17]]
**Equation 2**
[[IMAGE:cfe1f3d865067824_3_18]]
In both equations, [[IMAGE:cfe1f3d865067824_3_19]] refers to the shared parameters between the state-value network and
advantage network. [[IMAGE:cfe1f3d865067824_3_20]] and [[IMAGE:cfe1f3d865067824_3_21]] refer to the non-shared parameters of the state-value and
advantage network, respectively.
In both equations, [[IMAGE:cfe1f3d865067824_3_22]] is calculated as follows:
[[IMAGE:cfe1f3d865067824_3_23]]
Note: [[IMAGE:cfe1f3d865067824_3_24]] are target network parameters for online network parameters [[IMAGE:cfe1f3d865067824_3_25]]
Which among the following are correct statements?









Target in Equation 1 is characteristic of Dueling DQN algorithm and target in
equation 2 is characteristic of Dueling DDQN alogrithm
Target in Equation 2 is characteristic of DQN algorithm and target in equation
2 is characteristic of Double DQN alogrithm
Target in Equation 2 is characteristic of the DQN algorithm, and the target in
Equation 2 is characteristic of the Dueling DDQN alogrithm
Target in Equation 2 is characteristic of the Double Q-learning algorithm and
the target in Equation 2 is characteristic of DDQN
A published solution is not available for this question yet.
Question 5 MSQ · 2.0 marks
Consider the REINFORCE update with baseline for a stochastic policy [[IMAGE:cfe1f3d865067824_4_26]] :
[[IMAGE:cfe1f3d865067824_4_27]]
Which of the following statements are true?


Using a state-dependent baseline [[IMAGE:cfe1f3d865067824_4_28]] can reduce the variance of the gradient
estimate without introducing bias.

Choosing [[IMAGE:cfe1f3d865067824_4_29]] makes the update equivalent to using an advantage
function.

If [[IMAGE:cfe1f3d865067824_4_30]] depends on the action [[IMAGE:cfe1f3d865067824_4_31]] , the gradient estimate can become biased.


Any baseline [[IMAGE:cfe1f3d865067824_4_32]] that is independent of [[IMAGE:cfe1f3d865067824_4_33]] will necessarily make the gradient
estimate unbiased.


A published solution is not available for this question yet.
Question 6 NAT · 2.0 marks
Consider following MDP, assume [[IMAGE:cfe1f3d865067824_5_34]] :
[[IMAGE:cfe1f3d865067824_5_35]]
The edges have the value [[IMAGE:cfe1f3d865067824_5_36]] , where [[IMAGE:cfe1f3d865067824_5_37]] denotes the transition probability and [[IMAGE:cfe1f3d865067824_5_38]] is the immediate
expected reward. [[IMAGE:cfe1f3d865067824_5_39]] are non-terminal states and [[IMAGE:cfe1f3d865067824_5_40]] is a terminal state.
Based on the above data, answer the given
What is the probability that the accumulating eligibility trace for state [[IMAGE:cfe1f3d865067824_5_41]] is more than 3 at the end
of a trajectory?
**(Write your answer up to 4 decimal places)**








A published solution is not available for this question yet.
Question 7 NAT · 2.0 marks
Consider following MDP, assume [[IMAGE:cfe1f3d865067824_5_34]] :
[[IMAGE:cfe1f3d865067824_5_35]]
The edges have the value [[IMAGE:cfe1f3d865067824_5_36]] , where [[IMAGE:cfe1f3d865067824_5_37]] denotes the transition probability and [[IMAGE:cfe1f3d865067824_5_38]] is the immediate
expected reward. [[IMAGE:cfe1f3d865067824_5_39]] are non-terminal states and [[IMAGE:cfe1f3d865067824_5_40]] is a terminal state.
Based on the above data, answer the given
Consider the following trajectory that the agent follows:
[[IMAGE:cfe1f3d865067824_5_42]]
What will the eligibility trace of state [[IMAGE:cfe1f3d865067824_5_43]] just before the episode terminates?**(Write your answer**
**up to 4 decimal places)**









A published solution is not available for this question yet.
Question 8 NAT · 2.0 marks
Consider a 2D navigation environment (as shown in the figure below) where an autonomous
cleaning robot must locate and collect trash items (T1 and T2) scattered in the environment and
then deliver them to a designated disposal bin (G). The environment is shaped by curved, contour-
like walls, forming nested spaces with only one narrow passage connecting them, which requires
deliberate path planning.
[[IMAGE:cfe1f3d865067824_7_44]]
In this environment, the robot takes discrete actions from its action space [[IMAGE:cfe1f3d865067824_7_45]] ,
[[IMAGE:cfe1f3d865067824_7_46]] { NORTH, SOUTH, EAST, WEST }
Each action attempts to move the robot by [[IMAGE:cfe1f3d865067824_7_47]] cm in the corresponding direction. If the action
would cause a collision with a wall boundary, the robot remains in the same position and receives
a penalty.
The grey regions represent potholes. If the robot ends up in one, it cannot escape on its own and
requires human help, leading to a substantial penalty and relocation to the nearest safe position.
When the agent reaches a trash item, it automatically collects it and receives a reward. Likewise,
upon reaching a disposal bin, it automatically unloads the collected trash. Therefore, explicit
pickup and drop actions are not required.
The episode terminates either when the agent reaches the Goal (G) while carrying at least one
piece of trash, or when it has fallen into potholes a total of 10 times.
The position of the Robot ( [[IMAGE:cfe1f3d865067824_7_48]] ) can be represented as a tuple [[IMAGE:cfe1f3d865067824_7_49]] , where [[IMAGE:cfe1f3d865067824_7_50]] and [[IMAGE:cfe1f3d865067824_7_51]] are the distances
of the robot from the bottom-left corner of the environment.
For state representation, we would consider state discretisation by performing tile as well as
coarse coding, which converts [[IMAGE:cfe1f3d865067824_8_52]] into [[IMAGE:cfe1f3d865067824_8_53]] .
[[IMAGE:cfe1f3d865067824_8_54]]
Observe that there is 1 grid-tiling and 1 coarse-tiling here:
[[IMAGE:cfe1f3d865067824_8_55]]
[[IMAGE:cfe1f3d865067824_8_56]]
Based on the above data, answer the given
Using the coded state representation ( [[IMAGE:cfe1f3d865067824_8_57]] ) as input to a DQN with two hidden layers and a fully
connected output layer, what should be the dimensionality of the network’s input layer?














A published solution is not available for this question yet.
Question 9 NAT · 2.0 marks
Consider a 2D navigation environment (as shown in the figure below) where an autonomous
cleaning robot must locate and collect trash items (T1 and T2) scattered in the environment and
then deliver them to a designated disposal bin (G). The environment is shaped by curved, contour-
like walls, forming nested spaces with only one narrow passage connecting them, which requires
deliberate path planning.
[[IMAGE:cfe1f3d865067824_7_44]]
In this environment, the robot takes discrete actions from its action space [[IMAGE:cfe1f3d865067824_7_45]] ,
[[IMAGE:cfe1f3d865067824_7_46]] { NORTH, SOUTH, EAST, WEST }
Each action attempts to move the robot by [[IMAGE:cfe1f3d865067824_7_47]] cm in the corresponding direction. If the action
would cause a collision with a wall boundary, the robot remains in the same position and receives
a penalty.
The grey regions represent potholes. If the robot ends up in one, it cannot escape on its own and
requires human help, leading to a substantial penalty and relocation to the nearest safe position.
When the agent reaches a trash item, it automatically collects it and receives a reward. Likewise,
upon reaching a disposal bin, it automatically unloads the collected trash. Therefore, explicit
pickup and drop actions are not required.
The episode terminates either when the agent reaches the Goal (G) while carrying at least one
piece of trash, or when it has fallen into potholes a total of 10 times.
The position of the Robot ( [[IMAGE:cfe1f3d865067824_7_48]] ) can be represented as a tuple [[IMAGE:cfe1f3d865067824_7_49]] , where [[IMAGE:cfe1f3d865067824_7_50]] and [[IMAGE:cfe1f3d865067824_7_51]] are the distances
of the robot from the bottom-left corner of the environment.
For state representation, we would consider state discretisation by performing tile as well as
coarse coding, which converts [[IMAGE:cfe1f3d865067824_8_52]] into [[IMAGE:cfe1f3d865067824_8_53]] .
[[IMAGE:cfe1f3d865067824_8_54]]
Observe that there is 1 grid-tiling and 1 coarse-tiling here:
[[IMAGE:cfe1f3d865067824_8_55]]
[[IMAGE:cfe1f3d865067824_8_56]]
Based on the above data, answer the given
Using the coded state representation ( [[IMAGE:cfe1f3d865067824_9_58]] ) as input to a DQN with two hidden layers and a fully
connected output layer, what should be the size of the output layer?














A published solution is not available for this question yet.
Question 10 MSQ · 2.0 marks
Consider a 2D navigation environment (as shown in the figure below) where an autonomous
cleaning robot must locate and collect trash items (T1 and T2) scattered in the environment and
then deliver them to a designated disposal bin (G). The environment is shaped by curved, contour-
like walls, forming nested spaces with only one narrow passage connecting them, which requires
deliberate path planning.
[[IMAGE:cfe1f3d865067824_7_44]]
In this environment, the robot takes discrete actions from its action space [[IMAGE:cfe1f3d865067824_7_45]] ,
[[IMAGE:cfe1f3d865067824_7_46]] { NORTH, SOUTH, EAST, WEST }
Each action attempts to move the robot by [[IMAGE:cfe1f3d865067824_7_47]] cm in the corresponding direction. If the action
would cause a collision with a wall boundary, the robot remains in the same position and receives
a penalty.
The grey regions represent potholes. If the robot ends up in one, it cannot escape on its own and
requires human help, leading to a substantial penalty and relocation to the nearest safe position.
When the agent reaches a trash item, it automatically collects it and receives a reward. Likewise,
upon reaching a disposal bin, it automatically unloads the collected trash. Therefore, explicit
pickup and drop actions are not required.
The episode terminates either when the agent reaches the Goal (G) while carrying at least one
piece of trash, or when it has fallen into potholes a total of 10 times.
The position of the Robot ( [[IMAGE:cfe1f3d865067824_7_48]] ) can be represented as a tuple [[IMAGE:cfe1f3d865067824_7_49]] , where [[IMAGE:cfe1f3d865067824_7_50]] and [[IMAGE:cfe1f3d865067824_7_51]] are the distances
of the robot from the bottom-left corner of the environment.
For state representation, we would consider state discretisation by performing tile as well as
coarse coding, which converts [[IMAGE:cfe1f3d865067824_8_52]] into [[IMAGE:cfe1f3d865067824_8_53]] .
[[IMAGE:cfe1f3d865067824_8_54]]
Observe that there is 1 grid-tiling and 1 coarse-tiling here:
[[IMAGE:cfe1f3d865067824_8_55]]
[[IMAGE:cfe1f3d865067824_8_56]]
Based on the above data, answer the given
In the given environment, the state space is continuous while the action space is discrete. Now
consider a similar environment where the action space is also continuous, with the action defined
as the robot’s rotation angle relative to the positive horizontal axis.
You decide to use a **Deterministic Policy Gradient (DPG)** method with a deterministic policy
[[IMAGE:cfe1f3d865067824_9_59]] .
Which of the following statements are valid for ensuring adequate exploration and obtaining
unbiased policy-gradient estimates?














In a deterministic but stochastic environment, the environment’s own
randomness can provide sufficient exploration, so an explicit stochastic behaviour policy is not
strictly necessary for exploration.
In a deterministic and non-stochastic environment, to ensure exploration we
can use an off-policy actor–critic setup, where the behaviour policy [[IMAGE:cfe1f3d865067824_9_60]] is more exploratory
(e.g., [[IMAGE:cfe1f3d865067824_9_61]] , where [[IMAGE:cfe1f3d865067824_9_62]] is small noise) than the deterministic target policy [[IMAGE:cfe1f3d865067824_9_63]] .




In off-policy DPG, the policy gradient is taken with respect to the behaviour
policy [[IMAGE:cfe1f3d865067824_9_64]] , since it is the one generating the data, rather than with respect to the deterministic
target policy [[IMAGE:cfe1f3d865067824_9_65]] .


For DPG with a function-approximated critic [[IMAGE:cfe1f3d865067824_9_66]] , using a compatible
parametrisation between [[IMAGE:cfe1f3d865067824_9_67]] and [[IMAGE:cfe1f3d865067824_9_68]] can help ensure that the deterministic policy gradient
estimate remains unbiased, even when [[IMAGE:cfe1f3d865067824_10_69]] is approximate.




A published solution is not available for this question yet.
Question 11 MCQ · 2.0 marks
Consider a 2D navigation environment (as shown in the figure below) where an autonomous
cleaning robot must locate and collect trash items (T1 and T2) scattered in the environment and
then deliver them to a designated disposal bin (G). The environment is shaped by curved, contour-
like walls, forming nested spaces with only one narrow passage connecting them, which requires
deliberate path planning.
[[IMAGE:cfe1f3d865067824_7_44]]
In this environment, the robot takes discrete actions from its action space [[IMAGE:cfe1f3d865067824_7_45]] ,
[[IMAGE:cfe1f3d865067824_7_46]] { NORTH, SOUTH, EAST, WEST }
Each action attempts to move the robot by [[IMAGE:cfe1f3d865067824_7_47]] cm in the corresponding direction. If the action
would cause a collision with a wall boundary, the robot remains in the same position and receives
a penalty.
The grey regions represent potholes. If the robot ends up in one, it cannot escape on its own and
requires human help, leading to a substantial penalty and relocation to the nearest safe position.
When the agent reaches a trash item, it automatically collects it and receives a reward. Likewise,
upon reaching a disposal bin, it automatically unloads the collected trash. Therefore, explicit
pickup and drop actions are not required.
The episode terminates either when the agent reaches the Goal (G) while carrying at least one
piece of trash, or when it has fallen into potholes a total of 10 times.
The position of the Robot ( [[IMAGE:cfe1f3d865067824_7_48]] ) can be represented as a tuple [[IMAGE:cfe1f3d865067824_7_49]] , where [[IMAGE:cfe1f3d865067824_7_50]] and [[IMAGE:cfe1f3d865067824_7_51]] are the distances
of the robot from the bottom-left corner of the environment.
For state representation, we would consider state discretisation by performing tile as well as
coarse coding, which converts [[IMAGE:cfe1f3d865067824_8_52]] into [[IMAGE:cfe1f3d865067824_8_53]] .
[[IMAGE:cfe1f3d865067824_8_54]]
Observe that there is 1 grid-tiling and 1 coarse-tiling here:
[[IMAGE:cfe1f3d865067824_8_55]]
[[IMAGE:cfe1f3d865067824_8_56]]
Based on the above data, answer the given
You have decided to use the Dueling DQN architecture with the aggregation.
[[IMAGE:cfe1f3d865067824_10_70]]
What is the primary purpose of subtracting the term [[IMAGE:cfe1f3d865067824_10_71]] in this formulation?















To clip the magnitude of the advantage values and thereby prevent exploding
Q-values.
To guarantee that the action with the largest advantage always has a strictly
positive Q-value.
To force the average advantage over all actions in a state to be zero, making
the decomposition into [[IMAGE:cfe1f3d865067824_10_72]] and [[IMAGE:cfe1f3d865067824_10_73]] well-defined.


To allow the network to represent any Q-function using only the advantage
stream, making the value stream redundant.
A published solution is not available for this question yet.
Question 12 NAT · 2.0 marks
Consider the following Dueling DQN architecture given below that approximates the state-value
function ( [[IMAGE:cfe1f3d865067824_11_74]] ) and advantage for actions A and B for states [[IMAGE:cfe1f3d865067824_11_75]] . The hidden layer nodes (grey
colored nodes) employ the RELU activation function.
[[IMAGE:cfe1f3d865067824_11_76]]
The weight matrices of the network are given by:
[[IMAGE:cfe1f3d865067824_11_77]]
Here, [[IMAGE:cfe1f3d865067824_11_78]] contains the weights connecting the **input layer** to the **first hidden layer**.
Similarly, [[IMAGE:cfe1f3d865067824_11_79]] contains the weights connecting the **first hidden layer** [[IMAGE:cfe1f3d865067824_11_80]] to the **second hidden**
**layer** [[IMAGE:cfe1f3d865067824_11_81]] .
[[IMAGE:cfe1f3d865067824_12_82]]
[[IMAGE:cfe1f3d865067824_12_83]] contains the weights connecting the [[IMAGE:cfe1f3d865067824_12_84]] to the [[IMAGE:cfe1f3d865067824_12_85]] .
[[IMAGE:cfe1f3d865067824_12_86]] contains the weights connecting the layer [[IMAGE:cfe1f3d865067824_12_87]] to the layer [[IMAGE:cfe1f3d865067824_12_88]] .
Lastly, [[IMAGE:cfe1f3d865067824_12_89]] contains the weights connecting the layer [[IMAGE:cfe1f3d865067824_12_90]] to the layer [[IMAGE:cfe1f3d865067824_12_91]] .
Each entry in these matrices corresponds to a single scalar weight. For example, the entry [[IMAGE:cfe1f3d865067824_12_92]]
in [[IMAGE:cfe1f3d865067824_12_93]] denotes the weight from neuron [[IMAGE:cfe1f3d865067824_12_94]] in the first hidden layer to neuron [[IMAGE:cfe1f3d865067824_12_95]] in the second
hidden layer.
At some timestep [[IMAGE:cfe1f3d865067824_12_96]] , the weight matrices are as follows
[[IMAGE:cfe1f3d865067824_12_97]]
[[IMAGE:cfe1f3d865067824_12_98]]
Let [[IMAGE:cfe1f3d865067824_12_99]] denote the set containing the weight matrices [[IMAGE:cfe1f3d865067824_12_100]] and [[IMAGE:cfe1f3d865067824_12_101]] .
Let [[IMAGE:cfe1f3d865067824_12_102]] represent the set of weight matrices [[IMAGE:cfe1f3d865067824_12_103]] and [[IMAGE:cfe1f3d865067824_12_104]] , and let [[IMAGE:cfe1f3d865067824_12_105]] denote the set consisting of the
weight matrix [[IMAGE:cfe1f3d865067824_12_106]] .
Recall the dueling network formulation used in the programming assignment:
[[IMAGE:cfe1f3d865067824_12_107]]
We are interested in computing Q-values for the following state:
[[IMAGE:cfe1f3d865067824_12_108]]
Based on the above data, answer the given
Compute the **state-value** [[IMAGE:cfe1f3d865067824_13_109]] for state [[IMAGE:cfe1f3d865067824_13_110]] .





































A published solution is not available for this question yet.
Question 13 NAT · 3.0 marks
Consider the following Dueling DQN architecture given below that approximates the state-value
function ( [[IMAGE:cfe1f3d865067824_11_74]] ) and advantage for actions A and B for states [[IMAGE:cfe1f3d865067824_11_75]] . The hidden layer nodes (grey
colored nodes) employ the RELU activation function.
[[IMAGE:cfe1f3d865067824_11_76]]
The weight matrices of the network are given by:
[[IMAGE:cfe1f3d865067824_11_77]]
Here, [[IMAGE:cfe1f3d865067824_11_78]] contains the weights connecting the **input layer** to the **first hidden layer**.
Similarly, [[IMAGE:cfe1f3d865067824_11_79]] contains the weights connecting the **first hidden layer** [[IMAGE:cfe1f3d865067824_11_80]] to the **second hidden**
**layer** [[IMAGE:cfe1f3d865067824_11_81]] .
[[IMAGE:cfe1f3d865067824_12_82]]
[[IMAGE:cfe1f3d865067824_12_83]] contains the weights connecting the [[IMAGE:cfe1f3d865067824_12_84]] to the [[IMAGE:cfe1f3d865067824_12_85]] .
[[IMAGE:cfe1f3d865067824_12_86]] contains the weights connecting the layer [[IMAGE:cfe1f3d865067824_12_87]] to the layer [[IMAGE:cfe1f3d865067824_12_88]] .
Lastly, [[IMAGE:cfe1f3d865067824_12_89]] contains the weights connecting the layer [[IMAGE:cfe1f3d865067824_12_90]] to the layer [[IMAGE:cfe1f3d865067824_12_91]] .
Each entry in these matrices corresponds to a single scalar weight. For example, the entry [[IMAGE:cfe1f3d865067824_12_92]]
in [[IMAGE:cfe1f3d865067824_12_93]] denotes the weight from neuron [[IMAGE:cfe1f3d865067824_12_94]] in the first hidden layer to neuron [[IMAGE:cfe1f3d865067824_12_95]] in the second
hidden layer.
At some timestep [[IMAGE:cfe1f3d865067824_12_96]] , the weight matrices are as follows
[[IMAGE:cfe1f3d865067824_12_97]]
[[IMAGE:cfe1f3d865067824_12_98]]
Let [[IMAGE:cfe1f3d865067824_12_99]] denote the set containing the weight matrices [[IMAGE:cfe1f3d865067824_12_100]] and [[IMAGE:cfe1f3d865067824_12_101]] .
Let [[IMAGE:cfe1f3d865067824_12_102]] represent the set of weight matrices [[IMAGE:cfe1f3d865067824_12_103]] and [[IMAGE:cfe1f3d865067824_12_104]] , and let [[IMAGE:cfe1f3d865067824_12_105]] denote the set consisting of the
weight matrix [[IMAGE:cfe1f3d865067824_12_106]] .
Recall the dueling network formulation used in the programming assignment:
[[IMAGE:cfe1f3d865067824_12_107]]
We are interested in computing Q-values for the following state:
[[IMAGE:cfe1f3d865067824_12_108]]
Based on the above data, answer the given
In order to compute Q-values, we need the advantage of action A and action B for state [[IMAGE:cfe1f3d865067824_13_111]] ?
Compute the following advantage vector
[[IMAGE:cfe1f3d865067824_13_112]]
What is the value of [[IMAGE:cfe1f3d865067824_13_113]] + [[IMAGE:cfe1f3d865067824_13_114]] ?







































A published solution is not available for this question yet.
Question 14 NAT · 3.0 marks
Consider the following Dueling DQN architecture given below that approximates the state-value
function ( [[IMAGE:cfe1f3d865067824_11_74]] ) and advantage for actions A and B for states [[IMAGE:cfe1f3d865067824_11_75]] . The hidden layer nodes (grey
colored nodes) employ the RELU activation function.
[[IMAGE:cfe1f3d865067824_11_76]]
The weight matrices of the network are given by:
[[IMAGE:cfe1f3d865067824_11_77]]
Here, [[IMAGE:cfe1f3d865067824_11_78]] contains the weights connecting the **input layer** to the **first hidden layer**.
Similarly, [[IMAGE:cfe1f3d865067824_11_79]] contains the weights connecting the **first hidden layer** [[IMAGE:cfe1f3d865067824_11_80]] to the **second hidden**
**layer** [[IMAGE:cfe1f3d865067824_11_81]] .
[[IMAGE:cfe1f3d865067824_12_82]]
[[IMAGE:cfe1f3d865067824_12_83]] contains the weights connecting the [[IMAGE:cfe1f3d865067824_12_84]] to the [[IMAGE:cfe1f3d865067824_12_85]] .
[[IMAGE:cfe1f3d865067824_12_86]] contains the weights connecting the layer [[IMAGE:cfe1f3d865067824_12_87]] to the layer [[IMAGE:cfe1f3d865067824_12_88]] .
Lastly, [[IMAGE:cfe1f3d865067824_12_89]] contains the weights connecting the layer [[IMAGE:cfe1f3d865067824_12_90]] to the layer [[IMAGE:cfe1f3d865067824_12_91]] .
Each entry in these matrices corresponds to a single scalar weight. For example, the entry [[IMAGE:cfe1f3d865067824_12_92]]
in [[IMAGE:cfe1f3d865067824_12_93]] denotes the weight from neuron [[IMAGE:cfe1f3d865067824_12_94]] in the first hidden layer to neuron [[IMAGE:cfe1f3d865067824_12_95]] in the second
hidden layer.
At some timestep [[IMAGE:cfe1f3d865067824_12_96]] , the weight matrices are as follows
[[IMAGE:cfe1f3d865067824_12_97]]
[[IMAGE:cfe1f3d865067824_12_98]]
Let [[IMAGE:cfe1f3d865067824_12_99]] denote the set containing the weight matrices [[IMAGE:cfe1f3d865067824_12_100]] and [[IMAGE:cfe1f3d865067824_12_101]] .
Let [[IMAGE:cfe1f3d865067824_12_102]] represent the set of weight matrices [[IMAGE:cfe1f3d865067824_12_103]] and [[IMAGE:cfe1f3d865067824_12_104]] , and let [[IMAGE:cfe1f3d865067824_12_105]] denote the set consisting of the
weight matrix [[IMAGE:cfe1f3d865067824_12_106]] .
Recall the dueling network formulation used in the programming assignment:
[[IMAGE:cfe1f3d865067824_12_107]]
We are interested in computing Q-values for the following state:
[[IMAGE:cfe1f3d865067824_12_108]]
Based on the above data, answer the given
Compute the Q-value of action A and action B for state [[IMAGE:cfe1f3d865067824_13_115]] . What is the value of [[IMAGE:cfe1f3d865067824_13_116]] +
[[IMAGE:cfe1f3d865067824_14_117]] ?






































A published solution is not available for this question yet.
Question 15 MCQ · 1.0 marks
Consider the following navigation problem:
[[IMAGE:cfe1f3d865067824_15_118]]
The agent is spawned at the Start state, and has 4 actions on each state (Up, Down, Left, Right). Teleport to X
is a special state; if the agent lands on this state, it is **instantly** teleported to the state marked X.
Every action from every state, excluding the End state, incurs a reward of -1 (Episode terminates
once the agent reaches the End state)
A hierarchical learning approach is proposed, with the following hierarchy:
[[IMAGE:cfe1f3d865067824_15_119]]
where zone A and B correspond to rows 2-4 and 5-7 on the grid, respectively.
Based on the above data, answer the given
Which of the routes illustrated on the grid is taken when the Hierarchically Optimal policy is
executed?


Route 1
Route 2
Route 3
A published solution is not available for this question yet.
Question 16 MCQ · 1.0 marks
Consider the following navigation problem:
[[IMAGE:cfe1f3d865067824_15_118]]
The agent is spawned at the Start state, and has 4 actions on each state (Up, Down, Left, Right). Teleport to X
is a special state; if the agent lands on this state, it is **instantly** teleported to the state marked X.
Every action from every state, excluding the End state, incurs a reward of -1 (Episode terminates
once the agent reaches the End state)
A hierarchical learning approach is proposed, with the following hierarchy:
[[IMAGE:cfe1f3d865067824_15_119]]
where zone A and B correspond to rows 2-4 and 5-7 on the grid, respectively.
Based on the above data, answer the given
Which of the routes illustrated on the grid is taken when the Recursively Optimal policy is
executed?


Route 1
Route 2
Route 3
A published solution is not available for this question yet.
Question 17 NAT · 1.0 marks
Consider the following navigation problem:
[[IMAGE:cfe1f3d865067824_15_118]]
The agent is spawned at the Start state, and has 4 actions on each state (Up, Down, Left, Right). Teleport to X
is a special state; if the agent lands on this state, it is **instantly** teleported to the state marked X.
Every action from every state, excluding the End state, incurs a reward of -1 (Episode terminates
once the agent reaches the End state)
A hierarchical learning approach is proposed, with the following hierarchy:
[[IMAGE:cfe1f3d865067824_15_119]]
where zone A and B correspond to rows 2-4 and 5-7 on the grid, respectively.
Based on the above data, answer the given
Calculate the number of time-steps taken to finish the episode when the flat optimal policy is
executed.


A published solution is not available for this question yet.