da5007_2026T1_Q2_NA.pdf
Reinforcement Learning · Quiz 2 · Jan 2026
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 MCQ · 1.0 marks
What is a key advantage of using n-step TD prediction over one-step TD prediction?
n-step TD prediction requires less memory and computational resources.
n-step TD prediction can handle delayed rewards and credit assignment over
multiple time steps
n-step TD prediction converges faster to the optimal policy.
n-step TD prediction guarantees convergence to the optimal value function for
any choice of the learning rate.
A published solution is not available for this question yet.
Question 3 MCQ · 1.0 marks
In which of the following scenarios is Expected SARSA a good fit?
When function approximation is necessary due to large state spaces, requiring
deep networks for learning representations
When learning needs to prioritize separating state value and advantage
functions for better decision-making
When the environment has a high degree of stochasticity, making
bootstrapping unstable in value-based methods
When experience replay is essential for stable learning and sample efficiency
A published solution is not available for this question yet.
Question 4 MCQ · 2.0 marks
Consider the following target equations:
**Equation 1**
[[IMAGE:a489e35127f63335_3_2]]
**Equation 2**
[[IMAGE:a489e35127f63335_3_3]]
In both equations, [[IMAGE:a489e35127f63335_3_4]] refers to the shared parameters between the state-value network and
advantage network. [[IMAGE:a489e35127f63335_3_5]] and [[IMAGE:a489e35127f63335_3_6]] refer to the non-shared parameters of the state-value and advantage
network, respectively.
In both equations, [[IMAGE:a489e35127f63335_3_7]] is calculated using the following network.
[[IMAGE:a489e35127f63335_3_8]]
Note: [[IMAGE:a489e35127f63335_3_9]] are target network parameters for online network parameters [[IMAGE:a489e35127f63335_3_10]]
Which among the following are correct statements?









Target in Equation 1 is characteristic of Dueling DQN algorithm, and the target
in Equation 2 is characteristic of Dueling Double DQN algorithm
Target in Equation 1 is characteristic of the DQN algorithm, and the target in
Equation 2 is characteristic of the Double DQN algorithm
Target in Equation 1 is characteristic of the DQN algorithm, and the target in
Equation 2 is characteristic of the Dueling Double DQN algorithm
Target in Equation 1 is characteristic of the Double Q-learning algorithm, and
the target in Equation 2 is characteristic of Double DQN
A published solution is not available for this question yet.
Question 5 MSQ · 1.0 marks
Which of the following methods are a form of Generalized Policy Iteration?
Policy Improvement
Value Iteration
Q-Learning and SARSA
Monte Carlo control methods
A published solution is not available for this question yet.
Question 6 MSQ · 2.0 marks
In value-based deep reinforcement learning, the max operator in Q-learning can lead to
systematic overestimation of Q-values when neural networks are used for function approximation
Which option best explains why vanilla Deep Q-Netwok (DQN) tends to produce overoptimistic Q-
value estimates?
The same neural approximator [[IMAGE:a489e35127f63335_4_11]] is used to select the best action
[[IMAGE:a489e35127f63335_4_12]] and to evaluate that action [[IMAGE:a489e35127f63335_4_13]] , causing the max operator to
preferentially propagate positive approximation errors across correlated action-value predictions



Delayed target network polyak-averaging [[IMAGE:a489e35127f63335_4_14]] introduces
temporal mismatch in Bellman targets [[IMAGE:a489e35127f63335_4_15]] , causing overestimation


DQN assumes a learned model of the environment [[IMAGE:a489e35127f63335_4_16]] , and errors in
this model cause biased targets.

DQN uses Monte Carlo returns [[IMAGE:a489e35127f63335_5_17]] instead of bootstrapped
temporal-difference targets

A published solution is not available for this question yet.
Question 7 NAT · 1.0 marks
Consider the following MDP:
[[IMAGE:a489e35127f63335_5_18]]
Evaluate the value function [[IMAGE:a489e35127f63335_5_19]] , where [[IMAGE:a489e35127f63335_5_20]] is the equiprobable random policy. What is
[[IMAGE:a489e35127f63335_5_21]] ? (Set [[IMAGE:a489e35127f63335_5_22]] )





A published solution is not available for this question yet.
Question 8 NAT · 1.0 marks
Consider a 2D navigation environment (as shown in the figure below) where an autonomous
cleaning robot must locate and collect trash items (T1 and T2) scattered in the environment and
then deliver them to a designated disposal bin (G). The environment is shaped by curved, contour-
like walls, forming nested spaces with only one narrow passage connecting them, which requires
deliberate path planning.
[[IMAGE:a489e35127f63335_6_23]]
In this environment, the robot takes discrete actions from its action space [[IMAGE:a489e35127f63335_6_24]] ,
[[IMAGE:a489e35127f63335_6_25]] { NORTH, SOUTH, EAST, WEST }
Each action attempts to move the robot by 0.1 cm in the corresponding direction. If the action
would cause a collision with a wall boundary, the robot remains in the same position and receives
a penalty.
The grey regions represent potholes. If the robot ends up in one, it cannot escape on its own and
requires human help, leading to a substantial penalty and relocation to the nearest safe position.
When the agent reaches a trash item, it automatically collects it and receives a reward. Likewise,
upon reaching a disposal bin, it automatically unloads the collected trash. Therefore, explicit
pickup and drop actions are not required.
The episode terminates either when the agent reaches the Goal (G) while carrying at least one
piece of trash, or when it has fallen into potholes a total of 10 times.
The position of the Robot ( [[IMAGE:a489e35127f63335_6_26]] ) can be represented as a tuple [[IMAGE:a489e35127f63335_6_27]] , where [[IMAGE:a489e35127f63335_6_28]] and [[IMAGE:a489e35127f63335_6_29]] are the distances
of the robot from the bottom-left corner of the environment.
For state representation, we would consider state discretisation by performing tile as well as
coarse coding, which converts [[IMAGE:a489e35127f63335_6_30]] into [[IMAGE:a489e35127f63335_6_31]] { [[IMAGE:a489e35127f63335_6_32]] } [[IMAGE:a489e35127f63335_6_33]] .
[[IMAGE:a489e35127f63335_6_34]]
[[IMAGE:a489e35127f63335_7_35]]
Observe that there are 3 different grid-tilings here:
[[IMAGE:a489e35127f63335_7_36]]
Based on the above data, answer the given subquestions.
Using the coded state representation ( [[IMAGE:a489e35127f63335_7_37]] ) as input to a DQN with two hidden layers and a fully
connected output layer, what should be the dimensionality of the network’s input layer?















A published solution is not available for this question yet.
Question 9 NAT · 1.0 marks
Consider a 2D navigation environment (as shown in the figure below) where an autonomous
cleaning robot must locate and collect trash items (T1 and T2) scattered in the environment and
then deliver them to a designated disposal bin (G). The environment is shaped by curved, contour-
like walls, forming nested spaces with only one narrow passage connecting them, which requires
deliberate path planning.
[[IMAGE:a489e35127f63335_6_23]]
In this environment, the robot takes discrete actions from its action space [[IMAGE:a489e35127f63335_6_24]] ,
[[IMAGE:a489e35127f63335_6_25]] { NORTH, SOUTH, EAST, WEST }
Each action attempts to move the robot by 0.1 cm in the corresponding direction. If the action
would cause a collision with a wall boundary, the robot remains in the same position and receives
a penalty.
The grey regions represent potholes. If the robot ends up in one, it cannot escape on its own and
requires human help, leading to a substantial penalty and relocation to the nearest safe position.
When the agent reaches a trash item, it automatically collects it and receives a reward. Likewise,
upon reaching a disposal bin, it automatically unloads the collected trash. Therefore, explicit
pickup and drop actions are not required.
The episode terminates either when the agent reaches the Goal (G) while carrying at least one
piece of trash, or when it has fallen into potholes a total of 10 times.
The position of the Robot ( [[IMAGE:a489e35127f63335_6_26]] ) can be represented as a tuple [[IMAGE:a489e35127f63335_6_27]] , where [[IMAGE:a489e35127f63335_6_28]] and [[IMAGE:a489e35127f63335_6_29]] are the distances
of the robot from the bottom-left corner of the environment.
For state representation, we would consider state discretisation by performing tile as well as
coarse coding, which converts [[IMAGE:a489e35127f63335_6_30]] into [[IMAGE:a489e35127f63335_6_31]] { [[IMAGE:a489e35127f63335_6_32]] } [[IMAGE:a489e35127f63335_6_33]] .
[[IMAGE:a489e35127f63335_6_34]]
[[IMAGE:a489e35127f63335_7_35]]
Observe that there are 3 different grid-tilings here:
[[IMAGE:a489e35127f63335_7_36]]
Based on the above data, answer the given subquestions.
Using the coded state representation ( [[IMAGE:a489e35127f63335_8_38]] ) as input to a DQN with two hidden layers and a fully
connected output layer, what should be the size of the output layer?















A published solution is not available for this question yet.
Question 10 MSQ · 1.0 marks
Consider a 2D navigation environment (as shown in the figure below) where an autonomous
cleaning robot must locate and collect trash items (T1 and T2) scattered in the environment and
then deliver them to a designated disposal bin (G). The environment is shaped by curved, contour-
like walls, forming nested spaces with only one narrow passage connecting them, which requires
deliberate path planning.
[[IMAGE:a489e35127f63335_6_23]]
In this environment, the robot takes discrete actions from its action space [[IMAGE:a489e35127f63335_6_24]] ,
[[IMAGE:a489e35127f63335_6_25]] { NORTH, SOUTH, EAST, WEST }
Each action attempts to move the robot by 0.1 cm in the corresponding direction. If the action
would cause a collision with a wall boundary, the robot remains in the same position and receives
a penalty.
The grey regions represent potholes. If the robot ends up in one, it cannot escape on its own and
requires human help, leading to a substantial penalty and relocation to the nearest safe position.
When the agent reaches a trash item, it automatically collects it and receives a reward. Likewise,
upon reaching a disposal bin, it automatically unloads the collected trash. Therefore, explicit
pickup and drop actions are not required.
The episode terminates either when the agent reaches the Goal (G) while carrying at least one
piece of trash, or when it has fallen into potholes a total of 10 times.
The position of the Robot ( [[IMAGE:a489e35127f63335_6_26]] ) can be represented as a tuple [[IMAGE:a489e35127f63335_6_27]] , where [[IMAGE:a489e35127f63335_6_28]] and [[IMAGE:a489e35127f63335_6_29]] are the distances
of the robot from the bottom-left corner of the environment.
For state representation, we would consider state discretisation by performing tile as well as
coarse coding, which converts [[IMAGE:a489e35127f63335_6_30]] into [[IMAGE:a489e35127f63335_6_31]] { [[IMAGE:a489e35127f63335_6_32]] } [[IMAGE:a489e35127f63335_6_33]] .
[[IMAGE:a489e35127f63335_6_34]]
[[IMAGE:a489e35127f63335_7_35]]
Observe that there are 3 different grid-tilings here:
[[IMAGE:a489e35127f63335_7_36]]
Based on the above data, answer the given subquestions.
You have decided to use the Dueling DQN architecture with the aggregation.
[[IMAGE:a489e35127f63335_8_39]]
What is the primary purpose of subtracting the term [[IMAGE:a489e35127f63335_8_40]] in this formulation?
















It encourages the network to always choose the action with the highest
advantage, thereby directly implementing a greedy policy inside the Q-network.
It normalises the advantage values for a given state so that at least one action
has zero advantage, reducing identifiability issues between [[IMAGE:a489e35127f63335_8_41]] and [[IMAGE:a489e35127f63335_8_42]] .


It rescales the advantages to lie in the range [[IMAGE:a489e35127f63335_8_43]] , making training more
numerically stable.

It removes the need for a separate value stream [[IMAGE:a489e35127f63335_8_44]] , since the max operator
can recover the state value from the advantages alone.

A published solution is not available for this question yet.
Question 11 MSQ · 1.0 marks
In Deep Q-Networks (DQN), a separate target network is maintained whose weights are periodically
copied from the main (online) network rather than updated at every step.
Based on the above data, answer the given subquestions.
What is the primary reason for using a separate target network with periodically updated weights
in DQN?
To accelerate convergence by increasing the learning rate.
To stabilize learning by keeping target values fixed for several steps.
To ensure the Q-values are always up to date.
To share weights between the actor and the critic.
A published solution is not available for this question yet.
Question 12 MSQ · 1.0 marks
In Deep Q-Networks (DQN), a separate target network is maintained whose weights are periodically
copied from the main (online) network rather than updated at every step.
Based on the above data, answer the given subquestions.
What are the effects of periodically updating the weights of the target network in DQN?
Provides stable targets to the main network during training.
Causes the learning targets to be non-stationary.
Mitigates instabilities in Q-learning updates.
Avoids the need for a replay buffer.
A published solution is not available for this question yet.
Question 13 NAT · 0.5 marks
Consider a discounted ( [[IMAGE:a489e35127f63335_9_45]] ), episodic task that has two non-terminal states, A and B. The
rewards belongs to { 0,1,2,3 } set. The following are some episodes experienced by an agent
following a fixed policy. The terminal state is not explicitly mentioned for any of the episodes.
[[IMAGE:a489e35127f63335_10_46]]
Based on the above data, answer the given subquestions.
What is the estimate of [[IMAGE:a489e35127f63335_10_47]] returned by first-visit MC?



A published solution is not available for this question yet.
Question 14 NAT · 0.5 marks
Consider a discounted ( [[IMAGE:a489e35127f63335_9_45]] ), episodic task that has two non-terminal states, A and B. The
rewards belongs to { 0,1,2,3 } set. The following are some episodes experienced by an agent
following a fixed policy. The terminal state is not explicitly mentioned for any of the episodes.
[[IMAGE:a489e35127f63335_10_46]]
Based on the above data, answer the given subquestions.
What is the estimate of [[IMAGE:a489e35127f63335_10_48]] returned by first-visit MC?



A published solution is not available for this question yet.
Question 15 NAT · 0.5 marks
Consider a discounted ( [[IMAGE:a489e35127f63335_9_45]] ), episodic task that has two non-terminal states, A and B. The
rewards belongs to { 0,1,2,3 } set. The following are some episodes experienced by an agent
following a fixed policy. The terminal state is not explicitly mentioned for any of the episodes.
[[IMAGE:a489e35127f63335_10_46]]
Based on the above data, answer the given subquestions.
What is the estimate of [[IMAGE:a489e35127f63335_11_49]] returned by batch [[IMAGE:a489e35127f63335_11_50]] ?




A published solution is not available for this question yet.
Question 16 NAT · 0.5 marks
Consider a discounted ( [[IMAGE:a489e35127f63335_9_45]] ), episodic task that has two non-terminal states, A and B. The
rewards belongs to { 0,1,2,3 } set. The following are some episodes experienced by an agent
following a fixed policy. The terminal state is not explicitly mentioned for any of the episodes.
[[IMAGE:a489e35127f63335_10_46]]
Based on the above data, answer the given subquestions.
What is the estimate of [[IMAGE:a489e35127f63335_11_51]] returned by batch [[IMAGE:a489e35127f63335_11_52]] ?




A published solution is not available for this question yet.
Question 17 NAT · 1.0 marks
Consider the following Deep Q-Network given below that approximates the Q-values for actions
[[IMAGE:a489e35127f63335_11_53]] and [[IMAGE:a489e35127f63335_11_54]] for states [[IMAGE:a489e35127f63335_11_55]] . The hidden layer nodes (colored green) employ the [[IMAGE:a489e35127f63335_11_56]] activation
function. The weights of the connections are given directly above the edges connecting the nodes.
For instance, 3 is the weight for edge [[IMAGE:a489e35127f63335_12_57]] and -3 is the weight for edge [[IMAGE:a489e35127f63335_12_58]] .
[[IMAGE:a489e35127f63335_12_59]]
We are interested in the Q-values calculated by the network for the following states:
[[IMAGE:a489e35127f63335_12_60]]
Based on the above data, answer the given subquestions.
Recall the linear-approximator view for DQNs ( [[IMAGE:a489e35127f63335_12_61]] ). Evaluate [[IMAGE:a489e35127f63335_12_62]] i.e. the
state-representation for [[IMAGE:a489e35127f63335_12_63]] . Enter [[IMAGE:a489e35127f63335_12_64]] (the L-1 norm of the state-representation) as your
answer.












A published solution is not available for this question yet.
Question 18 MCQ · 1.0 marks
Consider the following Deep Q-Network given below that approximates the Q-values for actions
[[IMAGE:a489e35127f63335_11_53]] and [[IMAGE:a489e35127f63335_11_54]] for states [[IMAGE:a489e35127f63335_11_55]] . The hidden layer nodes (colored green) employ the [[IMAGE:a489e35127f63335_11_56]] activation
function. The weights of the connections are given directly above the edges connecting the nodes.
For instance, 3 is the weight for edge [[IMAGE:a489e35127f63335_12_57]] and -3 is the weight for edge [[IMAGE:a489e35127f63335_12_58]] .
[[IMAGE:a489e35127f63335_12_59]]
We are interested in the Q-values calculated by the network for the following states:
[[IMAGE:a489e35127f63335_12_60]]
Based on the above data, answer the given subquestions.
Obtain the state-representations for [[IMAGE:a489e35127f63335_12_65]] and [[IMAGE:a489e35127f63335_12_66]] as well. Which of the following pairs of states are
the closest (euclidean distance) in their respective representations?










[[IMAGE:a489e35127f63335_13_67]]

[[IMAGE:a489e35127f63335_13_68]]

[[IMAGE:a489e35127f63335_13_69]]

A published solution is not available for this question yet.
Question 19 NAT · 1.0 marks
Consider the following Deep Q-Network given below that approximates the Q-values for actions
[[IMAGE:a489e35127f63335_11_53]] and [[IMAGE:a489e35127f63335_11_54]] for states [[IMAGE:a489e35127f63335_11_55]] . The hidden layer nodes (colored green) employ the [[IMAGE:a489e35127f63335_11_56]] activation
function. The weights of the connections are given directly above the edges connecting the nodes.
For instance, 3 is the weight for edge [[IMAGE:a489e35127f63335_12_57]] and -3 is the weight for edge [[IMAGE:a489e35127f63335_12_58]] .
[[IMAGE:a489e35127f63335_12_59]]
We are interested in the Q-values calculated by the network for the following states:
[[IMAGE:a489e35127f63335_12_60]]
Based on the above data, answer the given subquestions.
What is [[IMAGE:a489e35127f63335_13_70]] ?









A published solution is not available for this question yet.
Question 20 NAT · 1.0 marks
Consider the following Deep Q-Network given below that approximates the Q-values for actions
[[IMAGE:a489e35127f63335_11_53]] and [[IMAGE:a489e35127f63335_11_54]] for states [[IMAGE:a489e35127f63335_11_55]] . The hidden layer nodes (colored green) employ the [[IMAGE:a489e35127f63335_11_56]] activation
function. The weights of the connections are given directly above the edges connecting the nodes.
For instance, 3 is the weight for edge [[IMAGE:a489e35127f63335_12_57]] and -3 is the weight for edge [[IMAGE:a489e35127f63335_12_58]] .
[[IMAGE:a489e35127f63335_12_59]]
We are interested in the Q-values calculated by the network for the following states:
[[IMAGE:a489e35127f63335_12_60]]
Based on the above data, answer the given subquestions.
What is [[IMAGE:a489e35127f63335_13_71]] ?









A published solution is not available for this question yet.
Question 21 NAT · 1.0 marks
Consider the following Deep Q-Network given below that approximates the Q-values for actions
[[IMAGE:a489e35127f63335_11_53]] and [[IMAGE:a489e35127f63335_11_54]] for states [[IMAGE:a489e35127f63335_11_55]] . The hidden layer nodes (colored green) employ the [[IMAGE:a489e35127f63335_11_56]] activation
function. The weights of the connections are given directly above the edges connecting the nodes.
For instance, 3 is the weight for edge [[IMAGE:a489e35127f63335_12_57]] and -3 is the weight for edge [[IMAGE:a489e35127f63335_12_58]] .
[[IMAGE:a489e35127f63335_12_59]]
We are interested in the Q-values calculated by the network for the following states:
[[IMAGE:a489e35127f63335_12_60]]
Based on the above data, answer the given subquestions.
Suppose you use [[IMAGE:a489e35127f63335_13_72]] -greedy with [[IMAGE:a489e35127f63335_13_73]] selection for balancing exploration with exploitation
during training. What is the probability that you'll pick action O2 given that the observed state is
[[IMAGE:a489e35127f63335_14_74]] ?











A published solution is not available for this question yet.