da5007_2026T1_ET_FN.pdf
Reinforcement Learning · End Term · Jan 2026 FN
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 MCQ · 0.0 marks
**INSTRUCTIONS:**
This paper contains bonus questions marked with **[Bonus]** at the beginning of the question. These
questions carry 0 marks toward the regular total of 30 and do not contribute to the End Term
score. However, correct answers will earn additional bonus marks, which will be added to your
final score in accordance with the grading policy. Attempting these questions is optional, and
there is no penalty for incorrect or unattempted bonus questions.
Instructions has been mentioned above.
This Instructions is just for a reference & not for an evaluation.
A published solution is not available for this question yet.
Question 3 MCQ · 1.0 marks
An option [[IMAGE:0d29d23c93e001a7_2_2]] is defined as a triple. Which of the following correctly lists **all three** components of an
option?

State space, action space, reward function
Initiation set, option policy, termination function
Value function, behaviour policy, discount factor
Transition model, reward model, termination condition
A published solution is not available for this question yet.
Question 4 MCQ · 1.0 marks
In DDPG, target networks are updated via soft updates [[IMAGE:0d29d23c93e001a7_3_3]] with [[IMAGE:0d29d23c93e001a7_3_4]] . The
primary reason for this slow update is:


To prevent the actor from updating faster than the critic
To stabilise training by preventing the TD target from changing too rapidly
To ensure the replay buffer contains on-policy data
To match the learning rate of the actor and critic networks
A published solution is not available for this question yet.
Question 5 MCQ · 1.0 marks
DDPG uses an experience replay buffer. What is the direct consequence of using an experience
replay buffer
The policy gradient estimate becomes unbiased
The algorithm becomes on-policy
Temporal correlations between consecutive samples are broken
The critic no longer requires a target network
A published solution is not available for this question yet.
Question 6 MCQ · 1.0 marks
Compared to the policy gradient, the DPG theorem requires integration over:
Both state and action spaces
State space only
Action space only
Neither — it requires only a single action sample per state
A published solution is not available for this question yet.
Question 7 MCQ · 0.0 marks
**[Bonus]** The TRPO surrogate objective is defined as
[[IMAGE:0d29d23c93e001a7_4_5]]
where
[[IMAGE:0d29d23c93e001a7_4_6]]
is an importance-sampling ratio.
Why is it justified to optimise [[IMAGE:0d29d23c93e001a7_4_7]] instead of the true objective [[IMAGE:0d29d23c93e001a7_4_8]]




[[IMAGE:0d29d23c93e001a7_4_9]] is exactly equal to [[IMAGE:0d29d23c93e001a7_4_10]] for all [[IMAGE:0d29d23c93e001a7_4_11]]



[[IMAGE:0d29d23c93e001a7_4_12]] is a first-order approximation of [[IMAGE:0d29d23c93e001a7_4_13]] , accurate when [[IMAGE:0d29d23c93e001a7_4_14]] is
close to [[IMAGE:0d29d23c93e001a7_4_15]] --- precisely the regime enforced by the KL constraint




[[IMAGE:0d29d23c93e001a7_4_16]] is an unbiased estimator of [[IMAGE:0d29d23c93e001a7_4_17]] under any policy


The surrogate eliminates the need for a value function baseline
A published solution is not available for this question yet.
Question 8 MSQ · 1.0 marks
Which are known limitations or failure modes of DDPG?
High sensitivity to hyperparameters such as learning rate and architecture
Inability to handle continuous action spaces
Brittle training that may diverge without careful tuning
Deterministic policy requires explicit exploration noise during training
A published solution is not available for this question yet.
Question 9 MSQ · 2.0 marks
Which are not valid estimators of the advantage function [[IMAGE:0d29d23c93e001a7_5_18]] ?

[[IMAGE:0d29d23c93e001a7_5_19]]

[[IMAGE:0d29d23c93e001a7_5_20]]

[[IMAGE:0d29d23c93e001a7_5_21]]

[[IMAGE:0d29d23c93e001a7_5_22]]

A published solution is not available for this question yet.
Question 10 MSQ · 3.0 marks
Recall standard control algorithms in RL, like Q-Learning, SARSA, Expected SARSA, etc., involve a
maximisation step in constructing their target policies.
For example:
In Q-Learning, the target is:
[[IMAGE:0d29d23c93e001a7_6_23]]
Which of the following statements are correct? **Select all that apply.**

[[IMAGE:0d29d23c93e001a7_6_24]]

[[IMAGE:0d29d23c93e001a7_6_25]]

[[IMAGE:0d29d23c93e001a7_7_26]]

[[IMAGE:0d29d23c93e001a7_7_27]]

[[IMAGE:0d29d23c93e001a7_7_28]]

[[IMAGE:0d29d23c93e001a7_7_29]]

A published solution is not available for this question yet.
Question 11 NAT · 1.0 marks
An option [[IMAGE:0d29d23c93e001a7_7_30]] is initiated in state [[IMAGE:0d29d23c93e001a7_7_31]] and executes for [[IMAGE:0d29d23c93e001a7_7_32]] steps with rewards
[[IMAGE:0d29d23c93e001a7_7_33]] and [[IMAGE:0d29d23c93e001a7_7_34]] . What is the cumulative discounted return of this option?





A published solution is not available for this question yet.
Question 12 NAT · 0.0 marks
**[Bonus]** Consider an agent in state [[IMAGE:0d29d23c93e001a7_7_35]] that executes action [[IMAGE:0d29d23c93e001a7_7_36]] . The transition is stochastic: the
environment transitions to one of two possible next states according to the following distribution:
[[IMAGE:0d29d23c93e001a7_7_37]]
with associated immediate rewards:
[[IMAGE:0d29d23c93e001a7_8_38]]
The Q-value estimates of two independent estimators [[IMAGE:0d29d23c93e001a7_8_39]] and [[IMAGE:0d29d23c93e001a7_8_40]] at each possible next state are
given in the tables below.
[[IMAGE:0d29d23c93e001a7_8_41]]
**Additional parameters:**
[[IMAGE:0d29d23c93e001a7_8_42]]
Using Double Q-learning with [[IMAGE:0d29d23c93e001a7_8_43]] for action selection and [[IMAGE:0d29d23c93e001a7_8_44]] for action evaluation, compute the
expected Double Q-learning target [[IMAGE:0d29d23c93e001a7_8_45]] for updating [[IMAGE:0d29d23c93e001a7_8_46]] :
[[IMAGE:0d29d23c93e001a7_8_47]]
Round your final answer to **2 decimal places**.













A published solution is not available for this question yet.
Question 13 MCQ · 2.0 marks
Ganesh is designing a robot arm controller with a 6-dimensional continuous action space. He
initially uses a stochastic policy [[IMAGE:0d29d23c93e001a7_9_48]] with the standard policy gradient(SPG):
[[IMAGE:0d29d23c93e001a7_9_49]]
Karthik suggests switching to the Deterministic Policy Gradient (DPG) with policy [[IMAGE:0d29d23c93e001a7_9_50]] :
[[IMAGE:0d29d23c93e001a7_9_51]]
Karthik argues that DPG is strictly more sample-efficient here. However, Ganesh notices that the
current DPG implementation uses only on-policy data sampled under [[IMAGE:0d29d23c93e001a7_9_52]] , and that the robot
operates in a noisy, stochastic environment where some states are rarely visited under the greedy
policy.
Based on the above data, answer the given subquestions.
Why is DPG generally more sample-efficient than SPG in high-dimensional continuous action
spaces?





DPG avoids computing [[IMAGE:0d29d23c93e001a7_9_53]] , which is numerically unstable in high
dimensions

DPG’s gradient integrates over the state space only, whereas SPG integrates
over both state and action spaces - requiring far more samples to estimate accurately when [[IMAGE:0d29d23c93e001a7_9_54]]
is large

DPG uses a target network which reduces gradient variance automatically
DPG enforces determinism, which eliminates exploration noise and thus
requires fewer environment interactions
A published solution is not available for this question yet.
Question 14 MSQ · 2.0 marks
Ganesh is designing a robot arm controller with a 6-dimensional continuous action space. He
initially uses a stochastic policy [[IMAGE:0d29d23c93e001a7_9_48]] with the standard policy gradient(SPG):
[[IMAGE:0d29d23c93e001a7_9_49]]
Karthik suggests switching to the Deterministic Policy Gradient (DPG) with policy [[IMAGE:0d29d23c93e001a7_9_50]] :
[[IMAGE:0d29d23c93e001a7_9_51]]
Karthik argues that DPG is strictly more sample-efficient here. However, Ganesh notices that the
current DPG implementation uses only on-policy data sampled under [[IMAGE:0d29d23c93e001a7_9_52]] , and that the robot
operates in a noisy, stochastic environment where some states are rarely visited under the greedy
policy.
Based on the above data, answer the given subquestions.
Ganesh’s on-policy DPG implementation is likely to suffer from which of the following problems?
Select all that apply.





Poor coverage of the state space — the greedy policy may never visit states
needed for accurate [[IMAGE:0d29d23c93e001a7_10_55]] estimation

Maximization bias, because the greedy policy always selects the action with
highest estimated [[IMAGE:0d29d23c93e001a7_10_56]]

The DPG theorem technically requires an off-policy behaviour policy to ensure
sufficient state coverage; on-policy data violates this
Insufficient exploration, since [[IMAGE:0d29d23c93e001a7_10_57]] is deterministic and produces no stochastic
variation in actions

The action-value gradient [[IMAGE:0d29d23c93e001a7_10_58]] cannot be computed on-policy

A published solution is not available for this question yet.
Question 15 NAT · 0.0 marks
Ganesh is designing a robot arm controller with a 6-dimensional continuous action space. He
initially uses a stochastic policy [[IMAGE:0d29d23c93e001a7_9_48]] with the standard policy gradient(SPG):
[[IMAGE:0d29d23c93e001a7_9_49]]
Karthik suggests switching to the Deterministic Policy Gradient (DPG) with policy [[IMAGE:0d29d23c93e001a7_9_50]] :
[[IMAGE:0d29d23c93e001a7_9_51]]
Karthik argues that DPG is strictly more sample-efficient here. However, Ganesh notices that the
current DPG implementation uses only on-policy data sampled under [[IMAGE:0d29d23c93e001a7_9_52]] , and that the robot
operates in a noisy, stochastic environment where some states are rarely visited under the greedy
policy.
Based on the above data, answer the given subquestions.
**[Bonus]** Ganesh adds Gaussian exploration noise [[IMAGE:0d29d23c93e001a7_10_59]] , [[IMAGE:0d29d23c93e001a7_10_60]] , with [[IMAGE:0d29d23c93e001a7_10_61]] .
The DPG actor update for a single sample is:
[[IMAGE:0d29d23c93e001a7_10_62]]
Given [[IMAGE:0d29d23c93e001a7_10_63]] , and for a particular state [[IMAGE:0d29d23c93e001a7_10_64]] :
[[IMAGE:0d29d23c93e001a7_10_65]]
Compute the [[IMAGE:0d29d23c93e001a7_10_66]] -norm of the parameter update vector [[IMAGE:0d29d23c93e001a7_10_67]] . Round to 4 decimal places.














A published solution is not available for this question yet.
Question 16 MSQ · 3.0 marks
A Bernoulli-logistic unit is a stochastic neuron whose output [[IMAGE:0d29d23c93e001a7_11_68]] follows a Bernoulli
distribution:
[[IMAGE:0d29d23c93e001a7_11_69]]
The unit receives a feature vector [[IMAGE:0d29d23c93e001a7_11_70]] as input. Let [[IMAGE:0d29d23c93e001a7_11_71]] denote the action preference
for action [[IMAGE:0d29d23c93e001a7_11_72]] in state [[IMAGE:0d29d23c93e001a7_11_73]] under policy parameter [[IMAGE:0d29d23c93e001a7_11_74]] .
Assume that difference in action preferences is a linear function of the input:
[[IMAGE:0d29d23c93e001a7_11_75]]
The softmax (exponential soft-max) policy converts preferences to probabilities:
[[IMAGE:0d29d23c93e001a7_11_76]]
The REINFORCE (Monte Carlo policy gradient) update rule is:
[[IMAGE:0d29d23c93e001a7_11_77]]
where [[IMAGE:0d29d23c93e001a7_11_78]] is the return from time [[IMAGE:0d29d23c93e001a7_11_79]] , and [[IMAGE:0d29d23c93e001a7_11_80]] is called the eligibility
vector.
Based on the above data, answer the given subquestions.
Applying the softmax distribution to the Bernoulli-logistic unit, which of the following correctly
expresses
[[IMAGE:0d29d23c93e001a7_12_81]] ?














[[IMAGE:0d29d23c93e001a7_12_82]]

[[IMAGE:0d29d23c93e001a7_12_83]]

[[IMAGE:0d29d23c93e001a7_12_84]]

[[IMAGE:0d29d23c93e001a7_12_85]]

A published solution is not available for this question yet.
Question 17 MSQ · 0.0 marks
A Bernoulli-logistic unit is a stochastic neuron whose output [[IMAGE:0d29d23c93e001a7_11_68]] follows a Bernoulli
distribution:
[[IMAGE:0d29d23c93e001a7_11_69]]
The unit receives a feature vector [[IMAGE:0d29d23c93e001a7_11_70]] as input. Let [[IMAGE:0d29d23c93e001a7_11_71]] denote the action preference
for action [[IMAGE:0d29d23c93e001a7_11_72]] in state [[IMAGE:0d29d23c93e001a7_11_73]] under policy parameter [[IMAGE:0d29d23c93e001a7_11_74]] .
Assume that difference in action preferences is a linear function of the input:
[[IMAGE:0d29d23c93e001a7_11_75]]
The softmax (exponential soft-max) policy converts preferences to probabilities:
[[IMAGE:0d29d23c93e001a7_11_76]]
The REINFORCE (Monte Carlo policy gradient) update rule is:
[[IMAGE:0d29d23c93e001a7_11_77]]
where [[IMAGE:0d29d23c93e001a7_11_78]] is the return from time [[IMAGE:0d29d23c93e001a7_11_79]] , and [[IMAGE:0d29d23c93e001a7_11_80]] is called the eligibility
vector.
Based on the above data, answer the given subquestions.
**[Bonus]** The Monte Carlo REINFORCE update for the Bernoulli-logistic unit upon receipt of return
[[IMAGE:0d29d23c93e001a7_12_86]] is:
[[IMAGE:0d29d23c93e001a7_12_87]]
Which of the following statements about this update are correct? Select all that apply.















When [[IMAGE:0d29d23c93e001a7_12_88]] , the eligibility vector is [[IMAGE:0d29d23c93e001a7_12_89]] , so the update reinforces
[[IMAGE:0d29d23c93e001a7_12_90]] in the direction of [[IMAGE:0d29d23c93e001a7_12_91]] when [[IMAGE:0d29d23c93e001a7_12_92]]





When [[IMAGE:0d29d23c93e001a7_12_93]] , the eligibility vector is [[IMAGE:0d29d23c93e001a7_12_94]] , meaning the update
decreases [[IMAGE:0d29d23c93e001a7_12_95]] (i.e. makes action 1 less likely) when [[IMAGE:0d29d23c93e001a7_12_96]]




The magnitude of the update is larger when the selected action is surprising
(low probability), since [[IMAGE:0d29d23c93e001a7_12_97]] is large when [[IMAGE:0d29d23c93e001a7_12_98]] and [[IMAGE:0d29d23c93e001a7_12_99]] is large when [[IMAGE:0d29d23c93e001a7_12_100]]




The update is an unbiased estimate of [[IMAGE:0d29d23c93e001a7_12_101]] only if the episode is generated
on-policy under [[IMAGE:0d29d23c93e001a7_12_102]]


Subtracting a baseline [[IMAGE:0d29d23c93e001a7_13_103]] from [[IMAGE:0d29d23c93e001a7_13_104]] changes the direction of the gradient
update, introducing bias to improve convergence speed


A published solution is not available for this question yet.
Question 18 MSQ · 2.0 marks
Vivek implements the REINFORCE (Monte Carlo policy gradient) algorithm for a text-generation
task modelled as an episodic MDP. The policy is a neural network [[IMAGE:0d29d23c93e001a7_13_105]] parameterised by [[IMAGE:0d29d23c93e001a7_13_106]] .
At each episode's end, he collects the full trajectory [[IMAGE:0d29d23c93e001a7_13_107]] and
computes the return:
[[IMAGE:0d29d23c93e001a7_13_108]]
The REINFORCE gradient estimator is:
[[IMAGE:0d29d23c93e001a7_13_109]]
After 1000 episodes, he observes that the gradient variance is extremely high and training is
unstable. Neel suggests adding a baseline [[IMAGE:0d29d23c93e001a7_13_110]] :
[[IMAGE:0d29d23c93e001a7_13_111]]
He claims: "Subtracting [[IMAGE:0d29d23c93e001a7_13_112]] reduces variance without biasing the gradient, as long as [[IMAGE:0d29d23c93e001a7_13_113]] does not
depend on [[IMAGE:0d29d23c93e001a7_13_114]] .''
Based on the above data, answer the given subquestions.
Why does REINFORCE exhibit high gradient variance, particularly in long-horizon tasks like text
generation?










The log-probability [[IMAGE:0d29d23c93e001a7_14_115]] is numerically unstable for large neural
networks

[[IMAGE:0d29d23c93e001a7_14_116]] accumulates rewards over the entire future horizon; in long episodes with
noisy rewards, [[IMAGE:0d29d23c93e001a7_14_117]] has large variance, which directly scales the gradient estimate


Monte Carlo rollouts are biased estimators of the true return
The policy gradient theorem does not hold for episodic tasks with discount
factor [[IMAGE:0d29d23c93e001a7_14_118]]

A published solution is not available for this question yet.
Question 19 NAT · 2.0 marks
Vivek implements the REINFORCE (Monte Carlo policy gradient) algorithm for a text-generation
task modelled as an episodic MDP. The policy is a neural network [[IMAGE:0d29d23c93e001a7_13_105]] parameterised by [[IMAGE:0d29d23c93e001a7_13_106]] .
At each episode's end, he collects the full trajectory [[IMAGE:0d29d23c93e001a7_13_107]] and
computes the return:
[[IMAGE:0d29d23c93e001a7_13_108]]
The REINFORCE gradient estimator is:
[[IMAGE:0d29d23c93e001a7_13_109]]
After 1000 episodes, he observes that the gradient variance is extremely high and training is
unstable. Neel suggests adding a baseline [[IMAGE:0d29d23c93e001a7_13_110]] :
[[IMAGE:0d29d23c93e001a7_13_111]]
He claims: "Subtracting [[IMAGE:0d29d23c93e001a7_13_112]] reduces variance without biasing the gradient, as long as [[IMAGE:0d29d23c93e001a7_13_113]] does not
depend on [[IMAGE:0d29d23c93e001a7_13_114]] .''
Based on the above data, answer the given subquestions.
For a single timestep [[IMAGE:0d29d23c93e001a7_14_119]] in an episode, the following values are observed:
[[IMAGE:0d29d23c93e001a7_14_120]]
Compute the parameter update [[IMAGE:0d29d23c93e001a7_14_121]] and report [[IMAGE:0d29d23c93e001a7_14_122]] .
Round to 4 decimal places.














A published solution is not available for this question yet.
Question 20 MSQ · 2.0 marks
Vivek implements the REINFORCE (Monte Carlo policy gradient) algorithm for a text-generation
task modelled as an episodic MDP. The policy is a neural network [[IMAGE:0d29d23c93e001a7_13_105]] parameterised by [[IMAGE:0d29d23c93e001a7_13_106]] .
At each episode's end, he collects the full trajectory [[IMAGE:0d29d23c93e001a7_13_107]] and
computes the return:
[[IMAGE:0d29d23c93e001a7_13_108]]
The REINFORCE gradient estimator is:
[[IMAGE:0d29d23c93e001a7_13_109]]
After 1000 episodes, he observes that the gradient variance is extremely high and training is
unstable. Neel suggests adding a baseline [[IMAGE:0d29d23c93e001a7_13_110]] :
[[IMAGE:0d29d23c93e001a7_13_111]]
He claims: "Subtracting [[IMAGE:0d29d23c93e001a7_13_112]] reduces variance without biasing the gradient, as long as [[IMAGE:0d29d23c93e001a7_13_113]] does not
depend on [[IMAGE:0d29d23c93e001a7_13_114]] .''
Based on the above data, answer the given subquestions.
Neel claims that any function [[IMAGE:0d29d23c93e001a7_15_123]] not depending on [[IMAGE:0d29d23c93e001a7_15_124]] is a valid zero-bias baseline. Which of the
following statements are correct justifications or counter-examples to this claim?












[[IMAGE:0d29d23c93e001a7_15_125]] for any [[IMAGE:0d29d23c93e001a7_15_126]] independent of [[IMAGE:0d29d23c93e001a7_15_127]] , since
[[IMAGE:0d29d23c93e001a7_15_128]]




If [[IMAGE:0d29d23c93e001a7_15_129]] , the baseline perfectly cancels the return and the gradient
becomes zero — a degenerate case

A baseline that depends on both [[IMAGE:0d29d23c93e001a7_15_130]] and [[IMAGE:0d29d23c93e001a7_15_131]] (e.g. [[IMAGE:0d29d23c93e001a7_15_132]] ) still leaves the
gradient unbiased



Using [[IMAGE:0d29d23c93e001a7_15_133]] introduces bias because [[IMAGE:0d29d23c93e001a7_15_134]] is estimated, not known
exactly


A published solution is not available for this question yet.
Question 21 MCQ · 2.0 marks
Lalit uses two algorithms to learn option values in a given programming assignment:
Algorithm A — SMDP Q-Learning: Updates [[IMAGE:0d29d23c93e001a7_15_135]] only at option termination using the full option
return:
[[IMAGE:0d29d23c93e001a7_15_136]]
Algorithm B — Intra-Option Q-Learning: Updates [[IMAGE:0d29d23c93e001a7_15_137]] at every primitive step using the one-step
consistency condition:
[[IMAGE:0d29d23c93e001a7_15_138]]
[[IMAGE:0d29d23c93e001a7_15_139]]
where
[[IMAGE:0d29d23c93e001a7_16_140]] is the termination probability at the next state.
The agent runs in a corridor with 5 states [[IMAGE:0d29d23c93e001a7_16_141]] . Option [[IMAGE:0d29d23c93e001a7_16_142]] has intra-option policy: move
right always. [[IMAGE:0d29d23c93e001a7_16_143]] if [[IMAGE:0d29d23c93e001a7_16_144]] , else [[IMAGE:0d29d23c93e001a7_16_145]] . Currently: [[IMAGE:0d29d23c93e001a7_16_146]] , [[IMAGE:0d29d23c93e001a7_16_147]] ,
[[IMAGE:0d29d23c93e001a7_16_148]] , [[IMAGE:0d29d23c93e001a7_16_149]] , [[IMAGE:0d29d23c93e001a7_16_150]] , [[IMAGE:0d29d23c93e001a7_16_151]] , agent moves from [[IMAGE:0d29d23c93e001a7_16_152]] to [[IMAGE:0d29d23c93e001a7_16_153]] .
Based on the above data, answer the given subquestions.
What is the key advantage of Intra-Option Q-Learning (Algorithm B) over SMDP Q-Learning
(Algorithm A)?



















It requires no knowledge of the termination function [[IMAGE:0d29d23c93e001a7_16_154]]

It updates [[IMAGE:0d29d23c93e001a7_16_155]] at every time step rather than only at termination, making
more efficient use of experience — especially for long-duration options

It converges to a globally optimal policy, whereas SMDP Q-learning only
converges locally
It eliminates the need for a high-level policy by learning all options
simultaneously
A published solution is not available for this question yet.
Question 22 NAT · 2.0 marks
Lalit uses two algorithms to learn option values in a given programming assignment:
Algorithm A — SMDP Q-Learning: Updates [[IMAGE:0d29d23c93e001a7_15_135]] only at option termination using the full option
return:
[[IMAGE:0d29d23c93e001a7_15_136]]
Algorithm B — Intra-Option Q-Learning: Updates [[IMAGE:0d29d23c93e001a7_15_137]] at every primitive step using the one-step
consistency condition:
[[IMAGE:0d29d23c93e001a7_15_138]]
[[IMAGE:0d29d23c93e001a7_15_139]]
where
[[IMAGE:0d29d23c93e001a7_16_140]] is the termination probability at the next state.
The agent runs in a corridor with 5 states [[IMAGE:0d29d23c93e001a7_16_141]] . Option [[IMAGE:0d29d23c93e001a7_16_142]] has intra-option policy: move
right always. [[IMAGE:0d29d23c93e001a7_16_143]] if [[IMAGE:0d29d23c93e001a7_16_144]] , else [[IMAGE:0d29d23c93e001a7_16_145]] . Currently: [[IMAGE:0d29d23c93e001a7_16_146]] , [[IMAGE:0d29d23c93e001a7_16_147]] ,
[[IMAGE:0d29d23c93e001a7_16_148]] , [[IMAGE:0d29d23c93e001a7_16_149]] , [[IMAGE:0d29d23c93e001a7_16_150]] , [[IMAGE:0d29d23c93e001a7_16_151]] , agent moves from [[IMAGE:0d29d23c93e001a7_16_152]] to [[IMAGE:0d29d23c93e001a7_16_153]] .
Based on the above data, answer the given subquestions.
Using Algorithm B (Intra-Option Q-Learning) and the values from the comprehension section,
compute the TD error [[IMAGE:0d29d23c93e001a7_16_156]] for the transition [[IMAGE:0d29d23c93e001a7_16_157]] . Round your answer to 4 decimal
places.
Note: [[IMAGE:0d29d23c93e001a7_16_158]] (the option does not terminate at [[IMAGE:0d29d23c93e001a7_16_159]] ).
[[IMAGE:0d29d23c93e001a7_16_160]]
























A published solution is not available for this question yet.
Question 23 NAT · 2.0 marks
Lalit uses two algorithms to learn option values in a given programming assignment:
Algorithm A — SMDP Q-Learning: Updates [[IMAGE:0d29d23c93e001a7_15_135]] only at option termination using the full option
return:
[[IMAGE:0d29d23c93e001a7_15_136]]
Algorithm B — Intra-Option Q-Learning: Updates [[IMAGE:0d29d23c93e001a7_15_137]] at every primitive step using the one-step
consistency condition:
[[IMAGE:0d29d23c93e001a7_15_138]]
[[IMAGE:0d29d23c93e001a7_15_139]]
where
[[IMAGE:0d29d23c93e001a7_16_140]] is the termination probability at the next state.
The agent runs in a corridor with 5 states [[IMAGE:0d29d23c93e001a7_16_141]] . Option [[IMAGE:0d29d23c93e001a7_16_142]] has intra-option policy: move
right always. [[IMAGE:0d29d23c93e001a7_16_143]] if [[IMAGE:0d29d23c93e001a7_16_144]] , else [[IMAGE:0d29d23c93e001a7_16_145]] . Currently: [[IMAGE:0d29d23c93e001a7_16_146]] , [[IMAGE:0d29d23c93e001a7_16_147]] ,
[[IMAGE:0d29d23c93e001a7_16_148]] , [[IMAGE:0d29d23c93e001a7_16_149]] , [[IMAGE:0d29d23c93e001a7_16_150]] , [[IMAGE:0d29d23c93e001a7_16_151]] , agent moves from [[IMAGE:0d29d23c93e001a7_16_152]] to [[IMAGE:0d29d23c93e001a7_16_153]] .
Based on the above data, answer the given subquestions.
Using the value of [[IMAGE:0d29d23c93e001a7_17_161]] computed in previous sub-question, calculate the updated value of [[IMAGE:0d29d23c93e001a7_17_162]] .
Round your answer to 4 decimal places.
[[IMAGE:0d29d23c93e001a7_17_163]]






















A published solution is not available for this question yet.