MauryaHub PYQ Practice

da5007_2026T1_ET_FN.pdf

Reinforcement Learning · End Term · Jan 2026 FN

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 2 MCQ · 0.0 marks

**INSTRUCTIONS:** This paper contains bonus questions marked with **[Bonus]** at the beginning of the question. These questions carry 0 marks toward the regular total of 30 and do not contribute to the End Term score. However, correct answers will earn additional bonus marks, which will be added to your final score in accordance with the grading policy. Attempting these questions is optional, and there is no penalty for incorrect or unattempted bonus questions.
  1. Instructions has been mentioned above.
  2. This Instructions is just for a reference & not for an evaluation.

A published solution is not available for this question yet.

Question 3 MCQ · 1.0 marks

An option [[IMAGE:0d29d23c93e001a7_2_2]] is defined as a triple. Which of the following correctly lists **all three** components of an option?
Source diagram or notation
  1. State space, action space, reward function
  2. Initiation set, option policy, termination function
  3. Value function, behaviour policy, discount factor
  4. Transition model, reward model, termination condition

A published solution is not available for this question yet.

Question 4 MCQ · 1.0 marks

In DDPG, target networks are updated via soft updates [[IMAGE:0d29d23c93e001a7_3_3]] with [[IMAGE:0d29d23c93e001a7_3_4]] . The primary reason for this slow update is:
Source diagram or notationSource diagram or notation
  1. To prevent the actor from updating faster than the critic
  2. To stabilise training by preventing the TD target from changing too rapidly
  3. To ensure the replay buffer contains on-policy data
  4. To match the learning rate of the actor and critic networks

A published solution is not available for this question yet.

Question 5 MCQ · 1.0 marks

DDPG uses an experience replay buffer. What is the direct consequence of using an experience replay buffer
  1. The policy gradient estimate becomes unbiased
  2. The algorithm becomes on-policy
  3. Temporal correlations between consecutive samples are broken
  4. The critic no longer requires a target network

A published solution is not available for this question yet.

Question 6 MCQ · 1.0 marks

Compared to the policy gradient, the DPG theorem requires integration over:
  1. Both state and action spaces
  2. State space only
  3. Action space only
  4. Neither — it requires only a single action sample per state

A published solution is not available for this question yet.

Question 7 MCQ · 0.0 marks

**[Bonus]** The TRPO surrogate objective is defined as [[IMAGE:0d29d23c93e001a7_4_5]] where [[IMAGE:0d29d23c93e001a7_4_6]] is an importance-sampling ratio. Why is it justified to optimise [[IMAGE:0d29d23c93e001a7_4_7]] instead of the true objective [[IMAGE:0d29d23c93e001a7_4_8]]
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
  1. [[IMAGE:0d29d23c93e001a7_4_9]] is exactly equal to [[IMAGE:0d29d23c93e001a7_4_10]] for all [[IMAGE:0d29d23c93e001a7_4_11]]
    Source diagram or notationSource diagram or notationSource diagram or notation
  2. [[IMAGE:0d29d23c93e001a7_4_12]] is a first-order approximation of [[IMAGE:0d29d23c93e001a7_4_13]] , accurate when [[IMAGE:0d29d23c93e001a7_4_14]] is close to [[IMAGE:0d29d23c93e001a7_4_15]] --- precisely the regime enforced by the KL constraint
    Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
  3. [[IMAGE:0d29d23c93e001a7_4_16]] is an unbiased estimator of [[IMAGE:0d29d23c93e001a7_4_17]] under any policy
    Source diagram or notationSource diagram or notation
  4. The surrogate eliminates the need for a value function baseline

A published solution is not available for this question yet.

Question 8 MSQ · 1.0 marks

Which are known limitations or failure modes of DDPG?
  1. High sensitivity to hyperparameters such as learning rate and architecture
  2. Inability to handle continuous action spaces
  3. Brittle training that may diverge without careful tuning
  4. Deterministic policy requires explicit exploration noise during training

A published solution is not available for this question yet.

Question 9 MSQ · 2.0 marks

Which are not valid estimators of the advantage function [[IMAGE:0d29d23c93e001a7_5_18]] ?
Source diagram or notation
  1. [[IMAGE:0d29d23c93e001a7_5_19]]
    Source diagram or notation
  2. [[IMAGE:0d29d23c93e001a7_5_20]]
    Source diagram or notation
  3. [[IMAGE:0d29d23c93e001a7_5_21]]
    Source diagram or notation
  4. [[IMAGE:0d29d23c93e001a7_5_22]]
    Source diagram or notation

A published solution is not available for this question yet.

Question 10 MSQ · 3.0 marks

Recall standard control algorithms in RL, like Q-Learning, SARSA, Expected SARSA, etc., involve a maximisation step in constructing their target policies. For example: In Q-Learning, the target is: [[IMAGE:0d29d23c93e001a7_6_23]] Which of the following statements are correct? **Select all that apply.**
Source diagram or notation
  1. [[IMAGE:0d29d23c93e001a7_6_24]]
    Source diagram or notation
  2. [[IMAGE:0d29d23c93e001a7_6_25]]
    Source diagram or notation
  3. [[IMAGE:0d29d23c93e001a7_7_26]]
    Source diagram or notation
  4. [[IMAGE:0d29d23c93e001a7_7_27]]
    Source diagram or notation
  5. [[IMAGE:0d29d23c93e001a7_7_28]]
    Source diagram or notation
  6. [[IMAGE:0d29d23c93e001a7_7_29]]
    Source diagram or notation

A published solution is not available for this question yet.

Question 11 NAT · 1.0 marks

An option [[IMAGE:0d29d23c93e001a7_7_30]] is initiated in state [[IMAGE:0d29d23c93e001a7_7_31]] and executes for [[IMAGE:0d29d23c93e001a7_7_32]] steps with rewards [[IMAGE:0d29d23c93e001a7_7_33]] and [[IMAGE:0d29d23c93e001a7_7_34]] . What is the cumulative discounted return of this option?
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

    A published solution is not available for this question yet.

    Question 12 NAT · 0.0 marks

    **[Bonus]** Consider an agent in state [[IMAGE:0d29d23c93e001a7_7_35]] that executes action [[IMAGE:0d29d23c93e001a7_7_36]] . The transition is stochastic: the environment transitions to one of two possible next states according to the following distribution: [[IMAGE:0d29d23c93e001a7_7_37]] with associated immediate rewards: [[IMAGE:0d29d23c93e001a7_8_38]] The Q-value estimates of two independent estimators [[IMAGE:0d29d23c93e001a7_8_39]] and [[IMAGE:0d29d23c93e001a7_8_40]] at each possible next state are given in the tables below. [[IMAGE:0d29d23c93e001a7_8_41]] **Additional parameters:** [[IMAGE:0d29d23c93e001a7_8_42]] Using Double Q-learning with [[IMAGE:0d29d23c93e001a7_8_43]] for action selection and [[IMAGE:0d29d23c93e001a7_8_44]] for action evaluation, compute the expected Double Q-learning target [[IMAGE:0d29d23c93e001a7_8_45]] for updating [[IMAGE:0d29d23c93e001a7_8_46]] : [[IMAGE:0d29d23c93e001a7_8_47]] Round your final answer to **2 decimal places**.
    Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

      A published solution is not available for this question yet.

      Question 13 MCQ · 2.0 marks

      Ganesh is designing a robot arm controller with a 6-dimensional continuous action space. He initially uses a stochastic policy [[IMAGE:0d29d23c93e001a7_9_48]] with the standard policy gradient(SPG): [[IMAGE:0d29d23c93e001a7_9_49]] Karthik suggests switching to the Deterministic Policy Gradient (DPG) with policy [[IMAGE:0d29d23c93e001a7_9_50]] : [[IMAGE:0d29d23c93e001a7_9_51]] Karthik argues that DPG is strictly more sample-efficient here. However, Ganesh notices that the current DPG implementation uses only on-policy data sampled under [[IMAGE:0d29d23c93e001a7_9_52]] , and that the robot operates in a noisy, stochastic environment where some states are rarely visited under the greedy policy. Based on the above data, answer the given subquestions.
      Why is DPG generally more sample-efficient than SPG in high-dimensional continuous action spaces?
      Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
      1. DPG avoids computing [[IMAGE:0d29d23c93e001a7_9_53]] , which is numerically unstable in high dimensions
        Source diagram or notation
      2. DPG’s gradient integrates over the state space only, whereas SPG integrates over both state and action spaces - requiring far more samples to estimate accurately when [[IMAGE:0d29d23c93e001a7_9_54]] is large
        Source diagram or notation
      3. DPG uses a target network which reduces gradient variance automatically
      4. DPG enforces determinism, which eliminates exploration noise and thus requires fewer environment interactions

      A published solution is not available for this question yet.

      Question 14 MSQ · 2.0 marks

      Ganesh is designing a robot arm controller with a 6-dimensional continuous action space. He initially uses a stochastic policy [[IMAGE:0d29d23c93e001a7_9_48]] with the standard policy gradient(SPG): [[IMAGE:0d29d23c93e001a7_9_49]] Karthik suggests switching to the Deterministic Policy Gradient (DPG) with policy [[IMAGE:0d29d23c93e001a7_9_50]] : [[IMAGE:0d29d23c93e001a7_9_51]] Karthik argues that DPG is strictly more sample-efficient here. However, Ganesh notices that the current DPG implementation uses only on-policy data sampled under [[IMAGE:0d29d23c93e001a7_9_52]] , and that the robot operates in a noisy, stochastic environment where some states are rarely visited under the greedy policy. Based on the above data, answer the given subquestions.
      Ganesh’s on-policy DPG implementation is likely to suffer from which of the following problems? Select all that apply.
      Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
      1. Poor coverage of the state space — the greedy policy may never visit states needed for accurate [[IMAGE:0d29d23c93e001a7_10_55]] estimation
        Source diagram or notation
      2. Maximization bias, because the greedy policy always selects the action with highest estimated [[IMAGE:0d29d23c93e001a7_10_56]]
        Source diagram or notation
      3. The DPG theorem technically requires an off-policy behaviour policy to ensure sufficient state coverage; on-policy data violates this
      4. Insufficient exploration, since [[IMAGE:0d29d23c93e001a7_10_57]] is deterministic and produces no stochastic variation in actions
        Source diagram or notation
      5. The action-value gradient [[IMAGE:0d29d23c93e001a7_10_58]] cannot be computed on-policy
        Source diagram or notation

      A published solution is not available for this question yet.

      Question 15 NAT · 0.0 marks

      Ganesh is designing a robot arm controller with a 6-dimensional continuous action space. He initially uses a stochastic policy [[IMAGE:0d29d23c93e001a7_9_48]] with the standard policy gradient(SPG): [[IMAGE:0d29d23c93e001a7_9_49]] Karthik suggests switching to the Deterministic Policy Gradient (DPG) with policy [[IMAGE:0d29d23c93e001a7_9_50]] : [[IMAGE:0d29d23c93e001a7_9_51]] Karthik argues that DPG is strictly more sample-efficient here. However, Ganesh notices that the current DPG implementation uses only on-policy data sampled under [[IMAGE:0d29d23c93e001a7_9_52]] , and that the robot operates in a noisy, stochastic environment where some states are rarely visited under the greedy policy. Based on the above data, answer the given subquestions.
      **[Bonus]** Ganesh adds Gaussian exploration noise [[IMAGE:0d29d23c93e001a7_10_59]] , [[IMAGE:0d29d23c93e001a7_10_60]] , with [[IMAGE:0d29d23c93e001a7_10_61]] . The DPG actor update for a single sample is: [[IMAGE:0d29d23c93e001a7_10_62]] Given [[IMAGE:0d29d23c93e001a7_10_63]] , and for a particular state [[IMAGE:0d29d23c93e001a7_10_64]] : [[IMAGE:0d29d23c93e001a7_10_65]] Compute the [[IMAGE:0d29d23c93e001a7_10_66]] -norm of the parameter update vector [[IMAGE:0d29d23c93e001a7_10_67]] . Round to 4 decimal places.
      Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

        A published solution is not available for this question yet.

        Question 16 MSQ · 3.0 marks

        A Bernoulli-logistic unit is a stochastic neuron whose output [[IMAGE:0d29d23c93e001a7_11_68]] follows a Bernoulli distribution: [[IMAGE:0d29d23c93e001a7_11_69]] The unit receives a feature vector [[IMAGE:0d29d23c93e001a7_11_70]] as input. Let [[IMAGE:0d29d23c93e001a7_11_71]] denote the action preference for action [[IMAGE:0d29d23c93e001a7_11_72]] in state [[IMAGE:0d29d23c93e001a7_11_73]] under policy parameter [[IMAGE:0d29d23c93e001a7_11_74]] . Assume that difference in action preferences is a linear function of the input: [[IMAGE:0d29d23c93e001a7_11_75]] The softmax (exponential soft-max) policy converts preferences to probabilities: [[IMAGE:0d29d23c93e001a7_11_76]] The REINFORCE (Monte Carlo policy gradient) update rule is: [[IMAGE:0d29d23c93e001a7_11_77]] where [[IMAGE:0d29d23c93e001a7_11_78]] is the return from time [[IMAGE:0d29d23c93e001a7_11_79]] , and [[IMAGE:0d29d23c93e001a7_11_80]] is called the eligibility vector. Based on the above data, answer the given subquestions.
        Applying the softmax distribution to the Bernoulli-logistic unit, which of the following correctly expresses [[IMAGE:0d29d23c93e001a7_12_81]] ?
        Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
        1. [[IMAGE:0d29d23c93e001a7_12_82]]
          Source diagram or notation
        2. [[IMAGE:0d29d23c93e001a7_12_83]]
          Source diagram or notation
        3. [[IMAGE:0d29d23c93e001a7_12_84]]
          Source diagram or notation
        4. [[IMAGE:0d29d23c93e001a7_12_85]]
          Source diagram or notation

        A published solution is not available for this question yet.

        Question 17 MSQ · 0.0 marks

        A Bernoulli-logistic unit is a stochastic neuron whose output [[IMAGE:0d29d23c93e001a7_11_68]] follows a Bernoulli distribution: [[IMAGE:0d29d23c93e001a7_11_69]] The unit receives a feature vector [[IMAGE:0d29d23c93e001a7_11_70]] as input. Let [[IMAGE:0d29d23c93e001a7_11_71]] denote the action preference for action [[IMAGE:0d29d23c93e001a7_11_72]] in state [[IMAGE:0d29d23c93e001a7_11_73]] under policy parameter [[IMAGE:0d29d23c93e001a7_11_74]] . Assume that difference in action preferences is a linear function of the input: [[IMAGE:0d29d23c93e001a7_11_75]] The softmax (exponential soft-max) policy converts preferences to probabilities: [[IMAGE:0d29d23c93e001a7_11_76]] The REINFORCE (Monte Carlo policy gradient) update rule is: [[IMAGE:0d29d23c93e001a7_11_77]] where [[IMAGE:0d29d23c93e001a7_11_78]] is the return from time [[IMAGE:0d29d23c93e001a7_11_79]] , and [[IMAGE:0d29d23c93e001a7_11_80]] is called the eligibility vector. Based on the above data, answer the given subquestions.
        **[Bonus]** The Monte Carlo REINFORCE update for the Bernoulli-logistic unit upon receipt of return [[IMAGE:0d29d23c93e001a7_12_86]] is: [[IMAGE:0d29d23c93e001a7_12_87]] Which of the following statements about this update are correct? Select all that apply.
        Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
        1. When [[IMAGE:0d29d23c93e001a7_12_88]] , the eligibility vector is [[IMAGE:0d29d23c93e001a7_12_89]] , so the update reinforces [[IMAGE:0d29d23c93e001a7_12_90]] in the direction of [[IMAGE:0d29d23c93e001a7_12_91]] when [[IMAGE:0d29d23c93e001a7_12_92]]
          Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
        2. When [[IMAGE:0d29d23c93e001a7_12_93]] , the eligibility vector is [[IMAGE:0d29d23c93e001a7_12_94]] , meaning the update decreases [[IMAGE:0d29d23c93e001a7_12_95]] (i.e. makes action 1 less likely) when [[IMAGE:0d29d23c93e001a7_12_96]]
          Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
        3. The magnitude of the update is larger when the selected action is surprising (low probability), since [[IMAGE:0d29d23c93e001a7_12_97]] is large when [[IMAGE:0d29d23c93e001a7_12_98]] and [[IMAGE:0d29d23c93e001a7_12_99]] is large when [[IMAGE:0d29d23c93e001a7_12_100]]
          Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
        4. The update is an unbiased estimate of [[IMAGE:0d29d23c93e001a7_12_101]] only if the episode is generated on-policy under [[IMAGE:0d29d23c93e001a7_12_102]]
          Source diagram or notationSource diagram or notation
        5. Subtracting a baseline [[IMAGE:0d29d23c93e001a7_13_103]] from [[IMAGE:0d29d23c93e001a7_13_104]] changes the direction of the gradient update, introducing bias to improve convergence speed
          Source diagram or notationSource diagram or notation

        A published solution is not available for this question yet.

        Question 18 MSQ · 2.0 marks

        Vivek implements the REINFORCE (Monte Carlo policy gradient) algorithm for a text-generation task modelled as an episodic MDP. The policy is a neural network [[IMAGE:0d29d23c93e001a7_13_105]] parameterised by [[IMAGE:0d29d23c93e001a7_13_106]] . At each episode's end, he collects the full trajectory [[IMAGE:0d29d23c93e001a7_13_107]] and computes the return: [[IMAGE:0d29d23c93e001a7_13_108]] The REINFORCE gradient estimator is: [[IMAGE:0d29d23c93e001a7_13_109]] After 1000 episodes, he observes that the gradient variance is extremely high and training is unstable. Neel suggests adding a baseline [[IMAGE:0d29d23c93e001a7_13_110]] : [[IMAGE:0d29d23c93e001a7_13_111]] He claims: "Subtracting [[IMAGE:0d29d23c93e001a7_13_112]] reduces variance without biasing the gradient, as long as [[IMAGE:0d29d23c93e001a7_13_113]] does not depend on [[IMAGE:0d29d23c93e001a7_13_114]] .'' Based on the above data, answer the given subquestions.
        Why does REINFORCE exhibit high gradient variance, particularly in long-horizon tasks like text generation?
        Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
        1. The log-probability [[IMAGE:0d29d23c93e001a7_14_115]] is numerically unstable for large neural networks
          Source diagram or notation
        2. [[IMAGE:0d29d23c93e001a7_14_116]] accumulates rewards over the entire future horizon; in long episodes with noisy rewards, [[IMAGE:0d29d23c93e001a7_14_117]] has large variance, which directly scales the gradient estimate
          Source diagram or notationSource diagram or notation
        3. Monte Carlo rollouts are biased estimators of the true return
        4. The policy gradient theorem does not hold for episodic tasks with discount factor [[IMAGE:0d29d23c93e001a7_14_118]]
          Source diagram or notation

        A published solution is not available for this question yet.

        Question 19 NAT · 2.0 marks

        Vivek implements the REINFORCE (Monte Carlo policy gradient) algorithm for a text-generation task modelled as an episodic MDP. The policy is a neural network [[IMAGE:0d29d23c93e001a7_13_105]] parameterised by [[IMAGE:0d29d23c93e001a7_13_106]] . At each episode's end, he collects the full trajectory [[IMAGE:0d29d23c93e001a7_13_107]] and computes the return: [[IMAGE:0d29d23c93e001a7_13_108]] The REINFORCE gradient estimator is: [[IMAGE:0d29d23c93e001a7_13_109]] After 1000 episodes, he observes that the gradient variance is extremely high and training is unstable. Neel suggests adding a baseline [[IMAGE:0d29d23c93e001a7_13_110]] : [[IMAGE:0d29d23c93e001a7_13_111]] He claims: "Subtracting [[IMAGE:0d29d23c93e001a7_13_112]] reduces variance without biasing the gradient, as long as [[IMAGE:0d29d23c93e001a7_13_113]] does not depend on [[IMAGE:0d29d23c93e001a7_13_114]] .'' Based on the above data, answer the given subquestions.
        For a single timestep [[IMAGE:0d29d23c93e001a7_14_119]] in an episode, the following values are observed: [[IMAGE:0d29d23c93e001a7_14_120]] Compute the parameter update [[IMAGE:0d29d23c93e001a7_14_121]] and report [[IMAGE:0d29d23c93e001a7_14_122]] . Round to 4 decimal places.
        Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

          A published solution is not available for this question yet.

          Question 20 MSQ · 2.0 marks

          Vivek implements the REINFORCE (Monte Carlo policy gradient) algorithm for a text-generation task modelled as an episodic MDP. The policy is a neural network [[IMAGE:0d29d23c93e001a7_13_105]] parameterised by [[IMAGE:0d29d23c93e001a7_13_106]] . At each episode's end, he collects the full trajectory [[IMAGE:0d29d23c93e001a7_13_107]] and computes the return: [[IMAGE:0d29d23c93e001a7_13_108]] The REINFORCE gradient estimator is: [[IMAGE:0d29d23c93e001a7_13_109]] After 1000 episodes, he observes that the gradient variance is extremely high and training is unstable. Neel suggests adding a baseline [[IMAGE:0d29d23c93e001a7_13_110]] : [[IMAGE:0d29d23c93e001a7_13_111]] He claims: "Subtracting [[IMAGE:0d29d23c93e001a7_13_112]] reduces variance without biasing the gradient, as long as [[IMAGE:0d29d23c93e001a7_13_113]] does not depend on [[IMAGE:0d29d23c93e001a7_13_114]] .'' Based on the above data, answer the given subquestions.
          Neel claims that any function [[IMAGE:0d29d23c93e001a7_15_123]] not depending on [[IMAGE:0d29d23c93e001a7_15_124]] is a valid zero-bias baseline. Which of the following statements are correct justifications or counter-examples to this claim?
          Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
          1. [[IMAGE:0d29d23c93e001a7_15_125]] for any [[IMAGE:0d29d23c93e001a7_15_126]] independent of [[IMAGE:0d29d23c93e001a7_15_127]] , since [[IMAGE:0d29d23c93e001a7_15_128]]
            Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
          2. If [[IMAGE:0d29d23c93e001a7_15_129]] , the baseline perfectly cancels the return and the gradient becomes zero — a degenerate case
            Source diagram or notation
          3. A baseline that depends on both [[IMAGE:0d29d23c93e001a7_15_130]] and [[IMAGE:0d29d23c93e001a7_15_131]] (e.g. [[IMAGE:0d29d23c93e001a7_15_132]] ) still leaves the gradient unbiased
            Source diagram or notationSource diagram or notationSource diagram or notation
          4. Using [[IMAGE:0d29d23c93e001a7_15_133]] introduces bias because [[IMAGE:0d29d23c93e001a7_15_134]] is estimated, not known exactly
            Source diagram or notationSource diagram or notation

          A published solution is not available for this question yet.

          Question 21 MCQ · 2.0 marks

          Lalit uses two algorithms to learn option values in a given programming assignment: Algorithm A — SMDP Q-Learning: Updates [[IMAGE:0d29d23c93e001a7_15_135]] only at option termination using the full option return: [[IMAGE:0d29d23c93e001a7_15_136]] Algorithm B — Intra-Option Q-Learning: Updates [[IMAGE:0d29d23c93e001a7_15_137]] at every primitive step using the one-step consistency condition: [[IMAGE:0d29d23c93e001a7_15_138]] [[IMAGE:0d29d23c93e001a7_15_139]] where [[IMAGE:0d29d23c93e001a7_16_140]] is the termination probability at the next state. The agent runs in a corridor with 5 states [[IMAGE:0d29d23c93e001a7_16_141]] . Option [[IMAGE:0d29d23c93e001a7_16_142]] has intra-option policy: move right always. [[IMAGE:0d29d23c93e001a7_16_143]] if [[IMAGE:0d29d23c93e001a7_16_144]] , else [[IMAGE:0d29d23c93e001a7_16_145]] . Currently: [[IMAGE:0d29d23c93e001a7_16_146]] , [[IMAGE:0d29d23c93e001a7_16_147]] , [[IMAGE:0d29d23c93e001a7_16_148]] , [[IMAGE:0d29d23c93e001a7_16_149]] , [[IMAGE:0d29d23c93e001a7_16_150]] , [[IMAGE:0d29d23c93e001a7_16_151]] , agent moves from [[IMAGE:0d29d23c93e001a7_16_152]] to [[IMAGE:0d29d23c93e001a7_16_153]] . Based on the above data, answer the given subquestions.
          What is the key advantage of Intra-Option Q-Learning (Algorithm B) over SMDP Q-Learning (Algorithm A)?
          Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
          1. It requires no knowledge of the termination function [[IMAGE:0d29d23c93e001a7_16_154]]
            Source diagram or notation
          2. It updates [[IMAGE:0d29d23c93e001a7_16_155]] at every time step rather than only at termination, making more efficient use of experience — especially for long-duration options
            Source diagram or notation
          3. It converges to a globally optimal policy, whereas SMDP Q-learning only converges locally
          4. It eliminates the need for a high-level policy by learning all options simultaneously

          A published solution is not available for this question yet.

          Question 22 NAT · 2.0 marks

          Lalit uses two algorithms to learn option values in a given programming assignment: Algorithm A — SMDP Q-Learning: Updates [[IMAGE:0d29d23c93e001a7_15_135]] only at option termination using the full option return: [[IMAGE:0d29d23c93e001a7_15_136]] Algorithm B — Intra-Option Q-Learning: Updates [[IMAGE:0d29d23c93e001a7_15_137]] at every primitive step using the one-step consistency condition: [[IMAGE:0d29d23c93e001a7_15_138]] [[IMAGE:0d29d23c93e001a7_15_139]] where [[IMAGE:0d29d23c93e001a7_16_140]] is the termination probability at the next state. The agent runs in a corridor with 5 states [[IMAGE:0d29d23c93e001a7_16_141]] . Option [[IMAGE:0d29d23c93e001a7_16_142]] has intra-option policy: move right always. [[IMAGE:0d29d23c93e001a7_16_143]] if [[IMAGE:0d29d23c93e001a7_16_144]] , else [[IMAGE:0d29d23c93e001a7_16_145]] . Currently: [[IMAGE:0d29d23c93e001a7_16_146]] , [[IMAGE:0d29d23c93e001a7_16_147]] , [[IMAGE:0d29d23c93e001a7_16_148]] , [[IMAGE:0d29d23c93e001a7_16_149]] , [[IMAGE:0d29d23c93e001a7_16_150]] , [[IMAGE:0d29d23c93e001a7_16_151]] , agent moves from [[IMAGE:0d29d23c93e001a7_16_152]] to [[IMAGE:0d29d23c93e001a7_16_153]] . Based on the above data, answer the given subquestions.
          Using Algorithm B (Intra-Option Q-Learning) and the values from the comprehension section, compute the TD error [[IMAGE:0d29d23c93e001a7_16_156]] for the transition [[IMAGE:0d29d23c93e001a7_16_157]] . Round your answer to 4 decimal places. Note: [[IMAGE:0d29d23c93e001a7_16_158]] (the option does not terminate at [[IMAGE:0d29d23c93e001a7_16_159]] ). [[IMAGE:0d29d23c93e001a7_16_160]]
          Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

            A published solution is not available for this question yet.

            Question 23 NAT · 2.0 marks

            Lalit uses two algorithms to learn option values in a given programming assignment: Algorithm A — SMDP Q-Learning: Updates [[IMAGE:0d29d23c93e001a7_15_135]] only at option termination using the full option return: [[IMAGE:0d29d23c93e001a7_15_136]] Algorithm B — Intra-Option Q-Learning: Updates [[IMAGE:0d29d23c93e001a7_15_137]] at every primitive step using the one-step consistency condition: [[IMAGE:0d29d23c93e001a7_15_138]] [[IMAGE:0d29d23c93e001a7_15_139]] where [[IMAGE:0d29d23c93e001a7_16_140]] is the termination probability at the next state. The agent runs in a corridor with 5 states [[IMAGE:0d29d23c93e001a7_16_141]] . Option [[IMAGE:0d29d23c93e001a7_16_142]] has intra-option policy: move right always. [[IMAGE:0d29d23c93e001a7_16_143]] if [[IMAGE:0d29d23c93e001a7_16_144]] , else [[IMAGE:0d29d23c93e001a7_16_145]] . Currently: [[IMAGE:0d29d23c93e001a7_16_146]] , [[IMAGE:0d29d23c93e001a7_16_147]] , [[IMAGE:0d29d23c93e001a7_16_148]] , [[IMAGE:0d29d23c93e001a7_16_149]] , [[IMAGE:0d29d23c93e001a7_16_150]] , [[IMAGE:0d29d23c93e001a7_16_151]] , agent moves from [[IMAGE:0d29d23c93e001a7_16_152]] to [[IMAGE:0d29d23c93e001a7_16_153]] . Based on the above data, answer the given subquestions.
            Using the value of [[IMAGE:0d29d23c93e001a7_17_161]] computed in previous sub-question, calculate the updated value of [[IMAGE:0d29d23c93e001a7_17_162]] . Round your answer to 4 decimal places. [[IMAGE:0d29d23c93e001a7_17_163]]
            Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

              A published solution is not available for this question yet.