MauryaHub PYQ Practice

da5007_2026T1_Q1_NA.pdf

Reinforcement Learning · Quiz 1 · Jan 2026

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 2 MCQ · 2.0 marks

Which of the following is/are correct and valid reasons to consider sampling actions from a softmax distribution instead of using an [[IMAGE:847d867b2fa2b5da_3_2]] greedy approach? 1.Under softmax exploration, the probability of selecting an action increases with its estimated action value, which reduces unnecessary exploration of clearly inferior actions. 2.Unlike the [[IMAGE:847d867b2fa2b5da_3_3]] -greedy method, softmax exploration does not require careful, gradual decay of the exploration parameter and still yields asymptotically correct behaviour even if the temperature is reduced sharply. 3.It enables more fine-grained discrimination among actions whose estimated Q-values are close to the maximum, allowing more nuanced preference for slightly better actions. Which of the above statements is/are correct?
Source diagram or notationSource diagram or notation
  1. 1, 2, 3
  2. only 3
  3. 1, 2
  4. 1, 3
  5. 3, 2
  6. Only 2

A published solution is not available for this question yet.

Question 3 MCQ · 2.0 marks

Consider a discounted return: [[IMAGE:847d867b2fa2b5da_4_4]] in an infinite-horizon MDP with bounded rewards [[IMAGE:847d867b2fa2b5da_4_5]] and discount factor [[IMAGE:847d867b2fa2b5da_4_6]] 1.For fixed rewards, the contribution of [[IMAGE:847d867b2fa2b5da_4_7]] to [[IMAGE:847d867b2fa2b5da_4_8]] decays geometrically as [[IMAGE:847d867b2fa2b5da_4_9]] . 2.If [[IMAGE:847d867b2fa2b5da_4_10]] and rewards are uniformly bounded, the infinite sum [[IMAGE:847d867b2fa2b5da_4_11]] is always finite. 3.For [[IMAGE:847d867b2fa2b5da_4_12]] , the discounted return [[IMAGE:847d867b2fa2b5da_4_13]] is bounded above in magnitude by [[IMAGE:847d867b2fa2b5da_4_14]] . Here, Rmax be the maximum reward for any transition. Which of the above statements is/are correct?
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
  1. 1, 2, 3
  2. 1, 3
  3. 2, 3
  4. only 1
  5. only 3

A published solution is not available for this question yet.

Question 4 MCQ · 1.0 marks

Consider the following assertion and reason pair and select the correct option: **Assertion:**In the UCB algorithm for multi-armed bandits, replacing the upper confidence bound with a lower confidence bound (and greedily selecting actions based on that lower bound) would still promote effective exploration or lead to optimal reward maximisation. **Reason:**Even when lower confidence bounds are used instead of upper bounds, if the algorithm greedily selects arms based on these lower bounds, it would still encourage exploration and ultimately achieve optimal reward maximisation.
  1. Both Assertion and Reason are true, and Reason is the correct explanation of Assertion.
  2. Both Assertion and Reason are true, but Reason is NOT the correct explanation of Assertion.
  3. Assertion is true, Reason is false
  4. Assertion is false, Reason is false

A published solution is not available for this question yet.

Question 5 MCQ · 1.0 marks

In the context of MDPs, what does it mean for a policy [[IMAGE:847d867b2fa2b5da_5_15]] to be greedy with respect to an action- value function [[IMAGE:847d867b2fa2b5da_5_16]] ?
Source diagram or notationSource diagram or notation
  1. For each state [[IMAGE:847d867b2fa2b5da_5_17]] , [[IMAGE:847d867b2fa2b5da_5_18]] selects actions that minimise [[IMAGE:847d867b2fa2b5da_5_19]]
    Source diagram or notationSource diagram or notationSource diagram or notation
  2. For each state [[IMAGE:847d867b2fa2b5da_5_20]] , [[IMAGE:847d867b2fa2b5da_5_21]] selects actions that maximise [[IMAGE:847d867b2fa2b5da_5_22]]
    Source diagram or notationSource diagram or notationSource diagram or notation
  3. For each state [[IMAGE:847d867b2fa2b5da_5_23]] , [[IMAGE:847d867b2fa2b5da_5_24]] ignores [[IMAGE:847d867b2fa2b5da_5_25]] and follows a fixed, pre-defined schedule.
    Source diagram or notationSource diagram or notationSource diagram or notation
  4. For each state [[IMAGE:847d867b2fa2b5da_5_26]] , [[IMAGE:847d867b2fa2b5da_5_27]] selects actions uniformly at random regardless of [[IMAGE:847d867b2fa2b5da_5_28]]
    Source diagram or notationSource diagram or notationSource diagram or notation

A published solution is not available for this question yet.

Question 6 MSQ · 2.0 marks

In policy iteration for finite MDPs, how are the Bellman equations used during the policy evaluation step?
  1. They are ignored; policy iteration does not rely on Bellman equations.
  2. They are used to compute [[IMAGE:847d867b2fa2b5da_5_29]] exactly (or approximately) for the current policy [[IMAGE:847d867b2fa2b5da_5_30]]
    Source diagram or notationSource diagram or notation
  3. They are used to directly compute the optimal policy without evaluating intermediate policies.
  4. They are used to compute immediate rewards without considering transitions.

A published solution is not available for this question yet.

Question 7 MSQ · 2.0 marks

In a [[IMAGE:847d867b2fa2b5da_6_31]] -armed bandit setting, we maintain a running estimate of the action-value for each arm [[IMAGE:847d867b2fa2b5da_6_32]] , denoted [[IMAGE:847d867b2fa2b5da_6_33]] . We now consider using an upper confidence bound (UCB) style rule for arm selection at time [[IMAGE:847d867b2fa2b5da_6_34]] , where [[IMAGE:847d867b2fa2b5da_6_35]] is the number of times arm [[IMAGE:847d867b2fa2b5da_6_36]] has been selected up to time [[IMAGE:847d867b2fa2b5da_6_37]] , and [[IMAGE:847d867b2fa2b5da_6_38]] is a tunable hyperparameter. Which of the following are good strategies for arm selection?
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
  1. [[IMAGE:847d867b2fa2b5da_6_39]]
    Source diagram or notation
  2. [[IMAGE:847d867b2fa2b5da_6_40]]
    Source diagram or notation
  3. [[IMAGE:847d867b2fa2b5da_6_41]]
    Source diagram or notation
  4. [[IMAGE:847d867b2fa2b5da_6_42]]
    Source diagram or notation
  5. None of these

A published solution is not available for this question yet.

Question 8 MSQ · 2.0 marks

Which of the following equations best represents the Bellman optimality equation for the optimal state-value function [[IMAGE:847d867b2fa2b5da_6_43]]
Source diagram or notation
  1. [[IMAGE:847d867b2fa2b5da_6_44]]
    Source diagram or notation
  2. [[IMAGE:847d867b2fa2b5da_6_45]]
    Source diagram or notation
  3. [[IMAGE:847d867b2fa2b5da_7_46]]
    Source diagram or notation
  4. [[IMAGE:847d867b2fa2b5da_7_47]]
    Source diagram or notation

A published solution is not available for this question yet.

Question 9 NAT · 2.0 marks

The sequence of rewards for a continuing task with [[IMAGE:847d867b2fa2b5da_7_48]] is given below: [[IMAGE:847d867b2fa2b5da_7_49]] Find the return [[IMAGE:847d867b2fa2b5da_7_50]] . Your answer should have exactly two places after the decimal point.
Source diagram or notationSource diagram or notationSource diagram or notation

    A published solution is not available for this question yet.

    Question 10 NAT · 2.0 marks

    You have a 3-armed bandit where each arm, when active, yields a Bernoulli reward with success probabilities [[IMAGE:847d867b2fa2b5da_7_51]] , [[IMAGE:847d867b2fa2b5da_7_52]] , and [[IMAGE:847d867b2fa2b5da_7_53]] . However, arm 1 is active in any given round with probability [[IMAGE:847d867b2fa2b5da_7_54]] , and otherwise it yields a reward of 0 regardless of its Bernoulli outcome. Arms 2 and 3 are always active. What is the expected reward from pulling arm 1 in a single round?
    Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

      A published solution is not available for this question yet.

      Question 11 NAT · 1.0 marks

      Consider a Markov Decision Process (MDP) with three states [[IMAGE:847d867b2fa2b5da_8_55]] , [[IMAGE:847d867b2fa2b5da_8_56]] , and [[IMAGE:847d867b2fa2b5da_8_57]] , where [[IMAGE:847d867b2fa2b5da_8_58]] is a terminal state and [[IMAGE:847d867b2fa2b5da_8_59]] . Each non-terminal state has a single available action, and the transitions are deterministic with associated rewards: the agent moves from [[IMAGE:847d867b2fa2b5da_8_60]] to [[IMAGE:847d867b2fa2b5da_8_61]] and receives a reward of 1, and from [[IMAGE:847d867b2fa2b5da_8_62]] to [[IMAGE:847d867b2fa2b5da_8_63]] and receives a reward of 2. Let the discount factor be [[IMAGE:847d867b2fa2b5da_8_64]] . Based on the above data, answer the given subquestions.
      Assume the initial value estimates are [[IMAGE:847d867b2fa2b5da_8_65]] , [[IMAGE:847d867b2fa2b5da_8_66]] , and [[IMAGE:847d867b2fa2b5da_8_67]] . After performing one synchronous policy evaluation update (under the fixed policy that follows the chain), determine [[IMAGE:847d867b2fa2b5da_8_68]] and [[IMAGE:847d867b2fa2b5da_8_69]] . Compute the ratio [[IMAGE:847d867b2fa2b5da_8_70]]
      Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

        A published solution is not available for this question yet.

        Question 12 NAT · 1.0 marks

        Consider a Markov Decision Process (MDP) with three states [[IMAGE:847d867b2fa2b5da_8_55]] , [[IMAGE:847d867b2fa2b5da_8_56]] , and [[IMAGE:847d867b2fa2b5da_8_57]] , where [[IMAGE:847d867b2fa2b5da_8_58]] is a terminal state and [[IMAGE:847d867b2fa2b5da_8_59]] . Each non-terminal state has a single available action, and the transitions are deterministic with associated rewards: the agent moves from [[IMAGE:847d867b2fa2b5da_8_60]] to [[IMAGE:847d867b2fa2b5da_8_61]] and receives a reward of 1, and from [[IMAGE:847d867b2fa2b5da_8_62]] to [[IMAGE:847d867b2fa2b5da_8_63]] and receives a reward of 2. Let the discount factor be [[IMAGE:847d867b2fa2b5da_8_64]] . Based on the above data, answer the given subquestions.
        Starting from [[IMAGE:847d867b2fa2b5da_9_71]] , [[IMAGE:847d867b2fa2b5da_9_72]] , and [[IMAGE:847d867b2fa2b5da_9_73]] . perform two successive synchronous policy evaluation sweeps (under the fixed policy that follows the chain). What is [[IMAGE:847d867b2fa2b5da_9_74]] after these two sweeps?
        Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

          A published solution is not available for this question yet.

          Question 13 NAT · 1.0 marks

          Consider a 2-state MDP with states [[IMAGE:847d867b2fa2b5da_9_75]] and [[IMAGE:847d867b2fa2b5da_9_76]] , discount factor [[IMAGE:847d867b2fa2b5da_9_77]] . There are two actions in [[IMAGE:847d867b2fa2b5da_9_78]] : [[IMAGE:847d867b2fa2b5da_9_79]] and [[IMAGE:847d867b2fa2b5da_9_80]] . Action [[IMAGE:847d867b2fa2b5da_9_81]] gives immediate reward [[IMAGE:847d867b2fa2b5da_9_82]] and moves deterministically to [[IMAGE:847d867b2fa2b5da_9_83]] ; action [[IMAGE:847d867b2fa2b5da_9_84]] gives immediate reward [[IMAGE:847d867b2fa2b5da_9_85]] and keeps the agent in [[IMAGE:847d867b2fa2b5da_9_86]] . In [[IMAGE:847d867b2fa2b5da_9_87]] , there is a single action that yields an immediate reward [[IMAGE:847d867b2fa2b5da_9_88]] and transitions back to [[IMAGE:847d867b2fa2b5da_9_89]] . Assume initial value estimates are [[IMAGE:847d867b2fa2b5da_9_90]] , [[IMAGE:847d867b2fa2b5da_9_91]] Based on the above data, answer the given subquestions.
          If you perform a single value iteration Bellman optimality update for [[IMAGE:847d867b2fa2b5da_9_92]] , what is the new value [[IMAGE:847d867b2fa2b5da_9_93]] ?
          Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

            A published solution is not available for this question yet.

            Question 14 NAT · 1.0 marks

            Consider a 2-state MDP with states [[IMAGE:847d867b2fa2b5da_9_75]] and [[IMAGE:847d867b2fa2b5da_9_76]] , discount factor [[IMAGE:847d867b2fa2b5da_9_77]] . There are two actions in [[IMAGE:847d867b2fa2b5da_9_78]] : [[IMAGE:847d867b2fa2b5da_9_79]] and [[IMAGE:847d867b2fa2b5da_9_80]] . Action [[IMAGE:847d867b2fa2b5da_9_81]] gives immediate reward [[IMAGE:847d867b2fa2b5da_9_82]] and moves deterministically to [[IMAGE:847d867b2fa2b5da_9_83]] ; action [[IMAGE:847d867b2fa2b5da_9_84]] gives immediate reward [[IMAGE:847d867b2fa2b5da_9_85]] and keeps the agent in [[IMAGE:847d867b2fa2b5da_9_86]] . In [[IMAGE:847d867b2fa2b5da_9_87]] , there is a single action that yields an immediate reward [[IMAGE:847d867b2fa2b5da_9_88]] and transitions back to [[IMAGE:847d867b2fa2b5da_9_89]] . Assume initial value estimates are [[IMAGE:847d867b2fa2b5da_9_90]] , [[IMAGE:847d867b2fa2b5da_9_91]] Based on the above data, answer the given subquestions.
            Suppose after some value iterations your current estimates are [[IMAGE:847d867b2fa2b5da_10_94]] , [[IMAGE:847d867b2fa2b5da_10_95]] . If you perform one value iteration update for [[IMAGE:847d867b2fa2b5da_10_96]] , what is the new value [[IMAGE:847d867b2fa2b5da_10_97]] ?
            Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

              A published solution is not available for this question yet.

              Question 15 MCQ · 2.0 marks

              Consider an N-armed bandit problem in which each arm [[IMAGE:847d867b2fa2b5da_10_98]] , for [[IMAGE:847d867b2fa2b5da_10_99]] , produces a reward of 1 with probability [[IMAGE:847d867b2fa2b5da_10_100]] and [[IMAGE:847d867b2fa2b5da_10_101]] otherwise, where [[IMAGE:847d867b2fa2b5da_10_102]] denotes the arm’s base success probability. Assume the bandit machine suffers from a hardware malfunction: when the agent attempts to pull arm [[IMAGE:847d867b2fa2b5da_10_103]] , the intended action is not always executed. Instead, with probability [[IMAGE:847d867b2fa2b5da_10_104]] a wiring fault causes the machine to randomly activate one arm chosen uniformly from the set [[IMAGE:847d867b2fa2b5da_10_105]] , regardless of the agent’s selection, and the reward is generated according to that arm’s reward distribution. We refer to this malfunction as [[IMAGE:847d867b2fa2b5da_10_106]] -noise activation Based on the above data, answer the given subquestions.
              Consider a faulty [[IMAGE:847d867b2fa2b5da_10_107]] -armed bandit with [[IMAGE:847d867b2fa2b5da_10_108]] -noise activation and base arm success probabilities [[IMAGE:847d867b2fa2b5da_11_109]] . If you always choose the arm with the highest base success probability [[IMAGE:847d867b2fa2b5da_11_110]] , what is the expected payoff per pull under this greedy strategy?
              Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
              1. [[IMAGE:847d867b2fa2b5da_11_111]]
                Source diagram or notation
              2. [[IMAGE:847d867b2fa2b5da_11_112]]
                Source diagram or notation
              3. [[IMAGE:847d867b2fa2b5da_11_113]]
                Source diagram or notation
              4. [[IMAGE:847d867b2fa2b5da_11_114]]
                Source diagram or notation

              A published solution is not available for this question yet.

              Question 16 MCQ · 2.0 marks

              Consider an N-armed bandit problem in which each arm [[IMAGE:847d867b2fa2b5da_10_98]] , for [[IMAGE:847d867b2fa2b5da_10_99]] , produces a reward of 1 with probability [[IMAGE:847d867b2fa2b5da_10_100]] and [[IMAGE:847d867b2fa2b5da_10_101]] otherwise, where [[IMAGE:847d867b2fa2b5da_10_102]] denotes the arm’s base success probability. Assume the bandit machine suffers from a hardware malfunction: when the agent attempts to pull arm [[IMAGE:847d867b2fa2b5da_10_103]] , the intended action is not always executed. Instead, with probability [[IMAGE:847d867b2fa2b5da_10_104]] a wiring fault causes the machine to randomly activate one arm chosen uniformly from the set [[IMAGE:847d867b2fa2b5da_10_105]] , regardless of the agent’s selection, and the reward is generated according to that arm’s reward distribution. We refer to this malfunction as [[IMAGE:847d867b2fa2b5da_10_106]] -noise activation Based on the above data, answer the given subquestions.
              Consider a faulty [[IMAGE:847d867b2fa2b5da_11_115]] -armed bandit with [[IMAGE:847d867b2fa2b5da_11_116]] -noise activation and distinct base success probabilities [[IMAGE:847d867b2fa2b5da_11_117]] . Suppose the agent follows a greedy strategy that always chooses the arm with the highest empirical mean reward. How does the presence of activation noise affect the asymptotic fraction of time for which the optimal arm (with success probability [[IMAGE:847d867b2fa2b5da_11_118]] ) is actually executed?
              Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
              1. Noise has no effect; the optimal arm is activated almost surely in the long run.
              2. Noise reduces the frequency with which the optimal arm is selected, but once it is selected, it is always activated.
              3. Noise eventually causes the greedy policy to oscillate indefinitely between arms, preventing convergence.
              4. Even when the greedy policy converges to always selecting the optimal arm, activation noise prevents the optimal arm from being activated more than a fraction [[IMAGE:847d867b2fa2b5da_11_119]] of the time.
                Source diagram or notation

              A published solution is not available for this question yet.

              Question 17 MCQ · 2.0 marks

              Consider an N-armed bandit problem in which each arm [[IMAGE:847d867b2fa2b5da_10_98]] , for [[IMAGE:847d867b2fa2b5da_10_99]] , produces a reward of 1 with probability [[IMAGE:847d867b2fa2b5da_10_100]] and [[IMAGE:847d867b2fa2b5da_10_101]] otherwise, where [[IMAGE:847d867b2fa2b5da_10_102]] denotes the arm’s base success probability. Assume the bandit machine suffers from a hardware malfunction: when the agent attempts to pull arm [[IMAGE:847d867b2fa2b5da_10_103]] , the intended action is not always executed. Instead, with probability [[IMAGE:847d867b2fa2b5da_10_104]] a wiring fault causes the machine to randomly activate one arm chosen uniformly from the set [[IMAGE:847d867b2fa2b5da_10_105]] , regardless of the agent’s selection, and the reward is generated according to that arm’s reward distribution. We refer to this malfunction as [[IMAGE:847d867b2fa2b5da_10_106]] -noise activation Based on the above data, answer the given subquestions.
              In a faulty [[IMAGE:847d867b2fa2b5da_12_120]] -armed bandit with [[IMAGE:847d867b2fa2b5da_12_121]] -noise activation, suppose you mistakenly model the environment as a standard bandit without activation noise. You estimate each arm’s success probability using sample averages of observed rewards. How does the [[IMAGE:847d867b2fa2b5da_12_122]] -noise affect your estimates of the true base success probabilities [[IMAGE:847d867b2fa2b5da_12_123]] ?
              Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
              1. Your estimates are biased toward the overall average success probability across all arms.
              2. Your estimates are biased upward, overestimating all [[IMAGE:847d867b2fa2b5da_12_124]]
                Source diagram or notation
              3. Your estimates remain unbiased for [[IMAGE:847d867b2fa2b5da_12_125]] because you still observe correct rewards for the arm you believe you pulled.
                Source diagram or notation
              4. Your estimates are biased downward, underestimating all [[IMAGE:847d867b2fa2b5da_12_126]]
                Source diagram or notation

              A published solution is not available for this question yet.

              Question 18 MCQ · 2.0 marks

              Consider a large-scale news recommendation platform that uses a contextual bandit approach to decide which headline to show to a user visiting the homepage. The context includes user features (e.g., location, device type, time of day, coarse interest profile) and article features (e.g., topic, recency, source reputation). Actions correspond to recommending one of the [[IMAGE:847d867b2fa2b5da_12_127]] candidate headlines, and the reward is [[IMAGE:847d867b2fa2b5da_12_128]] if the user clicks the recommended headline and [[IMAGE:847d867b2fa2b5da_12_129]] otherwise. Let the context vector [[IMAGE:847d867b2fa2b5da_12_130]] at time [[IMAGE:847d867b2fa2b5da_12_131]] encode both user and article features for the current recommendation opportunity. There are [[IMAGE:847d867b2fa2b5da_12_132]] discrete actions, each corresponding to selecting one of the candidate headlines for display. Based on the above data, answer the given subquestions.
              Which of the following best explains why a contextual bandit approach is preferred over a standard multi-armed bandit in this news recommendation scenario?
              Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
              1. Because user preferences are identical for all users, so context does not influence click behaviour.
              2. Because contextual bandits can tailor recommendations to each user and article pair using the context, thereby improving click-through rates across heterogeneous users.
              3. Because standard multi-armed bandits cannot perform any exploration at all.
              4. Because contextual bandits guarantee that the same headline is always optimal for every user.

              A published solution is not available for this question yet.

              Question 19 MCQ · 2.0 marks

              Consider a large-scale news recommendation platform that uses a contextual bandit approach to decide which headline to show to a user visiting the homepage. The context includes user features (e.g., location, device type, time of day, coarse interest profile) and article features (e.g., topic, recency, source reputation). Actions correspond to recommending one of the [[IMAGE:847d867b2fa2b5da_12_127]] candidate headlines, and the reward is [[IMAGE:847d867b2fa2b5da_12_128]] if the user clicks the recommended headline and [[IMAGE:847d867b2fa2b5da_12_129]] otherwise. Let the context vector [[IMAGE:847d867b2fa2b5da_12_130]] at time [[IMAGE:847d867b2fa2b5da_12_131]] encode both user and article features for the current recommendation opportunity. There are [[IMAGE:847d867b2fa2b5da_12_132]] discrete actions, each corresponding to selecting one of the candidate headlines for display. Based on the above data, answer the given subquestions.
              After many rounds of user interaction, what is the main optimality objective of the contextual bandit algorithm in this recommendation setting?
              Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
              1. Maximise the cumulative expected number of clicks by acting as closely as possible to the best policy that maps contexts to headlines.
              2. Ensure that on every single round, it always picks the empirically highest-CTR headline observed so far, regardless of context.
              3. Minimise the variance of rewards, even if that substantially reduces the total number of clicks.
              4. Focus only on the worst-performing user segment and ignore the rest.

              A published solution is not available for this question yet.