da5007_2023T3_Q1_NA.pdf
Reinforcement Learning · Quiz 1 · Sep 2023
← Course papers · Start practice / exam
This page contains the reliably extracted subset, not the complete original paper.
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 85 MCQ · 1.0 marks
Net profit for the month of April is:
$5920
$5275
$5322
$5800
**RL**
**Section Id :** 64065345317
**Section Number :** 5
**Section type :** Online
**Mandatory or Optional :** Mandatory
**Number of Questions :** 12
**Number of Questions to be attempted :** 12
**Section Marks :** 40
**Display Number Panel :** Yes
**Section Negative Marks :** 0
**Group All Questions :** No
**Enable Mark as Answered Mark for Review and**
Yes
**Clear Response :**
**Maximum Instruction Time :** 0
A published solution is not available for this question yet.
Question 87 MCQ · 0.0 marks
[[IMAGE:19e4e51306e3f6f5_3_0]]

Useful Data has been mentioned above.
This data attachment is just for a reference & not for an evaluation.
A published solution is not available for this question yet.
Question 88 MSQ · 2.0 marks
Select correct statements about UCB1 algorithm:
Each arm is selected atleast twice before any arm is selected third time.
Each arm is selected atleast once before any arm is selected second time.
Constant c controls degree of exploration, higher its value, higher the
exploration.
Constant c controls degree of exploration, higher its value, lower the
exploration.
A published solution is not available for this question yet.
Question 89 MSQ · 2.0 marks
For the value iteration algorithm, which of the following statements are correct?
[[IMAGE:19e4e51306e3f6f5_4_1]]

[[IMAGE:19e4e51306e3f6f5_4_2]]

[[IMAGE:19e4e51306e3f6f5_4_3]]

[[IMAGE:19e4e51306e3f6f5_4_4]]

A published solution is not available for this question yet.
Question 90 MCQ · 2.0 marks
Choose the correct statement(s) regarding explore-exploit dilemma:
Always exploring is not optimal.
Always exploiting is optimal.
None of these.
A published solution is not available for this question yet.
Question 91 MCQ · 2.0 marks
Which of the following is the correct Bellman equation for deterministic transitions? The symbols
have the usual meaning.
[[IMAGE:19e4e51306e3f6f5_5_5]]

[[IMAGE:19e4e51306e3f6f5_5_6]]

[[IMAGE:19e4e51306e3f6f5_5_7]]

[[IMAGE:19e4e51306e3f6f5_5_8]]

A published solution is not available for this question yet.
Question 92 MCQ · 2.0 marks
Consider the following statements and select the correct option.
**Assertion:** Monte Carlo value function approximation methods do not need knowledge of the
model to be implemented.
**Reason:** Monte Carlo value function approximation methods require only a way to sample
trajectories from the environment and aggregate the results.
Assertion and Reason are both true and Reason is a correct explanation of
Assertion.
Assertion and Reason are both true and Reason is not a correct explanation of
Assertion.
Assertion is true but Reason is false.
Assertion is false but Reason is true.
A published solution is not available for this question yet.
Question 93 NAT · 2.0 marks
[[IMAGE:19e4e51306e3f6f5_7_9]]
Based on the above data, answer the given subquestions.
What is the probability of choosing arm A2 at timestamp t = 6?

A published solution is not available for this question yet.
Question 94 MCQ · 2.0 marks
[[IMAGE:19e4e51306e3f6f5_7_9]]
Based on the above data, answer the given subquestions.
At time stamp t = 6, suppose the arm with least estimate so far is pulled and the reward is ln(4).
Which of the following is correct after timestamp t = 6?

Arm A1 is optimal arm.
Arm A2 is optimal arm.
Arm A3 is optimal arm.
An optimal arm can not be determined.
There is a tie for optimal arm.
A published solution is not available for this question yet.
Question 95 NAT · 2.0 marks
[[IMAGE:19e4e51306e3f6f5_8_10]]

A published solution is not available for this question yet.
Question 96 NAT · 3.0 marks
[[IMAGE:19e4e51306e3f6f5_9_11]]

A published solution is not available for this question yet.
Question 97 MSQ · 2.0 marks
[[IMAGE:19e4e51306e3f6f5_10_12]]
Based on the above data, answer the given subquestions.
Consider the following deterministic policies in table (1) :
[[IMAGE:19e4e51306e3f6f5_11_13]]
Which of the following are optimal policies?


π1
π2
π3
π4
Can not be determined.
A published solution is not available for this question yet.
Question 98 MSQ · 2.0 marks
[[IMAGE:19e4e51306e3f6f5_10_12]]
Based on the above data, answer the given subquestions.
Refer to table (1) and choose the correct statements from the following:
[[IMAGE:19e4e51306e3f6f5_12_14]]


π1 = π2
π2 < π3
π1 ≤ π4
π3 ≥ π4
None of these
A published solution is not available for this question yet.
Question 99 NAT · 2.0 marks
[[IMAGE:19e4e51306e3f6f5_10_12]]
Based on the above data, answer the given subquestions.
[[IMAGE:19e4e51306e3f6f5_12_15]]


A published solution is not available for this question yet.
Question 100 NAT · 3.0 marks
[[IMAGE:19e4e51306e3f6f5_10_12]]
Based on the above data, answer the given subquestions.
[[IMAGE:19e4e51306e3f6f5_13_16]]


A published solution is not available for this question yet.
Question 101 MSQ · 2.0 marks
[[IMAGE:19e4e51306e3f6f5_10_12]]
Based on the above data, answer the given subquestions.
[[IMAGE:19e4e51306e3f6f5_13_17]]


1 and 2.
1 and 3.
2 and 3.
2 and 5.
2 and 4.
4 and 5.
A published solution is not available for this question yet.
Question 102 NAT · 3.0 marks
[[IMAGE:19e4e51306e3f6f5_14_18]]
Based on the above data, answer the given subquestions.
What will be the value of v(hot) after one round of value iteration? Assuming v(hot) and v(cold) are
initialized with 0. Note the value function is updated synchronously.

A published solution is not available for this question yet.
Question 103 NAT · 3.0 marks
[[IMAGE:19e4e51306e3f6f5_14_18]]
Based on the above data, answer the given subquestions.
What will be the value of v(cold) after one round of value iteration? Assuming v(hot) and v(cold) are
initialized with 0. Note the value function is updated synchronously.

A published solution is not available for this question yet.
Question 104 NAT · 4.0 marks
[[IMAGE:19e4e51306e3f6f5_14_18]]
Based on the above data, answer the given subquestions.
What will be the value of v(hot) + v(cold) after **two** rounds of value iteration? Assuming v(hot) and
v(cold) are initialized with 0. Note the value function is updated synchronously.

A published solution is not available for this question yet.