WorksheetsRL-Quiz 2
Total questions: 6
Worksheet time: 4mins
Name
Class
Date
1.
The multi-armed bandit problem is a generalized use case for-
a)
Supervised Learning
b)
Reinforcement learning
c)
Unsupervised Learning
d)
Active Learning
2.
_________Reinforcement is defined as when an event, occurs due to a particular behavior.
a)
+ Ve
b)
- Ve
c)
0
d)
Neutral
3.
There are _______ types of reinforcement
a)
5
b)
4
c)
3
d)
2
4.
_______is all about making decisions sequentially
a)
supervised Learning
b)
Active learning
c)
Reinforcement Learning
d)
unsupervised Learning
5.
If the environment is completely observable, then its dynamic can be modeled as _________
a)
Q Learning
b)
Value
c)
Markov Process
d)
Poicy
6.
Full form of SARSA is (a)
100 %
