Epsilon-greedy simulator

Exploration vs. exploitation on a 5-armed Bernoulli bandit

Pulls
0
Avg reward
—
Optimal action
—
Regret
0.00
← Back to the chapter

The arms

This run

  • Average reward
  • % optimal action

Step at least twice to plot this run.

Compare epsilon values

A single run is mostly luck. This averages many independent runs so the curves separate.

Not run yet.

    View the simulation source