Learn from completed experience

Monte Carlo return explorer

Same episodes. Two ways to count visits.

← Back to the chapter

1. Include a completed episode

Replay three fixed examples. This illustrates averaging; it does not demonstrate convergence.

Changing γ recalculates every included episode. Reset keeps the selected discount.

No episodes included yet.

2. Inspect the latest episode

Waiting for a completed episode.

Returns and visit selection for the latest included episode
TimeStateNext rewardReturn GIncluded by
No visits included yet.

A repeated state contributes another sample only to every-visit MC.

3. Average across all included episodes

Values are unknown until a state contributes a return.

Values are rounded to three decimals. Each sample is one selected visit's return.