Episode Returns and Step Rewards
The racetrack is a discrete environment in which both position and velocity are represented using grid-based values.
A Race Measured in Steps
The racetrack turns driving around a bend into a reinforcement-learning problem. Reaching the finish is necessary, but it is not the only goal: the car should finish in as few time steps as possible while avoiding paths that leave the track. Moving aggressively can shorten the route, but a boundary collision sends the car back to the starting line and the episode continues.
The task balances speed and safety. A successful policy must learn both how to move quickly and how to avoid track-boundary collisions.
Position and Velocity on the Grid
The racetrack is discrete rather than continuous. The car occupies one cell from a finite collection of grid positions. Its velocity is also represented with discrete values: the horizontal and vertical components describe how many grid cells the car moves in those directions at a time step.
Reading One Discrete State
Suppose a car is represented as occupying one grid cell and has a horizontal velocity component of 2 and a vertical velocity component of 1.
Position: The car's location is one cell in the finite racetrack grid, not an arbitrary point in continuous space.
Velocity: The two components describe movement in the horizontal and vertical directions in terms of grid cells.
Next movement: The car's projected path is determined from its velocity before the environment updates its location.
A state combines a discrete grid location with discrete horizontal and vertical velocity components.
Nine Ways to Adjust Motion
An action does not directly select a destination cell. It changes the two velocity components. The horizontal component can receive an increment of minus 1, 0, or plus 1, and the vertical component has the same three choices. Combining those choices gives nine possible acceleration actions.
| Horizontal increment | Vertical increment | Resulting action meaning |
|---|---|---|
| -1 | -1 | Decrease both components |
| -1 | 0 | Decrease horizontal component only |
| -1 | +1 | Decrease horizontal and increase vertical component |
| 0 | -1 | Decrease vertical component only |
| 0 | 0 | Leave both intended increments at zero |
| 0 | +1 | Increase vertical component only |
| +1 | -1 | Increase horizontal and decrease vertical component |
| +1 | 0 | Increase horizontal component only |
| +1 | +1 | Increase both components |
The nine actions come from combining three horizontal increment choices with three vertical increment choices.
From Action to Episode Outcome
A time step follows an important order. First, the environment considers the velocity change, including possible noise that overrides the intended increments. It then projects the car's movement and checks whether that projected path intersects the finish line or the track boundary. Only after this check is the car's location updated.
Tracing a Boundary Collision
A car takes an action and its projected path intersects the track boundary away from the finish line.
Check the projected path: The environment checks the path before updating the car's location.
Classify the intersection: Because the intersection is with the boundary rather than the finish line, it is a nonterminal collision.
Reset the car: The car moves to a randomly selected position on the starting line, and both velocity components become zero.
Continue the episode: The collision does not end the episode. Later time steps continue to receive the minus 1 reward until the finish line is crossed.
A boundary collision costs additional time and restarts the car's motion without ending the episode.
Rewards Across a Complete Episode
Every time step before the finish crossing produces a reward of minus 1. The episode's accumulated reward therefore records how many penalized steps were required before termination: shorter successful episodes accumulate fewer penalties than longer ones. A collision adds more time steps because the episode continues after the reset.
Comparing Two Episode Outcomes
Compare a successful trajectory that reaches the finish without a collision with one that collides once before eventually finishing.
No collision: The car receives minus 1 at each time step until it crosses the finish line.
One collision: The car receives the step penalties before the collision, resets to the starting line with zero velocity, and receives additional minus 1 rewards during the continuing part of the episode.
Compare returns: The collision trajectory contains additional penalized time steps, so its accumulated reward is lower than that of an otherwise comparable shorter trajectory.
The step penalty makes fewer time steps preferable, while the reset consequence makes boundary collisions costly.
Randomness in the Driving Process
The racetrack is stochastic in several ways. Each episode begins from a randomly selected starting state. If a collision occurs, the reset places the car at a random position on the starting line. In addition, at every time step there is a probability of 0.1 that both velocity increments become zero, regardless of the increments intended by the action.
When inspecting an illustrative trajectory, distinguish between the policy's intended behavior and the environment's random interruptions. The task description allows noise to be turned off for example trajectories, making the displayed path easier to inspect.
Mistakes in Episode Tracing
Treating an action as a destination choice
An action changes the horizontal and vertical velocity components; it does not directly name a destination cell.
Fix:
Apply the component increments first, then reason about the projected path.Allowing velocity components outside the permitted range
The resulting velocity components must be nonnegative and less than 5.
Fix:
Check the component restrictions after considering the velocity adjustment.Ending the episode at every boundary intersection
Only a finish-line intersection ends the episode. A boundary collision resets the car and the episode continues.
Fix:
Move the car to a random starting-line position, set both velocity components to zero, and continue collecting step rewards.Ignoring the projected-path check
The path is checked before the location update, and that ordering distinguishes a finish from a nonterminal collision.
Fix:
Trace the sequence as velocity change, projected-path check, then outcome.Assuming the chosen increments always occur
With probability 0.1, both increments are replaced by zero increments.
Fix:
Include the possibility of the random zero-increment event when reasoning about actual trajectories.
Monte Carlo Control Connection
Monte Carlo control is a natural fit because the racetrack supplies complete episodes: an episode starts from a starting state and ends when the car crosses the finish line. The learner can evaluate the accumulated step rewards from complete trajectories and use those episodes to seek a policy for each starting state.
Trace this situation in words: the car takes an action, the projected path intersects the boundary rather than the finish line, and the reset occurs. Identify the car's new position, its velocity, whether the episode has ended, and what happens to rewards on subsequent steps.
Hints
- The reset position is randomly selected on the starting line.
- Both velocity components become zero.
- A boundary collision does not terminate the episode.
- Subsequent time steps continue to receive the minus 1 reward until the finish line is crossed.
Explain why a policy that reaches the finish quickly but often collides may be worse than a slightly slower policy that stays on the track. Use the step reward and the continuing nature of collision resets in your explanation.
Hints
- Every pre-finish time step receives minus 1.
- A collision adds further time steps to the same episode.
- The task balances speed against safety.
Key Takeaways
- The car occupies a discrete grid cell, and its horizontal and vertical velocity components are discrete as well.
- The nine actions come from combining horizontal and vertical increments of minus 1, 0, or plus 1, subject to nonnegative and less-than-5 velocity-component restrictions.
- The projected path is checked before the location update: a finish-line intersection ends the episode, while another boundary intersection causes a reset and continuation.
- Every pre-finish step receives minus 1, so shorter successful episodes accumulate fewer penalties and collisions add further cost.
- Random starting states, random collision-reset positions, and the probability of zero increments make trajectories stochastic and provide complete episodes for Monte Carlo control.
Key Takeaways
- The racetrack represents both position and velocity on a discrete grid.
- Each action adjusts two velocity components through one of nine increment combinations.
- The environment checks the projected path before updating location, distinguishing terminal finishes from continuing boundary collisions.
- Minus 1 step rewards favor shorter successful episodes, while collisions add further penalized steps.
- Random starts, random resets, and zero-increment noise make complete racetrack episodes suitable for Monte Carlo control.