Concepts / Episode Returns and Step Rewards

Episode Returns and Step Rewards

The racetrack is a discrete environment in which both position and velocity are represented using grid-based values.

  • Programming

A Race Measured in Steps

The racetrack turns driving around a bend into a reinforcement-learning problem. Reaching the finish is necessary, but it is not the only goal: the car should finish in as few time steps as possible while avoiding paths that leave the track. Moving aggressively can shorten the route, but a boundary collision sends the car back to the starting line and the episode continues.

The task balances speed and safety. A successful policy must learn both how to move quickly and how to avoid track-boundary collisions.

Position and Velocity on the Grid

The racetrack is discrete rather than continuous. The car occupies one cell from a finite collection of grid positions. Its velocity is also represented with discrete values: the horizontal and vertical components describe how many grid cells the car moves in those directions at a time step.

state also recordsstate also recordsGrid positionone track cellHorizontal velocitygrid-cell movementVertical velocitygrid-cell movement
How are the car's grid position and horizontal and vertical velocity represented at one moment?

Reading One Discrete State

Suppose a car is represented as occupying one grid cell and has a horizontal velocity component of 2 and a vertical velocity component of 1.

Position: The car's location is one cell in the finite racetrack grid, not an arbitrary point in continuous space.

Velocity: The two components describe movement in the horizontal and vertical directions in terms of grid cells.

Next movement: The car's projected path is determined from its velocity before the environment updates its location.

A state combines a discrete grid location with discrete horizontal and vertical velocity components.

Nine Ways to Adjust Motion

An action does not directly select a destination cell. It changes the two velocity components. The horizontal component can receive an increment of minus 1, 0, or plus 1, and the vertical component has the same three choices. Combining those choices gives nine possible acceleration actions.

Horizontal incrementVertical incrementResulting action meaning
-1-1Decrease both components
-10Decrease horizontal component only
-1+1Decrease horizontal and increase vertical component
0-1Decrease vertical component only
00Leave both intended increments at zero
0+1Increase vertical component only
+1-1Increase horizontal and decrease vertical component
+10Increase horizontal component only
+1+1Increase both components

The nine actions come from combining three horizontal increment choices with three vertical increment choices.

(-1, -1)horizontal -1, vertical -1(-1, 0)horizontal -1, vertical 0(-1, +1)horizontal -1, vertical +1(0, -1)horizontal 0, vertical -1(0, 0)horizontal 0, vertical 0(0, +1)horizontal 0, vertical +1(+1, -1)horizontal +1, vertical -1(+1, 0)horizontal +1, vertical 0(+1, +1)horizontal +1, vertical +1
How does each action change the horizontal and vertical velocity components?

From Action to Episode Outcome

A time step follows an important order. First, the environment considers the velocity change, including possible noise that overrides the intended increments. It then projects the car's movement and checks whether that projected path intersects the finish line or the track boundary. Only after this check is the car's location updated.

thencheck firstyesnoyesnoVelocity adjustmentintended change or noiseProjected pathbefore location updateFinish lineintersects?Episode endsfinish reachedTrack boundaryintersects?Starting linerandom position, zerovelocityContinue episodelocation updated
What happens when the projected path reaches the finish line or intersects the track boundary?

Tracing a Boundary Collision

A car takes an action and its projected path intersects the track boundary away from the finish line.

Check the projected path: The environment checks the path before updating the car's location.

Classify the intersection: Because the intersection is with the boundary rather than the finish line, it is a nonterminal collision.

Reset the car: The car moves to a randomly selected position on the starting line, and both velocity components become zero.

Continue the episode: The collision does not end the episode. Later time steps continue to receive the minus 1 reward until the finish line is crossed.

A boundary collision costs additional time and restarts the car's motion without ending the episode.

Rewards Across a Complete Episode

Every time step before the finish crossing produces a reward of minus 1. The episode's accumulated reward therefore records how many penalized steps were required before termination: shorter successful episodes accumulate fewer penalties than longer ones. A collision adds more time steps because the episode continues after the reset.

nextmay collidereset and continueeventuallycontributescontributescontributesTime step 1reward -1Time step 2reward -1Collisionepisode continuesFurther stepsmore rewards -1Finish crossingepisode endsEpisode returnaccumulated step rewards
How do rewards collected at successive time steps combine into the accumulated reward for a complete episode?

Comparing Two Episode Outcomes

Compare a successful trajectory that reaches the finish without a collision with one that collides once before eventually finishing.

No collision: The car receives minus 1 at each time step until it crosses the finish line.

One collision: The car receives the step penalties before the collision, resets to the starting line with zero velocity, and receives additional minus 1 rewards during the continuing part of the episode.

Compare returns: The collision trajectory contains additional penalized time steps, so its accumulated reward is lower than that of an otherwise comparable shorter trajectory.

The step penalty makes fewer time steps preferable, while the reset consequence makes boundary collisions costly.

Randomness in the Driving Process

The racetrack is stochastic in several ways. Each episode begins from a randomly selected starting state. If a collision occurs, the reset places the car at a random position on the starting line. In addition, at every time step there is a probability of 0.1 that both velocity increments become zero, regardless of the increments intended by the action.

then chooseapply with possible noisethen projectmay intersect boundarymay cross finishresetcontinue episodeEpisode startrandom starting stateChosen actionintended incrementsZero incrementsprobability 0.1Projected movementtrajectory depends onapplied changeBoundary collisionif boundary is intersectedRandom start-linepositionvelocity becomes zeroFinish crossingepisode ends
Where can randomness affect the car's trajectory, and how does that uncertainty unfold across an episode?

When inspecting an illustrative trajectory, distinguish between the policy's intended behavior and the environment's random interruptions. The task description allows noise to be turned off for example trajectories, making the displayed path easier to inspect.

Mistakes in Episode Tracing

  • Treating an action as a destination choice

    An action changes the horizontal and vertical velocity components; it does not directly name a destination cell.

    Fix: Apply the component increments first, then reason about the projected path.

  • Allowing velocity components outside the permitted range

    The resulting velocity components must be nonnegative and less than 5.

    Fix: Check the component restrictions after considering the velocity adjustment.

  • Ending the episode at every boundary intersection

    Only a finish-line intersection ends the episode. A boundary collision resets the car and the episode continues.

    Fix: Move the car to a random starting-line position, set both velocity components to zero, and continue collecting step rewards.

  • Ignoring the projected-path check

    The path is checked before the location update, and that ordering distinguishes a finish from a nonterminal collision.

    Fix: Trace the sequence as velocity change, projected-path check, then outcome.

  • Assuming the chosen increments always occur

    With probability 0.1, both increments are replaced by zero increments.

    Fix: Include the possibility of the random zero-increment event when reasoning about actual trajectories.

Monte Carlo Control Connection

Monte Carlo control is a natural fit because the racetrack supplies complete episodes: an episode starts from a starting state and ends when the car crosses the finish line. The learner can evaluate the accumulated step rewards from complete trajectories and use those episodes to seek a policy for each starting state.

EASY

Trace this situation in words: the car takes an action, the projected path intersects the boundary rather than the finish line, and the reset occurs. Identify the car's new position, its velocity, whether the episode has ended, and what happens to rewards on subsequent steps.

Hints
  • The reset position is randomly selected on the starting line.
  • Both velocity components become zero.
  • A boundary collision does not terminate the episode.
  • Subsequent time steps continue to receive the minus 1 reward until the finish line is crossed.
MEDIUM

Explain why a policy that reaches the finish quickly but often collides may be worse than a slightly slower policy that stays on the track. Use the step reward and the continuing nature of collision resets in your explanation.

Hints
  • Every pre-finish time step receives minus 1.
  • A collision adds further time steps to the same episode.
  • The task balances speed against safety.

Key Takeaways

  1. The car occupies a discrete grid cell, and its horizontal and vertical velocity components are discrete as well.
  2. The nine actions come from combining horizontal and vertical increments of minus 1, 0, or plus 1, subject to nonnegative and less-than-5 velocity-component restrictions.
  3. The projected path is checked before the location update: a finish-line intersection ends the episode, while another boundary intersection causes a reset and continuation.
  4. Every pre-finish step receives minus 1, so shorter successful episodes accumulate fewer penalties and collisions add further cost.
  5. Random starting states, random collision-reset positions, and the probability of zero increments make trajectories stochastic and provide complete episodes for Monte Carlo control.

Key Takeaways

  • The racetrack represents both position and velocity on a discrete grid.
  • Each action adjusts two velocity components through one of nine increment combinations.
  • The environment checks the projected path before updating location, distinguishing terminal finishes from continuing boundary collisions.
  • Minus 1 step rewards favor shorter successful episodes, while collisions add further penalized steps.
  • Random starts, random resets, and zero-increment noise make complete racetrack episodes suitable for Monte Carlo control.