← Paper overview

BVER: Bidirectional Voronoi-biased Exploration Curriculum for Reinforcement Learning

Iteration — 5 phases

iteration 0 goals reached 0 / 1
history live

Maze stage

Mode

Walk seeding — connect / voronoi
Policy simulation
Envs & horizons
Import / export config
Export writes the current maze + all settings as JSON. To make one the boot default, paste it into the SHIPPED_DEFAULT constant near the top of this file's script.

Legend

Race: forward-only vs backward-only vs bidirectional

connect ratios apply to the bidirectional pane; the single-direction panes force them to 0 (no opposite tree). Maze & env settings mirror the Explain tab — edit there, then return.

Forward-only

iter 0goals 0 / 1
all goals reached at iteration —

Backward-only

iter 0goals 0 / 1
all goals reached at iteration —

Bidirectional

iter 0goals 0 / 1
all goals reached at iteration —

goals reached vs. iteration (mean over seeds)

covered region vs. iteration: fraction of free cells covered (mean over seeds)

Algorithm 1 — BVER (equation numbers refer to the paper)
Algorithm 2 — Random walk frontier expansion
Algorithm 3 — Return-based thresholding of rollouts into frontier and solved buffers
Algorithm 4 — Rebuild of the mixed set from frontier and solved buffers