# HW2 experiment log (2026-10-04) `compare.py` was edited between rounds, so the candidate lists below are not all reproducible from the current file. Numbers are copied from the console output of each round in the session. Maps: `sim.make_maps(per_combo, seed=2026, goal_mode)`. Columns: open / dead_ends / loops_dense / ALL / failures / collisions. ## Round 1 — 40 maps per combo (360 maps), seed 2026 goal_mode=far | strategy | open | dead | loops | ALL | fail | coll | |---|---|---|---|---|---|---| | DFS | 84.28 | 91.76 | 86.57 | 87.53 | 0 | 0 | | nearest | 87.81 | 90.90 | 88.46 | 89.06 | 0 | 0 | | nearest+unk0.5 | 88.17 | 90.88 | 88.42 | 89.16 | 0 | 0 | | far0.2 | 89.94 | 91.29 | 89.80 | 90.34 | 0 | 0 | | far0.4 | 89.93 | 91.29 | 89.83 | 90.35 | 0 | 0 | | far0.7 | 90.04 | 91.29 | 89.85 | 90.39 | 0 | 0 | | far1.0 | 90.04 | 91.29 | 89.86 | 90.39 | 0 | 0 | goal_mode=uniform: DFS 88.35, nearest 88.70, nearest+unk0.5 88.87, far0.2–1.0 ≈ 89.01–89.02 (all 0 fail). ## Round 2 — 100 per combo (900 maps), seed 2026 far: nearest 89.20; far0.05 90.43; far0.1 90.43; far0.3 90.45; far1.0 90.50; far3.0 89.99; far0.3 nodir 90.23; far0.3+unk0.5 90.48; opt0.1 90.21; opt0.3 90.23; opt1.0 90.28 (all 0 fail). uniform: nearest 88.09; far0.05–1.0 88.15–88.17; far3.0 87.77; far0.3 nodir 88.32; far0.3+unk0.5 88.25; opt* 88.17–88.19. Decision: drop "optimistic" distance metric (no gain, slower). ## Round 3 — ValueAgent without commitment, 900 maps, seed 2026 | strategy | far ALL | far fail | uniform ALL | uniform fail | |---|---|---|---|---| | far0.3 | 90.45 | 0 | 88.16 | 0 | | far1.0 | 90.50 | 0 | 88.17 | 0 | | val p0 c2 | 89.02 | 2 | 88.37 | 4 | | val p2 c2 | 90.39 | 2 | 88.81 | 1 | | val p2 c6 | 89.67 | 5 | 87.82 | 11 | | val p4 c2 | 91.08 | 0 | 88.76 | 1 | | val p4 c6 | 90.35 | 4 | 87.74 | 11 | Failures came from target switching back and forth (oscillation). ## Round 4 — add commitment, 900 maps, seed 2026 | strategy | far ALL | uniform ALL | fail | |---|---|---|---| | far1.0 | 90.50 | 88.17 | 0 | | far1.0 commit | 90.50 | 88.17 | 0 | | val p4 c2 (no commit) | 91.08 | 88.76 | 0 / 1 | | val p2 c2 commit | 90.52 | 88.84 | 0 | | val p4 c2 commit | 91.01 | 88.80 | 0 | | val p4 c1 commit | 91.12 | 88.84 | 0 | | val p6 c2 commit | 91.31 | 88.76 | 0 | ## Round 5 — tuning, 900 maps, seed 2026 (current compare.py) | strategy | far ALL | uniform ALL | fail | |---|---|---|---| | val p6 c2 commit | 91.31 | 88.76 | 0 | | val p6 c1 commit | 91.33 | 88.74 | 0 | | val p6 c3 commit | 91.16 | 88.70 | 0 | | val p8 c2 commit | 91.21 | 88.68 | 0 | | val p10 c2 commit | 91.15 | 88.61 | 0 | Decision: POWER = 6, COST_OFFSET = 2. ## Final agent.py checks - `final_check.py 100 2026`: far 91.31 / uniform 88.76 (identical to "val p6 c2 commit"), 0 fail, max step use 44% / 42%. - `final_check.py 300 777` (different seed from tuning): far 91.14 / uniform 88.72, 0 fail, 0 collisions, max step use 42% / 40%, lowest score 81.24 / 80.09. - `report_numbers.py` and `distribution.py` also use seed 777. ## Round 6 — safety valve (after the adversarial review), 2026-10-05 Reviewer found a crafted 15x15 map where POWER=6 used 812/900 steps. `variants.py` (Variant(power, safe)): after `safe * 4*rows*cols` steps, choose the nearest frontier. "adv worst" = max step usage over the reviewer's 60 maps x all goals (`climb_all.jsonl`). | variant | adv worst | far ALL (seed 2026, 300/combo) | uniform ALL | |---|---|---|---| | current p6 | 90.2% | 91.10 | 88.86 | | p6 safe0.05 | 29.7% | 91.18 | 88.96 | | p6 safe0.10 | 35.5% | 91.16 | 88.90 | Same conclusion on seed 777 (100/combo: p3, safe 0.10–0.40, p4 safe0.30) and seed 4242 (300/combo: safe 0.05/0.075/0.10/0.15). Decision: SAFE_FRACTION = 0.05 added to agent.py (old version kept as ../agent_v1_before_safety.py). After the change: - `equiv_check.py`: agent.py == Variant(6.0, 0.05) on 720 maps (0 step differences). - `adv_check.py`: 29.7% worst on the 60 adversarial maps. - `final_check.py 300 9090` (fresh seed): far 91.16 / uniform 88.95, 0 fail, 0 collisions, max step use 27% / 26%, lowest 81.19 / 80.15. - `report_numbers.py` (seed 777): final far 91.24 / uniform 88.78. - `distribution.py`: median per-map 90.9, 11% perfect, 9-map exam p5 88.4 / median 91.2 / p95 94.2.