Preserve the completed teach-and-repeat laboratory stage: reference preparation, cascaded acquisition, local tracking and recovery, recording lifecycle, replay qualification, and persistent Rerun scene controls. Document the open grid-picking regression and Rerun upgrade contract. No autonomous driving or loop-closure optimization is claimed.
9.2 KiB
Causal planning replay — protocol frozen before execution
Owner scope: teach a known route manually, then repeat it. Approximately 30 s of initial preparation is acceptable; preparation latency and tracking latency are separate measurements. No new field capture is required for this increment.
Inputs and boundaries
Reference: the immutable forward 30.012 m A submap from saved run
820571ba-6076-482b-be34-29ceb5328c80, including its source hashes. Only its
reference cloud and reference trajectory may be reused, never its query cloud,
fitted transform, correspondence mask or final B heading. Query: original B raw
transport and real monotonic receipt metadata, verified against archived hashes.
Reference preparation precedes the playback clock and is reported separately.
One bounded functional replay at a time on the operator Mac, original 1× receipt cadence, at most 120 s / 40 m, one single-thread numeric child, at most 24 fits, 30 s worker timeout, existing bounded live cloud accumulator. This is not a load test or onboard timing qualification. No capture ingress, hardware command, controller, API server restart or product UI changes are needed.
Fixed comparisons
- Baseline: existing route-entry / first 3 m travel heading for every fit.
- Tracking candidate: after an accepted fresh fit, initialize the next fit from
that transform in the same K1 coordinate frame. GICP acceptance policy v1 is
unchanged. Reject a continuation changing the current scanner's mapped
position by more than 0.5 m or orientation by more than 5 degrees. Require
three consecutive compatible fresh candidates before lab state
tracking. These are predeclared experimental gates, not vehicle safety thresholds.
Both modes keep >=5 s between fit requests; only already received query events enter a snapshot. Results refer to the snapshot's last input timestamp, not to completion time. A result older than 8 s cannot establish current tracking. Receipt gaps >2 s clear the cloud window and all tracking authority; a fit from the prior segment cannot be accepted after the gap. No final offline fit is used for initialization. End of input removes current tracking authority.
Measure first available heading, first candidate, first sustained tracking, per-step extraction/snapshot/worker time, input-to-result age, lateness relative to original receipts, accepted/rejected windows, correction discontinuities, receipt gaps and retained input sequence/hash lineage. Do not call residual or agreement with another fit ground-truth pose accuracy.
Negative controls and decision
Use a deliberately wrong reference road patch selected by a fixed disjoint trajectory interval and a synthetic far-offset initial hypothesis. Exercise stale result, gap fencing, repeated inconsistent candidates and prefix-only causality with small deterministic fixtures. Retain rejected reports too.
Advance to a short physical test only if the replay explains acquisition, tracking and loss without future data or persistent false green. Failure yields a diagnosed next engineering step, not relaxed gates. The current first-3-m heading still requires movement: stationary acquisition and arbitrary-start recognition remain separate unqualified work.
Executed functional probes
Two sequential 1× replays completed on the local Darwin arm64 / Python 3.12.13 environment. Each consumed 974 pose/cloud events over approximately 63.12 s and stopped at the fixed 40 m bound (40.038 m using the live accumulator's 5 cm pose thinning). This distance differs slightly from the full archived pose sum and is not ground truth. The last two windows extend beyond the selected reference trajectory's 30 m end; they are coverage-limit diagnostics, not in-route acceptance. No K1 command, new capture, application server restart, Docker startup or inference was used. The source and old result checksums were unchanged.
| Measurement | Entry-hint baseline | Previous-accepted-result mode |
|---|---|---|
| Numerical windows | 6 | 6 |
| Accepted candidates | 0 | 0 |
| First eligible heading/window | 30.812 s | 30.812 s |
| GICP calculation range | 0.063–0.351 s | 0.072–0.333 s |
| Latest input to completed result | 0.290–0.863 s | 0.288–0.863 s |
| Maximum replay delivery lateness | 0.154 s | 0.145 s |
| Sustained tracking | Not established | Not established |
The second mode never reached its previous-result branch: no initial candidate passed the existing acceptance policy. These are two functional traces, not a throughput benchmark, full live-capture/UI qualification or Mini capacity claim. Loading the already prepared, hash-verified reference took 0.023–0.028 s; this does not replace the earlier measurement of preparing a reference from raw data.
Cause of rejection
B has a 9.973 s receipt gap ending at 30.803 s, with 10.407 m between adjacent received poses. The live accumulator correctly clears its query window and starts a new segment. The first usable direction appears only after that gap: there was no demonstrated stationary startup or stable acquisition in 30 s.
All six windows exceeded the fixed 3 m correction at the observed patch center: 3.146, 3.258, 3.469, 3.638, 3.770 and 4.105 m. The fourth also failed convergence. The corresponding heading correction is approximately 6.95–6.98 degrees. The 3 m policy is not a measured capture radius or the distance from the real rover to the reference start.
Despite rejection, nearest-surface coverage was 92.62–98.72%, with inlier RMSE 0.148–0.179 m. A post-run diagnostic of the rejected transforms found entry-point corrections around 2.745–2.759 m; the moving patch-center correction grows with distance under the initial yaw error. Consecutive rejected transforms change the current scanner's mapped position by at most 0.016 m and orientation by at most 0.143 degrees. This is consistency of rejected geometric fits, not measured pose accuracy, accepted tracking, or grounds for silently widening the gate. Rejected results never initialize the tracking mode.
Negative controls
The fixed third causal snapshot (22.046 m of thinned path, requested at 42.015 s) was reused in two separate controls. Neither uses the final offline B fit.
- Hint translated +1000 m on all axes: rejected, 0% overlap, residual unavailable, insufficient geometric information; calculation 1.928 s.
- Different A road patch at cumulative trajectory 130–155 m, seeded at that patch's entry using only the snapshot's preceding path: rejected by residual 0.279 m, overlap 56.16%; calculation 0.125 s. This narrowly passing overlap reinforces why green cannot depend on overlap alone. One wrong region is not a false-positive rate estimate.
Implementation and verification
device_plugins/xgrids_k1/planning_live.py: extracted the existing decoder so committed live input and archive replay use the same point/pose interpretation.planning_replay.py: lazy recorded events, mandatory aligned monotonic receipt metadata, no synthetic clock or hardware transport.missions/causal_replay.py: receipt-paced bounded replay, same live accumulator, one isolated fit worker, causal snapshots, per-step provenance and hashes.missions/causal_tracking.py: separate experimental freshness/consistency evaluator. It is not activated in the product live controller by this work.scripts/replay_planning_registration.pyandscripts/check_planning_replay_controls.py: explicit private inputs, checksum validation before/after, immutable experiment outputs.tests/test_causal_planning_replay.py: prefix-only geometry, stale and inconsistent fits, three-candidate qualification, gap fencing, mandatory real clock, and late completion cannot restore authority after input end.
All 32 focused causal/live/registration/native-replay tests passed; targeted Ruff checks passed. A subsequent terminal-state ordering fix clears authority at input end before waiting for the final child; a regression asserts this ordering. In both real traces all jobs completed before input end, so that fix does not alter their results. No frontend code changed or build was needed.
Private evidence is under data_dir/missions/causal-replays/:
20260919-baseline-001, 20260919-tracking-001, 20260919-controls-001 and
20260919-executed-source-001. The latter seals the executed source before
post-run formatting plus all three report hashes. Raw geometry stays out of Git.
Decision and next bounded experiment
Numerical latency is not the observed blocker at this bounded cloud size. Robust initial acquisition is. The owner-accepted startup wait gives time to evaluate several predeclared entry-position/heading hypotheses and reject ambiguity; it does not justify accepting a failed local fit. A new acquisition experiment must keep this failed baseline, distinguish the entry search area from the local tracking correction limit, and require consistent subsequent causal windows. Repeat wrong-region controls with that initializer before admitting field work. Do not use the known final B transformation as its seed or truth. The same saved recordings remain sufficient for the next engineering iteration. Stationary start, recovery after already-established tracking, broader locations and actual onboard compute remain unqualified.