Files
NODEDC_MISSION_CORE/docs/audits/2026-09-19-recovery-stationary.md
DCCONSTRUCTIONS e515ab1b8c feat(planning): consolidate recorded-route localization and spatial scene
Preserve the completed teach-and-repeat laboratory stage: reference preparation, cascaded acquisition, local tracking and recovery, recording lifecycle, replay qualification, and persistent Rerun scene controls. Document the open grid-picking regression and Rerun upgrade contract. No autonomous driving or loop-closure optimization is claimed.
2026-09-21 08:47:19 +03:00

184 lines
11 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Recovery and stationary entry — frozen protocol v1
This follow-up uses immutable A/B recordings only. No hardware commands, new
capture, vehicle authority or UI changes. All numerical jobs are sequential,
one child / CPU thread. Prior failed and successful traces remain unchanged.
## Receipt recovery
Prepare one fixed forward A submap from its first pose through the last pose
whose recorded path length is <=40 m. Same extraction v1 and 120-frame budget;
no B data used to select or build this map. Replay B at 1x, <=120 s / 40 m.
First run unchanged acquisition v1 on this reference. Then suppress all B
pose/cloud events with original relative receipt time in [44,47) seconds.
Do not shift timestamps, add poses, replay stale frames or change surviving
payloads. Save identities/times of every suppressed event. This is an injected
transport loss on real recorded geometry, not a new physical experiment.
Acceptance: qualified tracking must exist before 44 s; the resumed receipt gap
must clear old windows and qualification. Reacquisition must use a fresh bounded
search, never retain the pre-gap matrix as current, followed by three fresh
consistent windows from the new segment. Record the exact loss/recovery times,
seed, latency and in-route distances. Failure or insufficient time is retained,
not repaired by loosening thresholds or copying an old fitted transform.
The existing 8 s result freshness policy remains unchanged for this baseline;
a <8 s silence may retain the last fresh solution until the receipt gap is
observed. This diagnostic delay is not a navigation safety policy.
## Stationary prefix experiment
B's first 10 s contain approximately 0.005 m maximum recorded displacement. Freeze only
that prefix, rejecting it if displacement exceeds 0.10 m. Use the existing
bounded accumulator and fixed A40 reference. No B heading from later motion,
offline B registration, or future points. Positions establish only the prefix's
map-frame entry; no independent physical stillness/accuracy is claimed.
New stationary-entry/v1: nine offsets {-3,0,3} m along/across the reference
heading, and twelve yaw seeds 0..330 degrees at 30-degree intervals = 108 seeds.
Yaw rotates about query entry. Full yaw is searched because the K1 map-frame
heading is not assumed to match the taught route. Initial translation places
the prefix entry at the selected reference entry. Final translation bounds
remain 5 m XY / 1 m Z. No vertical or arbitrary-site search. Every seed uses
unchanged local GICP v1, including its 3 m /30-degree local correction limits.
Reuse complete-link solution clustering, support and ambiguity gates; at least
3 supporting seeds and 2 translation positions. Require all 108 hypotheses;
25 s soft/30 s hard deadline. Incomplete search is rejected. Run the same prefix
against the already fixed wrong A130–155 m region as a negative control. If the
correct prefix is ambiguous or rejected, record that limitation. Stationary
snapshot acceptance alone is not temporal tracking or operational readiness.
## Validation and decision gate
Focused fixtures cover fault boundaries and unchanged event identity/time,
post-gap three-window qualification, no old segment authority, prefix-only
selection, motion rejection, yaw-anchor preservation and full-search rejection.
Record UTC/monotonic times, input/code/output hashes, all seeds and clusters.
Do not install the stationary initializer into the product merely because one
snapshot looks good. Decide the next increment from positive and negative
evidence. No new operator scan is needed for these bounded experiments.
## Results: controlled loss and recovery
Private experiment root: `data_dir/missions/causal-replays/20260919-recovery-stationary-001`.
The fixed A interval is 39.926109 m and 52,965 retained points. A is session
`20260911T085226Z_viewer_live`; B is `20260911T134352Z_viewer_live`. All prior A/B
artifacts and failed/successful traces remain unchanged. The reference derives
solely from A; B geometry does not select reference frames.
Baseline: 2026-09-19T12:01:13.261Z–12:02:16.501Z; start monotonic ns
836950850212541. 974 delivered events, 63.125 s, 40.038 m. Initial candidate
33.100 s, first tracking 42.343 s. Six accepted windows. Maximum delivery lag
0.1672 s. This reproduces the previous outcome on the explicitly longer A map.
Loss probe: 2026-09-19T12:02:37.218Z–12:03:40.468Z; start monotonic ns
837034809904750. 60 events suppressed, 914 delivered, 63.123 s / 40.027 m;
maximum delivery lag 0.1525 s. No survivor payload or timestamp was changed.
Original 9.973 s gap remains; the injected interval produces a second observed
pose gap of 3.331 s / 2.569 m displacement.
| Event | Time from input start, s | State / evidence |
| --- | ---: | --- |
| Three-window qualification | 42.358 | tracking, segment 1 |
| Deliberately absent input | 44–47 | last result still within existing 8 s freshness limit |
| Resumed receipt reveals gap | 47.097 | lost; old matrix/windows cleared, segment 2 |
| New bounded search completes | 49.637 | acquiring, new-segment candidate 1 |
| Next fresh local refinement | 53.178 | acquiring, candidate 2 |
| Third new-segment window | 58.602 | tracking restored |
| Input ends at distance bound | 63.122 | lost; no current result retained |
The first post-gap candidate used a new 27-seed search, not the old fitted
transform. It completed 2.540 s after observed resumption, at input age 2.572 s.
Post-gap windows were at 26.138, 31.198 and 35.911 m, all within the selected A
path, with overlap 99.09%, 98.55%, 98.02% and RMSE 0.1453, 0.1536, 0.1571 m.
Requalification took 11.505 s from observed resumption. Thus recovery from this
injected receipt loss is demonstrated, with the diagnostic 5 s fit cadence.
This does not qualify a moving vehicle's reaction time. The short silence was
recognized on the next pose; immediate silence detection is not claimed.
## Results: heading-free stationary prefix
The prefix contained 168 pose/cloud events through 9.933688 s and 8,227 retained
points. Maximum recorded displacement was **0.0051505 m**. This is K1's own pose
estimate, not an independent physical stillness measurement. No later B points,
direction, fitted transform or correspondence mask was used. The wrong-region
entry and basis came directly from A's saved trajectory, not a previous B seed.
The two sequential snapshot probes ran 2026-09-19T12:04:04.176Z–12:04:46.942Z;
launcher monotonic interval 837121521919208–837164567013166 ns. The experiment
uses a frozen already-collected prefix; it is not a second wall-paced replay.
| Control | Hypotheses | Compute wall time, s | Result |
| --- | ---: | ---: | --- |
| Correct A entry | 108/108 | 24.548 | candidate; one eligible cluster, 7 seeds |
| Wrong A130–155 m entry | 108/108 | 16.560 | rejected; no eligible cluster |
Correct-entry overlap 89.815% within 0.5 m; inlier RMSE 0.17042 m. These are
surface-fit statistics, not localization accuracy. Seven agreeing initial
seeds are not seven independent observations. After both runs, a read-only
comparison with the earlier causal A40 candidate showed 0.10251 m difference
at the entry and 0.78413 degrees rotation. That comparison was not an input or
gate, and is not ground truth.
An operational implementation would need about 10 s accumulation plus 25 s
search on this host/input, before fresh validation. The successful search was
close to the unchanged 25 s soft budget; a slower/ambiguous search must reject.
Its old prefix already exceeds the current 8 s freshness limit on completion.
Therefore **the snapshot candidate is not activated as live tracking**. It
needs a separate provisional initialization result and fresh-window validation,
followed by the temporal gate. The three-metre-heading requirement remains in
the currently active product profile; stationary entry is an engineering probe.
## Implementation and verification
- `replay_faults.py`: explicit lazy time-window filtering, preserving survivor
identity/timestamps and an audit of dropped sequence numbers.
- `stationary_entry.py`: bounded prefix, motion/gap/completeness checks and
`stationary-entry/v1` full-yaw policy. No heading extraction from B motion.
- `entry_acquisition.py`: the existing clustering and search accept an explicit
policy; the default travel-entry constants remain unchanged. Complete-search
cardinality comes from the policy's translation/yaw grid.
- `entry_acquisition_worker.py`: optional stationary mode in the same isolated
worker, retaining one thread and a 30 s hard limit. Existing callers use the
default travel mode. No hardware path or product selector added.
- `check_planning_recovery.py`: fixed-reference preparation, separate baseline,
loss and stationary stages, provenance and before/after digest verification.
- `test_recovery_stationary.py`: fault boundaries, identity/time preservation,
prefix-only geometry, motion rejection, 108 anchor-preserving hypotheses,
partial-search rejection, gap reset and three new-segment candidates.
54 focused tests passed in 3.12 s, covering entry, causal replay, live profile,
project results, registration and viewer replay; one existing Starlette/httpx
deprecation warning. Ruff and `git diff --check` passed. Numerical jobs ran
sequentially on Darwin arm64 / Python 3.12.13 / small_gicp 1.0.1, one CPU thread.
No load test, Docker VM, new capture or vehicle commands. Frontend unchanged.
All probe processes completed. The existing canonical server was left running;
its `/api/health` returned operational `ok` after the experiments. No second
server was started or product connection restarted. The engineering results and
acceptance checker were saved in MISSIONCOR-81, preserving 40 structured blocks.
All input digests, reference artifact, executed-source hashes and 31/31/7
baseline/loss/stationary artifact hashes were rechecked. Exact executed sources
and focused tests are stored privately in `executed-source/`, with `seal.json`.
Report SHA-256:
- baseline: `0c4b3d4f096f0deb065bfd21e0917f0414772861c8c6c72f44380d3f6e794f5d`
- loss: `2ac13eae870a4c4eac5d0d95d35bdb72ab46c84e4a10bf29285fae593d6bcafb`
- stationary: `1390f4b5ceb3188ff89c9a70ff34fb20cb02140335483f13ba88571a80ff2815`
- seal: `619aec16a6f5921cdfbec2f2e7fda520bd6756c8c4caaf60f26a0f97b7083339`
## Decision
Controlled receipt-loss recovery passed on this recording. Stationary geometry
supports a bounded heading-free initialization candidate on this pair, with
the tested wrong region rejected. Neither result establishes arbitrary-site
localization, independent pose accuracy, repeatability across locations or
onboard resource qualification.
Next increment: make the stationary search result a provisional prior only;
validate against a current cloud before temporal qualification, while the input
continues. Separate startup waiting from maximum age while moving. Reduce
repeated fixed-reference preparation if measured helpful, retaining frozen
regression and ambiguity controls. Then qualify starts/orientations on held-out
data and the physical camera/UI path. No new operator scan is needed yet.