Preserve the completed teach-and-repeat laboratory stage: reference preparation, cascaded acquisition, local tracking and recovery, recording lifecycle, replay qualification, and persistent Rerun scene controls. Document the open grid-picking regression and Rerun upgrade contract. No autonomous driving or loop-closure optimization is claimed.
7.0 KiB
Physical stationary start 005: rejected acceptance
Date: 2026-09-19. Related cards: MISSIONCOR-81 and canonical K1 integration MISSIONCOR-3. This is a diagnosis, not a deployed correction or live acceptance.
Decision
The new physical attempt failed. The previous archive-only qualification is not sufficient evidence that planning works alongside the physical camera, acquisition publisher and browser. Do not request another field walk until the combined path has been qualified. No scanner command, network reconfiguration, runtime restart, threshold relaxation or product code change was performed during this diagnosis.
Exact evidence
Private sealed directory, outside Git:
NODEDC_MISSION_CORE/.runtime/mission-core/missions/diagnostics/20260919-physical-start-005.
It retains run/state snapshots, calculation input/output, raw MQTT with receipt
metadata, camera init/index/segments, selected browser/server diagnostics, offline
calculation outputs and the camera decode log. The root manifest hashes each file.
Physical run 47dc5323-5d2c-4743-a13f-b4f837cf9fea, query
20260919T150053Z_viewer_live (ja-sun-005), reference
20260911T085226Z_viewer_live (JA-SADOVAYA-001), first 30.0117 m.
The reference selection is supported; the user did not have to choose reference002.
Frozen registration input SHA-256:
b72e1f202a04c5455aa10d2907a3912bfb3fbf76b84a079746b23aaa0cc0f9eb.
Localization failure
The stationary ten-second prefix was valid: 200 source events, 9.9745 s span, maximum scanner movement 0.003334 m; 43,854 reference and 7,994 query points. Search started at 15:01:33.389 UTC; the initial-failure transition occurred at 15:02:00.336 UTC. The running code used the prepared-reference v2 fix.
Only 54 of the required 108 seeds completed within the 25-second search budget (25.4795 s). One supported candidate had 96.3758% geometric overlap, 0.151095 m inlier RMSE and five admitted seeds. These are geometric fit statistics, not a localization accuracy guarantee. An incomplete search cannot rule out a competing place hypothesis and must not be promoted to tracking.
The failed initialization is one-shot (maximum_initializations=1). No fresh
validation or tracking occurred. Waiting another two minutes could not restart
this run. The displayed initial-failure status was backed by backend state, but
its truncated action text failed to explain that waiting could no longer help.
Two sequential bounded offline checks of the exact saved input succeeded without changing seeds, fit thresholds or deadlines:
| Check | Seeds | Search time | Result |
|---|---|---|---|
| Physical combined run | 54/108 | 25.4795 s | incomplete / rejected |
| Isolated direct calculation | 108/108 | approximately 16.06 s | one candidate cluster |
| Production child worker, original thread limits | 108/108 | 16.0428 s | same candidate cluster |
The production worker wall time was 16.2001 s. The selected overlap/RMSE and five supporting seeds match the physical partial result. This proves a large execution slowdown in the combined physical run, not the exact source of contention. Neither isolated calculation supplies the three fresh validation windows required to authorize the product's tracking state.
Camera failure and stable reference
Compared against docs/lab/010_K1_MISSION_CORE_INTERNAL_LIVE_BASELINE_20260822.redacted.md,
docs/audits/2026-09-06-k1-live-reference-camera-root.md, ADR0007 and current
MISSIONCOR-3. The established path is acquisition-owned RTSP/TCP → FFmpeg H.264
copy/remux → durable fMP4 → disposable browser MSE reader. The per-lineage camera
fast path and verified physical-ledger snapshot cache are present in current code.
There is no evidence here that those fixes were simply removed.
In 005, a single archive epoch completed with 696 media segments / 42,239,662 valid container bytes, and no archive failure code. "Complete" describes durable container recording; it does not prove decodable H.264 content.
Browser diagnostics show 24 append/restart failures with video error code 3
(MEDIA_ERR_DECODE) and one late slow-reader retirement. Offline decoding of the
retained init plus all media segments in recorded index order confirms errors in
the archive itself: 166 corrupt decoded frame warnings, 207 macroblock decode
errors, plus 21 non-monotonic-DTS warnings. FFmpeg exited zero despite these errors;
its exit status alone must not be used as acceptance. Counts are diagnostic log
occurrences, not a claim of 166 distinct lost source frames.
The known baseline B had no H.264 corruption, zero browser restarts/slow-reader retirements and a 23 ms camera activation. Current activation took 2,377 ms, including 2,061 ms authority wait and 312 ms FFmpeg preparation; first-PCL admission was 448 ms versus baseline 65 ms. That earlier regression was caused by repeated ledger work delaying both MQTT publication and FFmpeg stdout. The present symptoms are consistent with contention, but no active-run CPU/lock/pipe profile was retained to identify the current blocking call. Source transport/encoded data corruption cannot be excluded from the existing archive alone.
The final planning-derived queue snapshot counted 446 lidar, 523 camera-frame and two pose overflow evictions. Those are derived-consumer counters, not proof of raw capture loss. Public metrics reported 1,775 preview drops and zero point decode errors. The distinction must remain explicit.
Next correction and acceptance
- Instrument bounded per-stage CPU/wall, scheduling/lock waits, raw-commit and FFmpeg drain latency. Preserve the canonical camera ownership, exact command lineage, copy/remux quality and durable recording. Identify the blocking path before changing camera retries, queue sizes or fit deadlines.
- Qualify planning with the ordinary acquisition publisher, camera/archive and UI consumption together. Use device-free retained input; synthetic load belongs on Worker006 per AGENTS, not the operator Mac. Already-corrupt 005 media is an error-handling fixture, never a clean input for claiming decoder recovery.
- Keep full entry coverage, ambiguity rejection and three disjoint fresh checks. Add the 005 frozen geometry to private regression evidence, together with wrong region/ambiguous/stale/changed-session controls. A fast isolated fit is not enough.
- Present distinct operator instructions without clipping: pending means remain stationary; confirmed means start the walk; terminal initial failure means the search has stopped and the device/recording must be ended before a new attempt. Do not introduce an unannounced automatic device START/STOP or retry.
- Only then schedule a physical stationary start with camera continuity, accepted fresh tracking, operator-visible next action and authoritative STOP/archive.
Local memory pressure was healthy for the sequential offline checks (49% reported free; swap stable at 4,816.81 MiB); no Docker VM or temporary worker was left running. Canonical8000 remained available. The unresolved physical STOP/recovery snapshot was retained; it was not overridden or converted into a successful device stop.