docs: record worker recovery acceptance and deployment boundaries

This commit is contained in:
DCCONSTRUCTIONS
2026-09-25 23:25:02 +03:00
parent e2cf3078b6
commit 0fb117fe7b
@@ -0,0 +1,124 @@
# Worker telemetry recovery, 2026-09-25
## Failure and evidence
Worker 006 is a remote compute host, not a device required to share the
operator's LAN. The owner explicitly requires Tailscale access across networks.
At the initial inspection the Worker was disconnected from its network; after
Wi-Fi returned the existing strict-pinned Tailscale SSH profile connected.
Read-only audit then established:
- Tailscale and OpenSSH were Running/Automatic; Tailscale ForceDaemon was true.
- Compute Docker containers were running. A failed historical StartCompute task
did not mean the entire Worker or every container was offline.
- Native Telegraf 1.38.4 was Stopped/Automatic, exit code zero. SCM restart actions
and non-crash recovery were already configured.
- Windows Application events on September 19, 21 and 23 showed Telegraf starting,
then terminating about 15 seconds later because its old LAN MQTT destination
was unreachable. Source inspection of the pinned upstream release confirms
that output retry handles only errors marked retryable; merely configuring
`startup_error_behavior=retry` did not cover the observed MQTT failure.
- Operator Docker was stopped, with login autostart disabled. No broker/query
listeners existed. Its saved broker bind belonged to an older DHCP lease.
- Private telemetry configuration remained in the original installation while
Core's network-apply handler assumed its current source checkout owned it.
No credentials, database volumes, Worker computation profiles or inference jobs
were replaced. Docker Desktop restored its existing unrelated restart-policy
containers when its engine started; those were not renamed or reconfigured.
## Implemented installation profile
Windows native Telegraf publishes to Worker `127.0.0.1:1883`. A dedicated
Mac-owned SSH reverse forward exposes that loopback listener and carries data
through the existing strictly pinned Tailscale SSH profile to operator
`127.0.0.1:1883`. Broker publication is loopback-only. Query remains
`127.0.0.1:18030`; Core remains `127.0.0.1:8000`. No `.local` name or LAN lease is
required by this transport. MQTT credentials and per-agent ACLs are preserved.
SSH supplies authenticated encryption; this does not claim public MQTT/mTLS
or multi-tenant enrollment acceptance.
`scripts/manage_telemetry_startup.py` owns a hash-bound plan/apply/rollback:
- preserved prepared stack directory, credentials and named volumes;
- explicit `MISSIONCORE_TELEMETRY_PLANE_ROOT` for the existing Core handler;
- `com.nodedc.telemetry-startup.local`: login startup and bounded reconciliation
every 30 seconds, including delayed Docker Desktop availability;
- `com.nodedc.telemetry-tunnel.local`: keepalive/reconnect and loopback-only
reverse forwarding, strict known-host verification;
- only broker, Timescale and normalizer are selected by Compose, no build or
image pull; existing database-bootstrap is an idempotent dependency;
- configuration backups outside Git; apply preserves canonical Core health;
- rollback restores declarations and credentials; it does not reset volumes or
stop Docker/unrelated containers. The restored broker declaration is applied
on subsequent reconciliation. Rollback rehearsal is not yet accepted.
`Install-NdcMissionCoreTelemetryRecovery.ps1` is included in the Windows agent
bundle. Loopback agent install/update invokes it. It installs one SYSTEM task
at boot and every minute, with no interactive-login requirement. Its script is
writable only by SYSTEM/Administrators. It starts only a stopped managed
Telegraf service, and only after the configured loopback endpoint is reachable.
A missing connection leaves the task retryable. It never restarts a healthy
agent or starts inference. Maintenance can disable this named task before an
intentional extended Telegraf stop; rollback restores its predecessor.
## Product status and refresh
Receiver failure now has `telemetry-receiver-unavailable`, distinct from stale
agent observations. Fleet contour, compute and network views show receiver
failure as an unconfirmed Worker state, not proof that its host is off. Visible
connectivity labels are consistently «В сети» / «Не в сети»; failure details
explain which part of the observation path failed. The
network view no longer paints the receiver green merely because the configured
source is agent-mqtt.
A responding Core with `recording_cache=capacity-pressure` is reachable with a
recording limitation. It is displayed as «В сети» with the limitation separately. This matters on the
operator host, where free space fell below the existing 2 GiB recording reserve
(about 1.8 GiB observed); no owner data was deleted.
The contour Refresh action now uses the canonical circular IconButton in the
outer window header, alongside expand/close. The body copy button is removed.
## Measured acceptance
- Fresh agent-mqtt data reached the canonical Core, with expected node identity.
- Native Windows service clean-stop test: automatic recovery in 51.7 s, new PID;
no manual Start-Service was needed (failure restoration path was not used).
- Dedicated tunnel SIGKILL: new tunnel process and new source observations in
10.43 s. The API's 30 s freshness window did not expire during this short cut.
- Normalizer stop: receiver-unavailable was observed, then automatic Compose
recovery and new source observations in 19.58 s.
- Tests are reproducible with `Test-NdcMissionCoreTelemetryRecovery.ps1` and
`scripts/check_telemetry_recovery.py`; they target telemetry only.
- 36 focused backend tests, Ruff, 984 frontend tests, TypeScript and production
build passed. Four installer/test PowerShell artifacts were parsed by the
native Windows parser; install/update hooks were not a clean-host rehearsal.
- Browser acceptance on canonical port 8000 confirms fresh Worker observations,
standardized connectivity labels and the circular Refresh control in the
window header (including its checking state).
## Qualification boundary
The accepted live profile is macOS operator session + Windows native telemetry
agent + the existing Tailscale trust relationship. Automatic recovery of the
three injected failures is proved. A full physical host power cycle, distinct
physical networks, extended network-loss soak and clean-host reinstall have
not been run in this increment. Tailscale may choose a direct encrypted path
when the peers happen to share a LAN; that does not reintroduce LAN addressing.
Mac LaunchAgents and Docker Desktop start after operator login, not before
macOS login. Windows Tailscale/SSH/telemetry recovery use system services/tasks.
Worker Docker computation still has its separate interactive-session lifecycle;
this work does not claim all GPU jobs should auto-resume after power loss.
Linux agent lifecycle and an unattended headless compute installation need a
separate qualified OS profile. Do not label the whole system universally
portable or cold-boot accepted from these component-level tests.
## Upstream references
- [Telegraf 1.38.4 output lifecycle](https://github.com/influxdata/telegraf/blob/v1.38.4/models/running_output.go)
- [Telegraf 1.38.4 MQTT connect](https://github.com/influxdata/telegraf/blob/v1.38.4/plugins/outputs/mqtt/mqtt.go)
- [Windows Tailscale unattended](https://tailscale.com/docs/how-to/run-unattended)
- [Docker Desktop login startup](https://docs.docker.com/desktop/settings-and-maintenance/settings/)