docs: record worker recovery acceptance and deployment boundaries
This commit is contained in:
@@ -0,0 +1,124 @@
|
|||||||
|
# Worker telemetry recovery, 2026-09-25
|
||||||
|
|
||||||
|
## Failure and evidence
|
||||||
|
|
||||||
|
Worker 006 is a remote compute host, not a device required to share the
|
||||||
|
operator's LAN. The owner explicitly requires Tailscale access across networks.
|
||||||
|
At the initial inspection the Worker was disconnected from its network; after
|
||||||
|
Wi-Fi returned the existing strict-pinned Tailscale SSH profile connected.
|
||||||
|
|
||||||
|
Read-only audit then established:
|
||||||
|
|
||||||
|
- Tailscale and OpenSSH were Running/Automatic; Tailscale ForceDaemon was true.
|
||||||
|
- Compute Docker containers were running. A failed historical StartCompute task
|
||||||
|
did not mean the entire Worker or every container was offline.
|
||||||
|
- Native Telegraf 1.38.4 was Stopped/Automatic, exit code zero. SCM restart actions
|
||||||
|
and non-crash recovery were already configured.
|
||||||
|
- Windows Application events on September 19, 21 and 23 showed Telegraf starting,
|
||||||
|
then terminating about 15 seconds later because its old LAN MQTT destination
|
||||||
|
was unreachable. Source inspection of the pinned upstream release confirms
|
||||||
|
that output retry handles only errors marked retryable; merely configuring
|
||||||
|
`startup_error_behavior=retry` did not cover the observed MQTT failure.
|
||||||
|
- Operator Docker was stopped, with login autostart disabled. No broker/query
|
||||||
|
listeners existed. Its saved broker bind belonged to an older DHCP lease.
|
||||||
|
- Private telemetry configuration remained in the original installation while
|
||||||
|
Core's network-apply handler assumed its current source checkout owned it.
|
||||||
|
|
||||||
|
No credentials, database volumes, Worker computation profiles or inference jobs
|
||||||
|
were replaced. Docker Desktop restored its existing unrelated restart-policy
|
||||||
|
containers when its engine started; those were not renamed or reconfigured.
|
||||||
|
|
||||||
|
## Implemented installation profile
|
||||||
|
|
||||||
|
Windows native Telegraf publishes to Worker `127.0.0.1:1883`. A dedicated
|
||||||
|
Mac-owned SSH reverse forward exposes that loopback listener and carries data
|
||||||
|
through the existing strictly pinned Tailscale SSH profile to operator
|
||||||
|
`127.0.0.1:1883`. Broker publication is loopback-only. Query remains
|
||||||
|
`127.0.0.1:18030`; Core remains `127.0.0.1:8000`. No `.local` name or LAN lease is
|
||||||
|
required by this transport. MQTT credentials and per-agent ACLs are preserved.
|
||||||
|
SSH supplies authenticated encryption; this does not claim public MQTT/mTLS
|
||||||
|
or multi-tenant enrollment acceptance.
|
||||||
|
|
||||||
|
`scripts/manage_telemetry_startup.py` owns a hash-bound plan/apply/rollback:
|
||||||
|
|
||||||
|
- preserved prepared stack directory, credentials and named volumes;
|
||||||
|
- explicit `MISSIONCORE_TELEMETRY_PLANE_ROOT` for the existing Core handler;
|
||||||
|
- `com.nodedc.telemetry-startup.local`: login startup and bounded reconciliation
|
||||||
|
every 30 seconds, including delayed Docker Desktop availability;
|
||||||
|
- `com.nodedc.telemetry-tunnel.local`: keepalive/reconnect and loopback-only
|
||||||
|
reverse forwarding, strict known-host verification;
|
||||||
|
- only broker, Timescale and normalizer are selected by Compose, no build or
|
||||||
|
image pull; existing database-bootstrap is an idempotent dependency;
|
||||||
|
- configuration backups outside Git; apply preserves canonical Core health;
|
||||||
|
- rollback restores declarations and credentials; it does not reset volumes or
|
||||||
|
stop Docker/unrelated containers. The restored broker declaration is applied
|
||||||
|
on subsequent reconciliation. Rollback rehearsal is not yet accepted.
|
||||||
|
|
||||||
|
`Install-NdcMissionCoreTelemetryRecovery.ps1` is included in the Windows agent
|
||||||
|
bundle. Loopback agent install/update invokes it. It installs one SYSTEM task
|
||||||
|
at boot and every minute, with no interactive-login requirement. Its script is
|
||||||
|
writable only by SYSTEM/Administrators. It starts only a stopped managed
|
||||||
|
Telegraf service, and only after the configured loopback endpoint is reachable.
|
||||||
|
A missing connection leaves the task retryable. It never restarts a healthy
|
||||||
|
agent or starts inference. Maintenance can disable this named task before an
|
||||||
|
intentional extended Telegraf stop; rollback restores its predecessor.
|
||||||
|
|
||||||
|
## Product status and refresh
|
||||||
|
|
||||||
|
Receiver failure now has `telemetry-receiver-unavailable`, distinct from stale
|
||||||
|
agent observations. Fleet contour, compute and network views show receiver
|
||||||
|
failure as an unconfirmed Worker state, not proof that its host is off. Visible
|
||||||
|
connectivity labels are consistently «В сети» / «Не в сети»; failure details
|
||||||
|
explain which part of the observation path failed. The
|
||||||
|
network view no longer paints the receiver green merely because the configured
|
||||||
|
source is agent-mqtt.
|
||||||
|
|
||||||
|
A responding Core with `recording_cache=capacity-pressure` is reachable with a
|
||||||
|
recording limitation. It is displayed as «В сети» with the limitation separately. This matters on the
|
||||||
|
operator host, where free space fell below the existing 2 GiB recording reserve
|
||||||
|
(about 1.8 GiB observed); no owner data was deleted.
|
||||||
|
|
||||||
|
The contour Refresh action now uses the canonical circular IconButton in the
|
||||||
|
outer window header, alongside expand/close. The body copy button is removed.
|
||||||
|
|
||||||
|
## Measured acceptance
|
||||||
|
|
||||||
|
- Fresh agent-mqtt data reached the canonical Core, with expected node identity.
|
||||||
|
- Native Windows service clean-stop test: automatic recovery in 51.7 s, new PID;
|
||||||
|
no manual Start-Service was needed (failure restoration path was not used).
|
||||||
|
- Dedicated tunnel SIGKILL: new tunnel process and new source observations in
|
||||||
|
10.43 s. The API's 30 s freshness window did not expire during this short cut.
|
||||||
|
- Normalizer stop: receiver-unavailable was observed, then automatic Compose
|
||||||
|
recovery and new source observations in 19.58 s.
|
||||||
|
- Tests are reproducible with `Test-NdcMissionCoreTelemetryRecovery.ps1` and
|
||||||
|
`scripts/check_telemetry_recovery.py`; they target telemetry only.
|
||||||
|
- 36 focused backend tests, Ruff, 984 frontend tests, TypeScript and production
|
||||||
|
build passed. Four installer/test PowerShell artifacts were parsed by the
|
||||||
|
native Windows parser; install/update hooks were not a clean-host rehearsal.
|
||||||
|
- Browser acceptance on canonical port 8000 confirms fresh Worker observations,
|
||||||
|
standardized connectivity labels and the circular Refresh control in the
|
||||||
|
window header (including its checking state).
|
||||||
|
|
||||||
|
## Qualification boundary
|
||||||
|
|
||||||
|
The accepted live profile is macOS operator session + Windows native telemetry
|
||||||
|
agent + the existing Tailscale trust relationship. Automatic recovery of the
|
||||||
|
three injected failures is proved. A full physical host power cycle, distinct
|
||||||
|
physical networks, extended network-loss soak and clean-host reinstall have
|
||||||
|
not been run in this increment. Tailscale may choose a direct encrypted path
|
||||||
|
when the peers happen to share a LAN; that does not reintroduce LAN addressing.
|
||||||
|
|
||||||
|
Mac LaunchAgents and Docker Desktop start after operator login, not before
|
||||||
|
macOS login. Windows Tailscale/SSH/telemetry recovery use system services/tasks.
|
||||||
|
Worker Docker computation still has its separate interactive-session lifecycle;
|
||||||
|
this work does not claim all GPU jobs should auto-resume after power loss.
|
||||||
|
Linux agent lifecycle and an unattended headless compute installation need a
|
||||||
|
separate qualified OS profile. Do not label the whole system universally
|
||||||
|
portable or cold-boot accepted from these component-level tests.
|
||||||
|
|
||||||
|
## Upstream references
|
||||||
|
|
||||||
|
- [Telegraf 1.38.4 output lifecycle](https://github.com/influxdata/telegraf/blob/v1.38.4/models/running_output.go)
|
||||||
|
- [Telegraf 1.38.4 MQTT connect](https://github.com/influxdata/telegraf/blob/v1.38.4/plugins/outputs/mqtt/mqtt.go)
|
||||||
|
- [Windows Tailscale unattended](https://tailscale.com/docs/how-to/run-unattended)
|
||||||
|
- [Docker Desktop login startup](https://docs.docker.com/desktop/settings-and-maintenance/settings/)
|
||||||
Reference in New Issue
Block a user