docs: record worker recovery acceptance and deployment boundaries
This commit is contained in:
@@ -0,0 +1,124 @@
|
||||
# Worker telemetry recovery, 2026-09-25
|
||||
|
||||
## Failure and evidence
|
||||
|
||||
Worker 006 is a remote compute host, not a device required to share the
|
||||
operator's LAN. The owner explicitly requires Tailscale access across networks.
|
||||
At the initial inspection the Worker was disconnected from its network; after
|
||||
Wi-Fi returned the existing strict-pinned Tailscale SSH profile connected.
|
||||
|
||||
Read-only audit then established:
|
||||
|
||||
- Tailscale and OpenSSH were Running/Automatic; Tailscale ForceDaemon was true.
|
||||
- Compute Docker containers were running. A failed historical StartCompute task
|
||||
did not mean the entire Worker or every container was offline.
|
||||
- Native Telegraf 1.38.4 was Stopped/Automatic, exit code zero. SCM restart actions
|
||||
and non-crash recovery were already configured.
|
||||
- Windows Application events on September 19, 21 and 23 showed Telegraf starting,
|
||||
then terminating about 15 seconds later because its old LAN MQTT destination
|
||||
was unreachable. Source inspection of the pinned upstream release confirms
|
||||
that output retry handles only errors marked retryable; merely configuring
|
||||
`startup_error_behavior=retry` did not cover the observed MQTT failure.
|
||||
- Operator Docker was stopped, with login autostart disabled. No broker/query
|
||||
listeners existed. Its saved broker bind belonged to an older DHCP lease.
|
||||
- Private telemetry configuration remained in the original installation while
|
||||
Core's network-apply handler assumed its current source checkout owned it.
|
||||
|
||||
No credentials, database volumes, Worker computation profiles or inference jobs
|
||||
were replaced. Docker Desktop restored its existing unrelated restart-policy
|
||||
containers when its engine started; those were not renamed or reconfigured.
|
||||
|
||||
## Implemented installation profile
|
||||
|
||||
Windows native Telegraf publishes to Worker `127.0.0.1:1883`. A dedicated
|
||||
Mac-owned SSH reverse forward exposes that loopback listener and carries data
|
||||
through the existing strictly pinned Tailscale SSH profile to operator
|
||||
`127.0.0.1:1883`. Broker publication is loopback-only. Query remains
|
||||
`127.0.0.1:18030`; Core remains `127.0.0.1:8000`. No `.local` name or LAN lease is
|
||||
required by this transport. MQTT credentials and per-agent ACLs are preserved.
|
||||
SSH supplies authenticated encryption; this does not claim public MQTT/mTLS
|
||||
or multi-tenant enrollment acceptance.
|
||||
|
||||
`scripts/manage_telemetry_startup.py` owns a hash-bound plan/apply/rollback:
|
||||
|
||||
- preserved prepared stack directory, credentials and named volumes;
|
||||
- explicit `MISSIONCORE_TELEMETRY_PLANE_ROOT` for the existing Core handler;
|
||||
- `com.nodedc.telemetry-startup.local`: login startup and bounded reconciliation
|
||||
every 30 seconds, including delayed Docker Desktop availability;
|
||||
- `com.nodedc.telemetry-tunnel.local`: keepalive/reconnect and loopback-only
|
||||
reverse forwarding, strict known-host verification;
|
||||
- only broker, Timescale and normalizer are selected by Compose, no build or
|
||||
image pull; existing database-bootstrap is an idempotent dependency;
|
||||
- configuration backups outside Git; apply preserves canonical Core health;
|
||||
- rollback restores declarations and credentials; it does not reset volumes or
|
||||
stop Docker/unrelated containers. The restored broker declaration is applied
|
||||
on subsequent reconciliation. Rollback rehearsal is not yet accepted.
|
||||
|
||||
`Install-NdcMissionCoreTelemetryRecovery.ps1` is included in the Windows agent
|
||||
bundle. Loopback agent install/update invokes it. It installs one SYSTEM task
|
||||
at boot and every minute, with no interactive-login requirement. Its script is
|
||||
writable only by SYSTEM/Administrators. It starts only a stopped managed
|
||||
Telegraf service, and only after the configured loopback endpoint is reachable.
|
||||
A missing connection leaves the task retryable. It never restarts a healthy
|
||||
agent or starts inference. Maintenance can disable this named task before an
|
||||
intentional extended Telegraf stop; rollback restores its predecessor.
|
||||
|
||||
## Product status and refresh
|
||||
|
||||
Receiver failure now has `telemetry-receiver-unavailable`, distinct from stale
|
||||
agent observations. Fleet contour, compute and network views show receiver
|
||||
failure as an unconfirmed Worker state, not proof that its host is off. Visible
|
||||
connectivity labels are consistently «В сети» / «Не в сети»; failure details
|
||||
explain which part of the observation path failed. The
|
||||
network view no longer paints the receiver green merely because the configured
|
||||
source is agent-mqtt.
|
||||
|
||||
A responding Core with `recording_cache=capacity-pressure` is reachable with a
|
||||
recording limitation. It is displayed as «В сети» with the limitation separately. This matters on the
|
||||
operator host, where free space fell below the existing 2 GiB recording reserve
|
||||
(about 1.8 GiB observed); no owner data was deleted.
|
||||
|
||||
The contour Refresh action now uses the canonical circular IconButton in the
|
||||
outer window header, alongside expand/close. The body copy button is removed.
|
||||
|
||||
## Measured acceptance
|
||||
|
||||
- Fresh agent-mqtt data reached the canonical Core, with expected node identity.
|
||||
- Native Windows service clean-stop test: automatic recovery in 51.7 s, new PID;
|
||||
no manual Start-Service was needed (failure restoration path was not used).
|
||||
- Dedicated tunnel SIGKILL: new tunnel process and new source observations in
|
||||
10.43 s. The API's 30 s freshness window did not expire during this short cut.
|
||||
- Normalizer stop: receiver-unavailable was observed, then automatic Compose
|
||||
recovery and new source observations in 19.58 s.
|
||||
- Tests are reproducible with `Test-NdcMissionCoreTelemetryRecovery.ps1` and
|
||||
`scripts/check_telemetry_recovery.py`; they target telemetry only.
|
||||
- 36 focused backend tests, Ruff, 984 frontend tests, TypeScript and production
|
||||
build passed. Four installer/test PowerShell artifacts were parsed by the
|
||||
native Windows parser; install/update hooks were not a clean-host rehearsal.
|
||||
- Browser acceptance on canonical port 8000 confirms fresh Worker observations,
|
||||
standardized connectivity labels and the circular Refresh control in the
|
||||
window header (including its checking state).
|
||||
|
||||
## Qualification boundary
|
||||
|
||||
The accepted live profile is macOS operator session + Windows native telemetry
|
||||
agent + the existing Tailscale trust relationship. Automatic recovery of the
|
||||
three injected failures is proved. A full physical host power cycle, distinct
|
||||
physical networks, extended network-loss soak and clean-host reinstall have
|
||||
not been run in this increment. Tailscale may choose a direct encrypted path
|
||||
when the peers happen to share a LAN; that does not reintroduce LAN addressing.
|
||||
|
||||
Mac LaunchAgents and Docker Desktop start after operator login, not before
|
||||
macOS login. Windows Tailscale/SSH/telemetry recovery use system services/tasks.
|
||||
Worker Docker computation still has its separate interactive-session lifecycle;
|
||||
this work does not claim all GPU jobs should auto-resume after power loss.
|
||||
Linux agent lifecycle and an unattended headless compute installation need a
|
||||
separate qualified OS profile. Do not label the whole system universally
|
||||
portable or cold-boot accepted from these component-level tests.
|
||||
|
||||
## Upstream references
|
||||
|
||||
- [Telegraf 1.38.4 output lifecycle](https://github.com/influxdata/telegraf/blob/v1.38.4/models/running_output.go)
|
||||
- [Telegraf 1.38.4 MQTT connect](https://github.com/influxdata/telegraf/blob/v1.38.4/plugins/outputs/mqtt/mqtt.go)
|
||||
- [Windows Tailscale unattended](https://tailscale.com/docs/how-to/run-unattended)
|
||||
- [Docker Desktop login startup](https://docs.docker.com/desktop/settings-and-maintenance/settings/)
|
||||
Reference in New Issue
Block a user