NODEDC_PLATFORM/services/dc-amd-proxy/OPERATIONS.md

36 KiB
Raw Permalink Blame History

DC AMD Proxy — production operations and recovery canon

This document is the reproducible operations contract for the restricted Cesium/Bing egress path through the neighbouring Windows AMD workstation. It describes the implementation that is in source now. It is not a proposal for turning the NAS or the workstation into a general-purpose proxy.

The two non-negotiable boundaries are:

  1. the NAS keeps its own routes, DNS, VPN, Tailscale and Internet connectivity unchanged;
  2. only approved Cesium/Bing HTTPS traffic may leave through the AMD host.

Do not solve an incident by adding a system proxy, default route, transparent NAT, browser-visible token, permissive CONNECT endpoint, 0.0.0.0 Windows bind, or direct NAS provider fallback.

1. Canonical architecture

Browser
  -> Foundry BFF
  -> Platform Map Gateway
       |-- warm object: NAS live/offline cache -> browser
       `-- cold/refresh object:
             HTTP + x-proxy-token
             -> NAS  <NAS_LAN_IP>:<NAS_PROXY_PORT>  dc-amd-proxy
                  -> authenticated HTTP CONNECT
                  -> AMD <AMD_LAN_IP>:<CONNECTOR_PORT> dc-amd-connector
                       -> AMD/Docker DNS
                       -> Windows VPN/default egress route
                       -> approved Cesium/Bing host:443
                  <- end-to-end provider TLS is terminated by dc-amd-proxy
             <- streamed provider response
       -> browser and, when enabled, atomic NAS cache fill

The AMD connector is a byte-forwarding CONNECT boundary. It does not parse an Ion bearer and cannot read the provider TLS stream. dc-amd-proxy establishes TLS through that tunnel, validates the provider certificate and SNI, follows only approved redirects, and returns the response to Map Gateway.

This distinction is operationally important:

  • a warm cache hit does not traverse AMD, VPN or Cesium;
  • a cold miss or refresh does traverse AMD and depends on the workstation's Docker Desktop, DNS, routing and VPN;
  • objects already cached on NAS must remain usable when AMD or its VPN is unavailable;
  • dc-amd-proxy is not in the browser-to-NAS cache read path.

Current components and ownership

Component Runtime Canonical state Responsibility
Map Gateway Synology Docker /volume1/docker/nodedc-platform/platform/services/map-gateway Provider credential scope, cache-first reads, streaming cache fill and private egress calls
dc-amd-proxy Synology Docker /volume1/docker/dc-amd-proxy Request authorization, target allowlist, CONNECT/TLS pool, redirect/retry policy and telemetry
Proxy pair state Synology bind mount /volume1/docker/dc-amd-proxy/runtime/connector-access Paired connector credential; persistent across container recreation
Map egress credential Synology root-owned file /volume1/docker/nodedc-platform/secrets/map-egress-proxy-token Authenticates Map Gateway to dc-amd-proxy
dc-amd-connector AMD Windows Docker Desktop C:\NODEDC\dc-amd-connector by default Restricted authenticated CONNECT to approved hosts on port 443
Connector credential AMD local file C:\NODEDC\dc-amd-connector\runtime\connector-access Generated locally once; never printed or stored in source
Docker Desktop launcher Current Windows user's Startup folder NODE.DC DC AMD Connector - Docker Desktop.cmd Starts Docker Desktop after that user signs in

Current source defaults are NAS listener port 8790, AMD connector port 8791, eight active sockets and four idle sockets per origin. IP addresses are deployment values and must be treated as configuration, not universal constants.

2. Trust, secrets and traffic boundaries

There are three different credentials. They are not interchangeable.

  1. The Cesium Ion credential belongs to Map Gateway provider configuration. It must not be placed in Foundry browser code, an artifact or this proxy's Compose file.
  2. The Map egress credential authenticates Map Gateway to dc-amd-proxy using x-proxy-token. The canonical runner synchronizes it into the root-owned file mounted read-only into the two services.
  3. The connector credential authenticates dc-amd-proxy to the Windows connector using Proxy-Authorization: Bearer .... The Windows installer generates it locally; the one-time pair request stores the same value in the NAS runtime bind mount.

No command in this runbook prints, copies to chat, or accepts a secret as a command-line argument.

Current runner dependency that must not be mistaken for a traffic route

The deploy runner currently seeds the shared Map egress credential from the root-owned PROXY_TOKEN in /volume1/docker/proxy-contur/.env. Therefore a valid, safely permissioned historical proxy-contur environment file is a deploy-time precondition for dc-amd-proxy and Map Gateway applies.

This does not mean Cesium traffic passes through proxy-contur. Runtime traffic goes directly from Map Gateway to dc-amd-proxy, then to the AMD connector. Removing this legacy secret-seeding dependency is an explicit evolution item later in this document.

NAS network boundary

dc-amd-proxy uses network_mode: host so the one-time pair endpoint can see the real AMD source address. Synology's published-port bridge path would hide that source behind Docker NAT. Host mode does not make the service broad:

  • the process binds only <NAS_LAN_IP>:<NAS_PROXY_PORT>;
  • /api/pair accepts only <AMD_LAN_IP> and only while unpaired;
  • /proxy/cesium/fetch accepts only GET/HEAD, requires the private Map egress credential and validates the complete target URL;
  • /healthz and /status expose safe state/counters but no credentials or target URLs;
  • the container is read-only, drops all capabilities, enables no-new-privileges, and has only a small no-exec tmpfs plus the two explicit bind mounts.

The current Map Gateway URL is a NAS-LAN address, not a public hostname and not a Docker service discovery name.

Windows network boundary

The connector publishes only <AMD_LAN_IP>:<CONNECTOR_PORT>. The installer creates one inbound Windows Firewall rule with all of these restrictions:

  • TCP only;
  • local address <AMD_LAN_IP>;
  • local port <CONNECTOR_PORT>;
  • remote address <NAS_LAN_IP>;
  • no bind to the VPN adapter and no bind to all interfaces.

The connector accepts only authenticated CONNECT <allowed-host>:443. Ordinary forward-proxy HTTP methods, other ports, IP literals and arbitrary hosts are rejected. Its container is read-only, drops capabilities and is bounded to one CPU, 256 MB memory, 64 processes and 4096 file descriptors. These are circuit breakers for a shared production workstation, not bandwidth throttles.

The installer does not change the Windows default route, DNS, system proxy, VPN configuration or unrelated firewall rules. The connector's DNS lookup and TCP connection occur inside Docker Desktop on the AMD host and therefore use the egress path exposed to Docker by Windows/VPN.

Exact provider allowlist

Both NAS proxy and Windows connector currently allow only:

  • api.cesium.com
  • assets.ion.cesium.com
  • dev.virtualearth.net
  • ecn.t0.tiles.virtualearth.net
  • ecn.t1.tiles.virtualearth.net
  • ecn.t2.tiles.virtualearth.net
  • ecn.t3.tiles.virtualearth.net

Adding a provider hostname requires one reviewed change to both allowlists, updated smoke coverage, a new Windows connector package, a new NAS app-overlay artifact and the normal plan/apply acceptance. Never add a wildcard such as *.cesium.com or *.virtualearth.net.

3. Request lifecycle

For a cache miss, the sequence is:

  1. Map Gateway checks the NAS cache before acquiring provider credentials or contacting egress.
  2. Map Gateway calls http://<NAS_LAN_IP>:<NAS_PROXY_PORT>/proxy/cesium/fetch?url=<encoded-url> and authenticates with the private Map egress credential.
  3. dc-amd-proxy validates method, URL scheme, host, port and credentials.
  4. A per-origin HTTP/1.1 keep-alive agent either reuses an existing TLS tunnel, opens a new one, or queues behind the bounded active-socket limit.
  5. For a new connection, the NAS opens an authenticated CONNECT request to <AMD_LAN_IP>:<CONNECTOR_PORT>.
  6. The connector validates its credential, exact hostname and port 443, then resolves/connects through the AMD network/VPN.
  7. The NAS completes TLS over the byte tunnel and validates the provider certificate.
  8. Response headers and body stream back to Map Gateway. Client cancellation propagates through queue, CONNECT, TLS, header wait, redirect drain, retry and response streaming.
  9. Map Gateway can stream the first bytes to the caller while its independent branch completes an atomic cache fill. A failed/partial fill must not become a valid cache object.

Provider headers forwarded upstream are deliberately narrow: Accept, Range, If-None-Match, If-Modified-Since, scoped provider authorization and the Map referer. Accept-Encoding is set to identity. Hop-by-hop headers are not forwarded back downstream.

Credential scope across redirects

  • same-origin redirects retain the provider Authorization header;
  • cross-origin redirects remove Authorization before the next request;
  • every redirect target is parsed again through the exact HTTPS/443/host allowlist;
  • redirect bodies are drained with abort and idle-timeout protection;
  • the chain is bounded and terminates with amd_upstream_redirect_limit rather than looping indefinitely.

This is why the Ion API bearer cannot leak when the API redirects to the assets or Bing origin.

Pool, retry and timeout policy

Control Current default Meaning
Active sockets 8 per origin Maximum concurrent provider tunnels for each approved origin
Idle sockets 4 per origin Reusable warm TLS tunnels retained for each origin
TCP keep-alive initial delay 30 seconds Socket keep-alive probe setting; actual idle lifetime is also controlled by provider/server socket closure
Connector/TLS timeout 20 seconds Bounds NAS-to-AMD CONNECT and provider TLS setup
Upstream header timeout 30 seconds Bounds one attempt until provider response headers
Response/redirect body idle timeout 30 seconds Resets on body progress; does not cap a healthy long stream
Slow request threshold 2 seconds Emits safe slow-request telemetry; does not cancel the request
Windows connector upstream connect timeout 20 seconds Bounds AMD-to-provider TCP connect
Windows connector tunnel idle timeout 90 seconds Closes a tunnel with no traffic

The NAS duration values accept 1120 seconds and pool sizes accept 132. Do not increase them to hide a route/VPN problem: first interpret queue, connection, TTFB and body timings separately.

There is at most one transport retry. It is allowed only when all conditions are true:

  • method is idempotent GET or HEAD;
  • the failed attempt used a previously reused keep-alive socket;
  • failure is ECONNRESET, EPIPE, ETIMEDOUT or ECONNABORTED;
  • the request has not already retried and the client has not aborted.

A failure on a newly opened socket is not blindly replayed. A retry may open a replacement socket after the stale reused socket is discarded. Both attempts are included in telemetry and terminal accounting.

Runtime configuration reference

All changes to these values require a reviewed Compose/source artifact; do not inject an ad-hoc live container environment.

NAS proxy variable Default Contract
PORT 8790 Internal/listener port
DC_AMD_PROXY_BIND_ADDRESS deployment NAS LAN IP Exact listener address; IPv4 only
DC_AMD_CONNECTOR_HOST deployment AMD LAN IP Stable connector LAN address; IPv4 only
DC_AMD_CONNECTOR_PORT 8791 Windows published connector port
DC_AMD_PAIR_ALLOWED_SOURCE deployment AMD LAN IP Only source allowed to perform first pair
DC_AMD_CONNECTOR_TOKEN_FILE /var/lib/dc-amd-proxy/connector-access Persistent paired connector state
DC_AMD_MAP_EGRESS_TOKEN_FILE /run/nodedc-secrets/map-egress-proxy-token Read-only Map Gateway authentication file
DC_AMD_CONNECT_TIMEOUT_SECONDS 20 CONNECT and TLS setup timeout
DC_AMD_UPSTREAM_TIMEOUT_SECONDS 30 Per-attempt response-header timeout
DC_AMD_BODY_IDLE_TIMEOUT_SECONDS 30 Progress-based response/redirect body timeout
DC_AMD_POOL_MAX_SOCKETS 8 Active sockets per origin
DC_AMD_POOL_MAX_FREE_SOCKETS 4 Idle sockets per origin
DC_AMD_SLOW_REQUEST_SECONDS 2 Safe slow-event threshold

The Windows Compose .env contains only AMD_CONNECTOR_BIND_IP and AMD_CONNECTOR_PORT. The connector server also supports bounded AMD_CONNECTOR_CONNECT_TIMEOUT_SECONDS (default 20) and AMD_CONNECTOR_IDLE_TIMEOUT_SECONDS (default 90); changing them should be a versioned connector-package change.

4. Reproducible initial installation

Replace every angle-bracket placeholder before executing a command. Do not paste a credential into any placeholder.

4.1 Preconditions

  • AMD workstation has a stable LAN address reserved in DHCP or statically configured.
  • NAS and AMD can reach each other on the trusted LAN.
  • Docker Desktop is installed and its Linux engine is running.
  • The intended Windows user can sign in after a reboot.
  • The VPN is connected and exposes its route to Docker Desktop workloads.
  • The reviewed Windows connector package contains a valid SHA256SUMS.
  • The canonical NAS runner is installed at /usr/local/sbin/nodedc-deploy.
  • No other host is using the selected connector LAN address/port.

4.2 Install the Windows connector

Extract the reviewed connector package to a temporary local folder. Open PowerShell as Administrator, enter that folder and run:

powershell.exe -NoProfile -ExecutionPolicy Bypass -File ".\install.ps1" `
  -Destination "C:\NODEDC\dc-amd-connector" `
  -NasAddress "<NAS_LAN_IP>" `
  -BindAddress "<AMD_LAN_IP>" `
  -Port <CONNECTOR_PORT> `
  -EnableAutoStart $true

The installer is fail-closed. It verifies SHA256SUMS, refuses to overwrite an existing destination or firewall rule, generates a random connector credential without printing it, builds/starts the container, waits for healthy, creates the NAS-only firewall rule and writes the current-user Docker Desktop Startup launcher. On failure it removes only the state it created.

Verify locally without displaying the credential:

Set-Location "C:\NODEDC\dc-amd-connector"
docker compose ps
docker inspect dc-amd-connector --format '{{.State.Status}} health={{if .State.Health}}{{.State.Health.Status}}{{end}} restart={{.HostConfig.RestartPolicy.Name}}'
docker compose logs --tail 100

Expected state is running health=healthy and restart policy unless-stopped.

This safe CONNECT probe reads the credential only inside the container and prints only the HTTP status line:

docker exec dc-amd-connector node -e "const fs=require('fs'),net=require('net');const t=fs.readFileSync('/run/dc-amd-secrets/connector-access','utf8').trim();const s=net.connect(8791,'127.0.0.1',()=>s.write('CONNECT api.cesium.com:443 HTTP/1.1\r\nHost: api.cesium.com:443\r\nProxy-Authorization: Bearer '+t+'\r\n\r\n'));s.once('data',d=>{const v=d.toString('ascii');console.log(v.split('\r\n')[0]);s.destroy();process.exit(v.startsWith('HTTP/1.1 200')?0:1)});s.on('error',e=>{console.error(e.message);process.exit(1)})"

Expected output is HTTP/1.1 200 Connection Established. This proves that the connector can currently reach the API host; it does not validate an Ion token or prove the VPN exit country.

4.3 Build and transfer the NAS artifact

From a clean platform repository checkout on the trusted build Mac:

node infra/deploy-runner/build-dc-amd-proxy-artifact.mjs \
  dc-amd-proxy-<CHANGE_SLUG>-<YYYYMMDD>-<NNN>

cd infra/deploy-artifacts
shasum -a 256 nodedc-dc-amd-proxy-<CHANGE_SLUG>-<YYYYMMDD>-<NNN>.tgz \
  > nodedc-dc-amd-proxy-<CHANGE_SLUG>-<YYYYMMDD>-<NNN>.tgz.sha256

Copy exactly the .tgz and .tgz.sha256 through SMB into:

/volume1/docker/nodedc-deploy/inbox

The builder includes only Dockerfile, README.md, docker-compose.yml, package.json and server.mjs. Runtime, .env, credentials, logs, Windows packages and this operations document are not deploy payload members.

4.4 Canonical NAS plan/apply

On Synology, from the canonical inbox:

cd /volume1/docker/nodedc-deploy/inbox

sha256sum -c \
  nodedc-dc-amd-proxy-<CHANGE_SLUG>-<YYYYMMDD>-<NNN>.tgz.sha256

sudo /usr/local/sbin/nodedc-deploy verify-install

sudo /usr/local/sbin/nodedc-deploy plan \
  /volume1/docker/nodedc-deploy/inbox/nodedc-dc-amd-proxy-<CHANGE_SLUG>-<YYYYMMDD>-<NNN>.tgz

Accept the plan only when all are true:

  • component=dc-amd-proxy
  • type=app-overlay
  • payload_root=/volume1/docker/dc-amd-proxy
  • services=dc-amd-proxy
  • runtime_secret=runner-synced:.../map-egress-proxy-token
  • runtime_state=runner-prepared:/volume1/docker/dc-amd-proxy/runtime
  • state=new
  • the file list is exactly the five allowlisted source files.

Then apply the exact planned filename; never use “latest” discovery:

sudo /usr/local/sbin/nodedc-deploy apply \
  /volume1/docker/nodedc-deploy/inbox/nodedc-dc-amd-proxy-<CHANGE_SLUG>-<YYYYMMDD>-<NNN>.tgz

The runner preserves/prepares the private pair-state directory, synchronizes the Map egress credential, creates a source backup, builds with Compose, force-recreates only dc-amd-proxy, waits for container health, moves the artifact to applied, and records the result. A new service should report state=awaiting_pair.

4.5 Pair Windows to NAS

Pair only after NAS status says awaiting_pair. Place the reviewed pairing tool on the AMD machine and run it from an elevated PowerShell window:

powershell.exe -NoProfile -ExecutionPolicy Bypass `
  -File ".\pair-dc-amd-proxy.ps1" `
  -NasAddress "<NAS_LAN_IP>" `
  -Port <NAS_PROXY_PORT>

Expected output:

AMD connector paired with NAS. No secret value was displayed.

The pair call is a one-time HTTP LAN handoff with the Windows proxy explicitly disabled. Its security boundary is the trusted LAN plus exact AMD source IP. After the first successful pair, later attempts return 409 already_paired.

Verify on NAS:

curl -fsS http://<NAS_LAN_IP>:<NAS_PROXY_PORT>/status

sudo /usr/local/bin/docker inspect dc-amd-proxy \
  --format 'status={{.State.Status}} health={{if .State.Health}}{{.State.Health.Status}}{{end}} network={{.HostConfig.NetworkMode}} read_only={{.HostConfig.ReadonlyRootfs}} restart={{.HostConfig.RestartPolicy.Name}}'

Expected properties are state=paired, forwarding=amd_connector_only, directEgress=false, container healthy, network host, read-only true, and restart unless-stopped.

Finally verify one real Map Gateway request using the product health workflow. Do not construct a shell command that reads or prints either private runtime credential.

For a first installation, switch Map Gateway to http://<NAS_LAN_IP>:<NAS_PROXY_PORT> only through a separate reviewed Platform artifact and only after the proxy is paired and its tunnel test passes. This separation keeps a failed connector bootstrap from breaking an already-working provider route.

5. Restart and reboot behavior

Synology

The NAS container has restart: unless-stopped. A Docker daemon or NAS reboot restarts it. Pair state survives because ./runtime is a bind mount outside the container. The Map egress credential is a separate root-owned bind mount.

If an operator intentionally runs docker compose down, the container is removed and restart policy cannot recreate it. Restore it from the canonical source directory:

cd /volume1/docker/dc-amd-proxy
sudo /usr/local/bin/docker compose -p dc-amd-proxy up -d --build --no-deps dc-amd-proxy

Use that command only for recovery of the already-deployed version. Source changes still require a runner artifact.

Windows AMD host

Docker Desktop on the WSL 2 backend is not guaranteed to be an unattended Windows boot service. The installed launcher starts it when the configured Windows user signs in. After Docker Engine starts, restart: unless-stopped restores the existing connector container.

Therefore the actual recovery sequence is:

  1. Windows boots;
  2. the configured user signs in;
  3. Startup launches Docker Desktop;
  4. Docker Engine becomes ready;
  5. the existing connector container restarts;
  6. the VPN must be connected and usable from Docker;
  7. cold Cesium requests recover.

Before user sign-in, or while Docker/VPN is unavailable, live misses fail and Map Gateway must rely on existing NAS cache. If the connector container was removed with docker compose down, restart policy alone cannot recreate it; run from elevated PowerShell:

Set-Location "C:\NODEDC\dc-amd-connector"
docker compose up -d --build

6. Routine operations and observability

NAS commands

curl -fsS http://<NAS_LAN_IP>:<NAS_PROXY_PORT>/status

sudo /usr/local/bin/docker inspect dc-amd-proxy \
  --format 'status={{.State.Status}} health={{if .State.Health}}{{.State.Health.Status}}{{end}} restart={{.HostConfig.RestartPolicy.Name}}'

sudo /usr/local/bin/docker logs --tail 200 dc-amd-proxy

cd /volume1/docker/dc-amd-proxy
sudo /usr/local/bin/docker compose -p dc-amd-proxy ps

Windows commands

Set-Location "C:\NODEDC\dc-amd-connector"
docker compose ps
docker compose logs --tail 200
docker inspect dc-amd-connector --format '{{.State.Status}} health={{if .State.Health}}{{.State.Health.Status}}{{end}} restart={{.HostConfig.RestartPolicy.Name}}'

What /status means

/healthz and /status currently return the same safe operational body. They prove process/local state, not end-to-end provider reachability.

Important fields:

  • state: awaiting_pair or paired;
  • forwarding: disabled until paired, then amd_connector_only;
  • directEgress: must always be false;
  • pool.activeSockets, idleSockets, queuedRequests, maxQueuedRequests;
  • pool.byOrigin: safe active/idle/queued counts per allowlisted origin;
  • metrics.requests, terminalRequests, inFlightRequests;
  • completedResponses, failedRequests, clientAborts, streamFailures;
  • openedTunnels, tunnelFailures, reusedSocketRequests, retries;
  • complete and partial byte counts;
  • average/max queue, connection, TTFB, attempt, retry and total duration;
  • per-host counters and the last short transport error.

At all times:

requests = terminalRequests + inFlightRequests

At quiescence, inFlightRequests should return to zero. A violation indicates a lifecycle/accounting defect. Counters are process-local and reset after a container restart; pair state does not reset.

Interpret latency by stage:

Symptom Likely stage
High queuedRequests / queueMs Per-origin eight-socket pool saturated or requests not terminating
High connectionMs / tunnel failures NAS-to-AMD LAN, Windows connector, DNS, CONNECT or TLS setup
Low connection but high ttfbMs VPN/provider latency or provider throttling
Low TTFB but high total duration Slow/stalled response body or downstream backpressure
Rising partialBytes / streamFailures Provider/VPN body died after headers
Rising clientAborts Browser/navigation cancellation or caller timeout; not automatically an upstream failure
Low reusedSocketRequests with high openedTunnels Keep-alive churn, route instability or origin fan-out
Retry spikes Reused sockets are being reset; one bounded recovery is working but transport is unstable

JSON logs intentionally contain event names, approved hostname, short error code and numeric timing only. They must not include a full URL, query string, Ion bearer, connector credential or Map egress credential. Normal fast tile requests are not logged individually. Useful events include:

  • cesium_egress_tunnel_opened
  • cesium_egress_reused_socket_retry
  • cesium_egress_retry_completed
  • cesium_egress_retry_failed
  • cesium_egress_slow_or_failed
  • cesium_egress_stream_failed
  • cesium_egress_request_failed
  • amd_connector_paired

A warm Map Gateway cache hit should not increase dc-amd-proxy request counters. If it does, audit the cache-first ordering before tuning the proxy.

7. Failure matrix

Observable result Probable cause Safe action
NAS status unavailable Container stopped/unhealthy or wrong NAS bind address Inspect container/Compose state and logs; do not change NAS routes
state=awaiting_pair No durable connector pair state Verify intended AMD host, then run the one-time pairing tool
Pair HTTP 403 pair_source_not_allowed Request did not originate from configured AMD LAN IP Verify stable AMD IP, NAS host-mode listener and no proxy/NAT in the pair call
Pair HTTP 409 already_paired Pair state already exists Do not overwrite it; use the controlled migration/rotation procedure
map_egress_unauthorized Map Gateway and NAS proxy do not share the runner-synced Map egress credential Re-run canonical preflight/deploy investigation; never copy a token into .env by hand
amd_connector_not_paired NAS runtime pair file absent/invalid Pair from the allowed AMD source; inspect ownership only through canonical runner checks
Connector refused/timeout Windows off, no user sign-in, Docker stopped, firewall/IP mismatch or connector unhealthy Restore Windows/Docker/connector; warm cache should continue
amd_connector_rejected Connector credential mismatch or CONNECT rejected Stop; do not retry pairing blindly. Treat as pair-state/migration incident
cesium_target_not_allowed Provider redirected/requested an unreviewed host or wrong scheme/port Capture hostname safely, review necessity, update both allowlists through source
TLS timeout/certificate error VPN path, DNS, time, CA or provider interception problem Verify AMD route/VPN and NAS container time/CA; never disable TLS verification in production
amd_upstream_timeout No provider headers within attempt timeout Compare connection vs TTFB metrics; verify VPN/provider rather than increasing timeout first
Body/redirect idle timeout Stream stopped making progress Check VPN packet loss/provider; partial object must not enter Map cache
Provider HTTP 401/403 Ion token, asset permission, referer restriction or provider account issue Diagnose Map Gateway/provider configuration; transport is functioning if status is passed through
Provider HTTP 429 Provider rate/throughput policy Reduce cold fan-out and use cache; do not add unbounded sockets/retries
Rendering survives but cold areas fail Expected cache-first degradation while AMD/VPN is down Restore egress; preserve existing cache and avoid destructive cache resets
Queue grows and never drains Saturation, hung lifecycle or insufficient abort propagation Inspect in-flight/accounting invariant, body idle errors and client aborts before changing pool size
AMD host rebooted and does not recover No Windows sign-in, Startup launcher missing, Docker not ready, container removed or VPN disconnected Follow the explicit Windows reboot sequence

Important VPN caveat

The connector trusts the route Docker Desktop receives from Windows. It does not cryptographically prove that a request used a particular VPN or country. If Windows allows direct Internet egress when the VPN is off, the connector may use that direct AMD-host route. It will still never fall back to direct NAS egress, but that is a different guarantee.

If VPN-only provider egress is mandatory, enforce it on the AMD host/VPN with an outbound kill switch scoped to the connector workload/approved destinations and test it with VPN on and off. The current installer creates only an inbound NAS-only firewall rule; it does not install this outbound policy.

8. Migration to another AMD machine

Migration has two independent identities: stable LAN endpoint and connector credential. A clean install on another machine creates a new credential, while the current NAS pair endpoint is intentionally write-once. Changing only the IP in Compose is therefore insufficient.

Required zero-guesswork migration sequence

  1. Keep the old connector active while preparing the replacement.
  2. Assign the replacement a reserved <NEW_AMD_LAN_IP> and verify VPN/Docker routing.
  3. Install the reviewed connector package on the replacement using <NEW_AMD_LAN_IP> and the same NAS address/connector port.
  4. Verify local connector health and the safe CONNECT probe.
  5. Create a reviewed source change that updates NAS connector host and allowed pair source to <NEW_AMD_LAN_IP>.
  6. Add a runner-supported, audited pair-rotation transition that can retire the old NAS pair state and admit exactly one new pair. The transition must back up state, never print either credential, fail closed and be idempotent.
  7. Build, checksum, SMB-transfer, plan, and apply the exact versioned artifact/runner change through the deploy canon.
  8. Pair from the replacement machine, verify an end-to-end cold request and confirm directEgress=false plus clean accounting.
  9. Disable the old connector, then uninstall it only after acceptance.

The current implementation does not yet provide step 6 as a first-class runner operation. Until it exists, a fresh-secret migration is not fully canonical. Do not work around this by deleting /volume1/docker/dc-amd-proxy/runtime/connector-access, copying the credential through chat/SMB, weakening /api/pair, or running two active pair sources.

If an urgent migration is required before controlled rotation exists, open an Ops change and implement/review that transition first. This limitation is preferable to an undocumented secret reset.

To remove a retired connector from Windows after the new path is accepted:

powershell.exe -NoProfile -ExecutionPolicy Bypass `
  -File "C:\NODEDC\dc-amd-connector\uninstall.ps1" `
  -Destination "C:\NODEDC\dc-amd-connector"

The uninstaller removes only this container, its NAS-only firewall rule, Startup launcher and installation directory. It does not remove Docker Desktop or unrelated containers.

9. Deploy and rollback canon

Source/change acceptance before packaging

Run from the service directory in a clean checkout:

node --check server.mjs
npm run test:connection-pool
docker build --no-cache -t nodedc/dc-amd-proxy:<VERIFY_TAG> .
git status --short

The smoke suite must cover pool reuse/saturation, aborts in each lifecycle stage, exact terminal accounting, one safe retry, body idle timeout, redirect credential boundaries and absence of secret markers in logs/status.

Commit source and documentation together. Build an artifact from the commit, not from an unknown dirty tree. The normal deployment sequence is:

source review -> tests -> commit -> artifact -> checksum -> SMB inbox
-> verify-install -> exact plan -> human plan review -> exact apply
-> container health -> proxy status -> end-to-end Map acceptance

Never put .env, runtime, tokens, logs, backup archives, Windows package state or shell hooks inside a dc-amd-proxy app-overlay.

What the runner records

Before applying, the runner creates:

/volume1/docker/nodedc-deploy/backups/<BACKUP_ID>/

with manifest, file lists and source-before.tgz. Successful artifacts move to /volume1/docker/nodedc-deploy/applied; failed artifacts move to /volume1/docker/nodedc-deploy/failed. Applied and failed JSONL state lives under /volume1/docker/nodedc-deploy/state.

For this component the current runner does not perform generic automatic source rollback after a failed post-copy build/health check. The backup is evidence/recovery input, not permission for an improvised live extraction.

Canonical forward rollback

Use a new patch ID and package the known-good service source from a separate clean worktree. Do not reset or overwrite the active developer worktree:

git worktree add "../NODEDC-rollback" <KNOWN_GOOD_COMMIT>
cd "../NODEDC-rollback"

node infra/deploy-runner/build-dc-amd-proxy-artifact.mjs \
  dc-amd-proxy-rollback-<YYYYMMDD>-<NNN>

cd infra/deploy-artifacts
shasum -a 256 nodedc-dc-amd-proxy-rollback-<YYYYMMDD>-<NNN>.tgz \
  > nodedc-dc-amd-proxy-rollback-<YYYYMMDD>-<NNN>.tgz.sha256

Transfer the exact pair to the inbox and run the same checksum, verify-install, exact plan, review and exact apply sequence. The runtime pair-state directory remains outside the source overlay and must be preserved.

If the failed version changed a runtime contract, Compose contract or secret format, stop and design an explicit rollback transition instead of assuming a source-only forward rollback is safe.

10. Capacity and product behavior

The default eight active sockets are per approved origin, not global. A Cesium session can use more than one origin, but every origin remains bounded. Twenty users do not overwrite one another's transport state; they share the pool and may queue. Cache objects are owned by Map Gateway, not by this proxy.

The proxy has no tile index, no cache eviction and no “do not overwrite cache” switch. Those belong to Map Gateway. It streams bytes and records transport telemetry. Consequently:

  • proxy/VPN speed affects only cold/refresh acquisition;
  • NAS cache read speed should be independent from AMD/VPN speed;
  • increasing pool size cannot fix slow warm cache reads;
  • cache fill concurrency and capacity policy must be tuned in Map Gateway;
  • a browser reload should reuse NAS cache even if local browser cache is empty.

11. Known limitations and evolution path

The current path is production-usable but not yet a general infrastructure egress product. Track these changes explicitly rather than accumulating one-off live fixes:

  1. Controlled pair rotation/migration. Add runner-owned one-time rotation, revocation and rollback state; optionally support overlapping old/new credentials for a bounded cutover window.
  2. Remove the legacy proxy-contur secret source. Generate/manage a dedicated Map egress credential as first-class runner state.
  3. VPN-route enforcement. Add a tested Windows outbound kill switch or a dedicated egress appliance so VPN-off cannot become direct AMD egress.
  4. Unattended host recovery. Replace user-sign-in-dependent Docker Desktop with a managed Windows service, signed installer or dedicated Linux egress node if 24/7 cold-cache availability is required.
  5. Pair-channel hardening. Replace source-IP-only HTTP pairing with a short-lived nonce plus authenticated/encrypted handoff or mutual TLS.
  6. Two-level health. Keep local /healthz, add a rate-limited synthetic egress readiness check that proves DNS/CONNECT/TLS without consuming or exposing the master Ion credential.
  7. Durable metrics. Export the safe counters to the platform metrics stack; process-local /status currently resets on restart.
  8. Central allowlist contract. Generate identical proxy/connector host policies from one reviewed provider contract and test that they cannot drift.
  9. Circuit breaking and backoff. Add per-origin fail-fast state for a dead AMD/VPN path so many cold misses do not occupy all queues until timeout.
  10. High availability. Define an explicit active/standby connector model, health-based selection and credential isolation before adding a second workstation.
  11. Signed Windows distribution/update. Provide a versioned package builder, signature verification and in-place upgrade/rollback rather than manual archive handling.
  12. SLOs and load tests. Establish cache-hit, cold-TTFB, queue-depth, abort-rate and recovery-time objectives using realistic concurrent Map Page sessions.

Until those changes land, preserve the narrow boundary: official allowlisted HTTPS only, no direct NAS fallback, no global workstation proxy, no secret in source, and every NAS mutation through a reviewed data-only artifact plus the canonical runner.