feat(simulation): add S1 worker lifecycle ownership
This commit is contained in:
parent
12fba7126d
commit
630d1ae819
|
|
@ -151,6 +151,15 @@ a new linked run/episode. This is repository-level lifecycle acceptance only;
|
||||||
PX4 command delivery, watchdog/failsafe behavior, navigation behavior and real
|
PX4 command delivery, watchdog/failsafe behavior, navigation behavior and real
|
||||||
actuator control remain unaccepted.
|
actuator control remain unaccepted.
|
||||||
|
|
||||||
|
The next S1B repository increment adds the transport-neutral
|
||||||
|
`SimulationApplicationService`, an exact-S0/D-only worker admission guard and a
|
||||||
|
POSIX process supervisor. Start is idempotent, a second active owner is
|
||||||
|
rejected, provider processes receive dedicated PGIDs and run-scoped logs, and
|
||||||
|
shutdown follows the S0 dependency order with residue reported as failure.
|
||||||
|
Pause/resume/step/reset/stop are owned by the application service through a
|
||||||
|
worker port. Stock-rover provider specs, target-worker execution and PX4/Gazebo
|
||||||
|
control mapping are still pending factual acceptance.
|
||||||
|
|
||||||
See the [Polygon product/SRS](docs/12_SIMULATION_POLYGON_PRODUCT_AND_SRS.md),
|
See the [Polygon product/SRS](docs/12_SIMULATION_POLYGON_PRODUCT_AND_SRS.md),
|
||||||
[ADR 0015](docs/adr/0015-simulation-polygon-qualification-boundary.md) and the
|
[ADR 0015](docs/adr/0015-simulation-polygon-qualification-boundary.md) and the
|
||||||
[SIM S0 worker runbook](docs/runbooks/SIM_S0_AI_WORKER.md).
|
[SIM S0 worker runbook](docs/runbooks/SIM_S0_AI_WORKER.md).
|
||||||
|
|
|
||||||
|
|
@ -24,7 +24,7 @@ Each gate produces evidence and an explicit GO, PAUSE or BLOCKED result.
|
||||||
| Plugin isolation | GO (laboratory control plane) — vendor backend/frontend and optional scene controls are plugin-owned; manifest/runtime descriptor parity, versioned handshake, lifecycle health and transport correlation fail closed while execution remains in-process |
|
| Plugin isolation | GO (laboratory control plane) — vendor backend/frontend and optional scene controls are plugin-owned; manifest/runtime descriptor parity, versioned handshake, lifecycle health and transport correlation fail closed while execution remains in-process |
|
||||||
| K1 application control | GO (physical staged cycle) — after fixing the PCAP-proven `sint64` time field, one explicit UI launch completed all 14 canonical operations on one control session, reached live `SCANNING + project + init_ready`, displayed real points, then one explicit STOP returned K1 to unbound `READY`. No retry or fallback command was sent. Native-project reuse through LixelGO/USB remains an independent verification |
|
| K1 application control | GO (physical staged cycle) — after fixing the PCAP-proven `sint64` time field, one explicit UI launch completed all 14 canonical operations on one control session, reached live `SCANNING + project + init_ready`, displayed real points, then one explicit STOP returned K1 to unbound `READY`. No retry or fallback command was sent. Native-project reuse through LixelGO/USB remains an independent verification |
|
||||||
| Stage 8 product storage | PAUSE — retention, replication, encryption, capacity monitoring and long-run browser/WASM stress remain deployment gates |
|
| Stage 8 product storage | PAUSE — retention, replication, encryption, capacity monitoring and long-run browser/WASM stress remain deployment gates |
|
||||||
| Simulation Polygon | SIM S0 GO; S1A repository boundary implemented — dedicated Docker-free D-only worker and exact stack/clock/resource evidence are accepted. Typed immutable run identity, lifecycle journal, canonical rover command envelopes, authority/TTL checks, artifact index, reset identity and fail-closed restart recovery are implemented locally. Provider process orchestration, PX4 command delivery, watchdog/failsafe evidence, navigation/safety acceptance and real actuator authority remain absent |
|
| Simulation Polygon | SIM S0 GO; S1A complete; S1B ownership foundation implemented locally — immutable run/journal/command contracts, exact-S0/D-only admission, idempotent application lifecycle, one-active-run authority and PGID process supervision with deterministic shutdown are implemented and tested. Target stock-rover provider specs/execution, PX4 command delivery, telemetry/frames, watchdog/failsafe evidence, navigation/safety acceptance and real actuator authority remain absent |
|
||||||
|
|
||||||
USB project copying remains optional ground truth rather than a blocker for the
|
USB project copying remains optional ground truth rather than a blocker for the
|
||||||
now-verified network path. Owner-operated LixelGO traffic verifies the MQTT
|
now-verified network path. Owner-operated LixelGO traffic verifies the MQTT
|
||||||
|
|
@ -152,9 +152,11 @@ Without the exact admitted target-worker evidence its correct result remains
|
||||||
`INCOMPLETE`. S1 starts from the accepted S0 version/storage/time/process
|
`INCOMPLETE`. S1 starts from the accepted S0 version/storage/time/process
|
||||||
boundary. Its S1A increment implements server-owned immutable run identity,
|
boundary. Its S1A increment implements server-owned immutable run identity,
|
||||||
legal lifecycle transitions, canonical rover command envelopes, authority/TTL
|
legal lifecycle transitions, canonical rover command envelopes, authority/TTL
|
||||||
admission, append-only artifacts and restart reconciliation. Process ownership,
|
admission, append-only artifacts and restart reconciliation. The local S1B
|
||||||
PX4 delivery, heartbeat/watchdog execution and captured failsafe outcomes remain
|
increment adds idempotent application lifecycle, single active ownership, an
|
||||||
the next S1 increments.
|
exact-profile/D-only worker guard and deterministic POSIX PGID supervision.
|
||||||
|
Target provider launch/control, PX4 delivery, heartbeat/watchdog execution and
|
||||||
|
captured failsafe outcomes remain the next S1 increments.
|
||||||
|
|
||||||
## Stage 0 — repository and host baseline
|
## Stage 0 — repository and host baseline
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -534,6 +534,36 @@ portable. The target Simulation Orchestrator must bind that root under the
|
||||||
accepted `/mnt/d/NDC_MISSIONCORE` storage boundary; S1 provider orchestration
|
accepted `/mnt/d/NDC_MISSIONCORE` storage boundary; S1 provider orchestration
|
||||||
and target evidence are not implied by this increment.
|
and target evidence are not implied by this increment.
|
||||||
|
|
||||||
|
#### S1B ownership foundation status
|
||||||
|
|
||||||
|
The next local increment implements the provider-neutral ownership boundary:
|
||||||
|
|
||||||
|
- `SimulationApplicationService` is the only caller allowed to drive persisted
|
||||||
|
start, pause, resume, single-step, stop, reset and restart reconciliation;
|
||||||
|
- start and stop requests are journal-bound to idempotency keys, and a
|
||||||
|
different request cannot rebind the same run operation;
|
||||||
|
- a second `starting|running|paused|stopping` run cannot acquire the same worker
|
||||||
|
authority;
|
||||||
|
- the worker port exposes lifecycle operations without exposing a browser-to-PX4
|
||||||
|
channel or choosing a remote transport;
|
||||||
|
- `S0WorkerGuard` requires the run's exact accepted S0 profile ID/SHA-256, the
|
||||||
|
canonical `/mnt/d/NDC_MISSIONCORE/simulation/artifacts/s1/runs/<run-id>`
|
||||||
|
artifact layout and the actual
|
||||||
|
`/mnt/d/NDC_MISSIONCORE/simulation/runtime/s1/runs/<run-id>` supervisor
|
||||||
|
working layout before provider start;
|
||||||
|
- `PosixProcessSupervisor` launches argv without a shell, creates one dedicated
|
||||||
|
session/PGID per provider, writes run-scoped stdout/stderr and a process
|
||||||
|
registry, rolls back partial starts, and stops in ascending S0 shutdown order;
|
||||||
|
- failure, profile drift or process residue becomes a persisted terminal
|
||||||
|
failure rather than a successful stop.
|
||||||
|
|
||||||
|
This is a tested local foundation, not S1B target acceptance. Concrete
|
||||||
|
Gazebo/PX4/XRCE/ROS provider specs, loopback namespace creation, target
|
||||||
|
health probes, post-restart PID identity reconciliation and live
|
||||||
|
pause/step/reset evidence remain open. The application service's port can be
|
||||||
|
implemented worker-locally without turning SSH or an operator shell into the
|
||||||
|
product data plane.
|
||||||
|
|
||||||
### S2
|
### S2
|
||||||
|
|
||||||
- Gazebo LiDAR/odometry feed the ROS 2/Nav2 baseline.
|
- Gazebo LiDAR/odometry feed the ROS 2/Nav2 baseline.
|
||||||
|
|
|
||||||
|
|
@ -125,3 +125,25 @@ own provider processes, transmit a PX4 setpoint, execute heartbeat/watchdog
|
||||||
failsafe behavior, or qualify navigation/safety behavior. The target
|
failsafe behavior, or qualify navigation/safety behavior. The target
|
||||||
orchestrator must place the caller-supplied repository root under the accepted
|
orchestrator must place the caller-supplied repository root under the accepted
|
||||||
D-only worker boundary.
|
D-only worker boundary.
|
||||||
|
|
||||||
|
## S1B ownership foundation
|
||||||
|
|
||||||
|
`k1link.simulation.orchestrator` now owns the application lifecycle through a
|
||||||
|
transport-neutral worker port. Start and stop requests are persisted with
|
||||||
|
idempotency identity, one active run owns a worker at a time, and
|
||||||
|
pause/resume/step/reset cannot bypass run-state validation. Provider start
|
||||||
|
failure, S0 profile drift, stop residue and restart interruption all become
|
||||||
|
explicit run events and terminal failures.
|
||||||
|
|
||||||
|
`k1link.simulation.process_supervisor` provides the worker-local POSIX ownership
|
||||||
|
primitive: shell-free argv, dedicated provider PGIDs, run-scoped logs, atomic
|
||||||
|
process registry, partial-start rollback and ascending shutdown order matching
|
||||||
|
the accepted S0 dependency graph. `k1link.simulation.worker` binds that
|
||||||
|
primitive to the exact accepted S0 profile digest and canonical D-only S1 run
|
||||||
|
artifact and process-runtime layouts.
|
||||||
|
|
||||||
|
This is not target-worker acceptance. Provider specifications, loopback network
|
||||||
|
namespace creation, live readiness probes, restart-safe PID/start-token
|
||||||
|
reconciliation and Gazebo world-control/PX4 command adapters remain open. The
|
||||||
|
worker port deliberately fixes ownership semantics without prematurely fixing
|
||||||
|
the eventual transport.
|
||||||
|
|
|
||||||
|
|
@ -16,6 +16,21 @@ from k1link.simulation.contracts import (
|
||||||
RunState,
|
RunState,
|
||||||
SimulationContractError,
|
SimulationContractError,
|
||||||
)
|
)
|
||||||
|
from k1link.simulation.orchestrator import (
|
||||||
|
ActiveQualificationRunError,
|
||||||
|
SimulationApplicationService,
|
||||||
|
SimulationOrchestratorError,
|
||||||
|
SimulationWorkerPort,
|
||||||
|
WorkerStartResult,
|
||||||
|
WorkerStopResult,
|
||||||
|
)
|
||||||
|
from k1link.simulation.process_supervisor import (
|
||||||
|
OwnedProcess,
|
||||||
|
PosixProcessSupervisor,
|
||||||
|
ProcessSpec,
|
||||||
|
ProcessStopResult,
|
||||||
|
ProcessSupervisorError,
|
||||||
|
)
|
||||||
from k1link.simulation.run_store import (
|
from k1link.simulation.run_store import (
|
||||||
QualificationRunConflictError,
|
QualificationRunConflictError,
|
||||||
QualificationRunIntegrityError,
|
QualificationRunIntegrityError,
|
||||||
|
|
@ -34,9 +49,17 @@ from k1link.simulation.s0 import (
|
||||||
load_s0_profile,
|
load_s0_profile,
|
||||||
run_s0_doctor,
|
run_s0_doctor,
|
||||||
)
|
)
|
||||||
|
from k1link.simulation.worker import (
|
||||||
|
LocalProcessWorkerAdapter,
|
||||||
|
S0WorkerGuard,
|
||||||
|
SimulationWorldControl,
|
||||||
|
WorkerAdmission,
|
||||||
|
WorkerAdmissionError,
|
||||||
|
)
|
||||||
|
|
||||||
__all__ = [
|
__all__ = [
|
||||||
"AckermannControlSetpoint",
|
"AckermannControlSetpoint",
|
||||||
|
"ActiveQualificationRunError",
|
||||||
"AuthorityProfile",
|
"AuthorityProfile",
|
||||||
"CheckStatus",
|
"CheckStatus",
|
||||||
"CommandAuthorityScope",
|
"CommandAuthorityScope",
|
||||||
|
|
@ -44,6 +67,12 @@ __all__ = [
|
||||||
"ControlSetpoint",
|
"ControlSetpoint",
|
||||||
"DifferentialControlSetpoint",
|
"DifferentialControlSetpoint",
|
||||||
"DoctorVerdict",
|
"DoctorVerdict",
|
||||||
|
"LocalProcessWorkerAdapter",
|
||||||
|
"OwnedProcess",
|
||||||
|
"PosixProcessSupervisor",
|
||||||
|
"ProcessSpec",
|
||||||
|
"ProcessStopResult",
|
||||||
|
"ProcessSupervisorError",
|
||||||
"ProviderPin",
|
"ProviderPin",
|
||||||
"QualificationArtifact",
|
"QualificationArtifact",
|
||||||
"QualificationEvent",
|
"QualificationEvent",
|
||||||
|
|
@ -58,10 +87,19 @@ __all__ = [
|
||||||
"RuntimeAcceptance",
|
"RuntimeAcceptance",
|
||||||
"RunKind",
|
"RunKind",
|
||||||
"RunState",
|
"RunState",
|
||||||
|
"S0WorkerGuard",
|
||||||
"S0DoctorReport",
|
"S0DoctorReport",
|
||||||
"S0Profile",
|
"S0Profile",
|
||||||
"S0ProfileError",
|
"S0ProfileError",
|
||||||
"SimulationContractError",
|
"SimulationContractError",
|
||||||
|
"SimulationApplicationService",
|
||||||
|
"SimulationOrchestratorError",
|
||||||
|
"SimulationWorkerPort",
|
||||||
|
"SimulationWorldControl",
|
||||||
|
"WorkerAdmission",
|
||||||
|
"WorkerAdmissionError",
|
||||||
|
"WorkerStartResult",
|
||||||
|
"WorkerStopResult",
|
||||||
"load_s0_profile",
|
"load_s0_profile",
|
||||||
"run_s0_doctor",
|
"run_s0_doctor",
|
||||||
]
|
]
|
||||||
|
|
|
||||||
|
|
@ -0,0 +1,544 @@
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import threading
|
||||||
|
from contextlib import suppress
|
||||||
|
from dataclasses import dataclass
|
||||||
|
from typing import Protocol
|
||||||
|
|
||||||
|
from k1link.simulation.contracts import QualificationEvent, QualificationRun, RunState
|
||||||
|
from k1link.simulation.run_store import (
|
||||||
|
ACTIVE_RECOVERY_STATES,
|
||||||
|
QualificationRunConflictError,
|
||||||
|
QualificationRunStore,
|
||||||
|
QualificationRunTransitionError,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
class SimulationOrchestratorError(RuntimeError):
|
||||||
|
"""The server-owned S1 lifecycle could not be completed safely."""
|
||||||
|
|
||||||
|
|
||||||
|
class ActiveQualificationRunError(SimulationOrchestratorError):
|
||||||
|
"""Another qualification run already owns the worker authority."""
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True, slots=True)
|
||||||
|
class WorkerStartResult:
|
||||||
|
provider_ids: tuple[str, ...]
|
||||||
|
profile_sha256: str
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True, slots=True)
|
||||||
|
class WorkerStopResult:
|
||||||
|
stopped_provider_ids: tuple[str, ...]
|
||||||
|
residue_provider_ids: tuple[str, ...]
|
||||||
|
|
||||||
|
@property
|
||||||
|
def clean(self) -> bool:
|
||||||
|
return not self.residue_provider_ids
|
||||||
|
|
||||||
|
|
||||||
|
class SimulationWorkerPort(Protocol):
|
||||||
|
"""Transport-neutral worker boundary; no browser/PX4 shortcut exists."""
|
||||||
|
|
||||||
|
def start(self, run: QualificationRun) -> WorkerStartResult: ...
|
||||||
|
|
||||||
|
def pause(self, run: QualificationRun) -> None: ...
|
||||||
|
|
||||||
|
def resume(self, run: QualificationRun) -> None: ...
|
||||||
|
|
||||||
|
def step(self, run: QualificationRun, step_count: int) -> int: ...
|
||||||
|
|
||||||
|
def stop(self, run: QualificationRun) -> WorkerStopResult: ...
|
||||||
|
|
||||||
|
def reconcile(self, run: QualificationRun) -> WorkerStopResult: ...
|
||||||
|
|
||||||
|
|
||||||
|
class SimulationApplicationService:
|
||||||
|
"""Sole owner of one worker's persisted qualification-run lifecycle."""
|
||||||
|
|
||||||
|
def __init__(self, store: QualificationRunStore, worker: SimulationWorkerPort) -> None:
|
||||||
|
self.store = store
|
||||||
|
self.worker = worker
|
||||||
|
self._lock = threading.RLock()
|
||||||
|
|
||||||
|
def start(
|
||||||
|
self,
|
||||||
|
run_id: str,
|
||||||
|
*,
|
||||||
|
idempotency_key: str,
|
||||||
|
observed_at_utc: str,
|
||||||
|
host_monotonic_ns: int,
|
||||||
|
) -> QualificationRun:
|
||||||
|
key = _idempotency_key(idempotency_key)
|
||||||
|
with self._lock:
|
||||||
|
run = self.store.load(run_id)
|
||||||
|
existing = self._start_request(run_id, key)
|
||||||
|
if existing:
|
||||||
|
if run.state is RunState.ADMITTED:
|
||||||
|
return self._continue_start(
|
||||||
|
run,
|
||||||
|
observed_at_utc=observed_at_utc,
|
||||||
|
host_monotonic_ns=host_monotonic_ns,
|
||||||
|
)
|
||||||
|
return run
|
||||||
|
if self._has_other_start_request(run_id):
|
||||||
|
raise QualificationRunConflictError(
|
||||||
|
"qualification run is already bound to another start request"
|
||||||
|
)
|
||||||
|
active = tuple(
|
||||||
|
candidate
|
||||||
|
for candidate in self.store.list_runs()
|
||||||
|
if candidate.run_id != run_id and candidate.state in ACTIVE_RECOVERY_STATES
|
||||||
|
)
|
||||||
|
if active:
|
||||||
|
raise ActiveQualificationRunError(
|
||||||
|
f"worker authority is owned by {active[0].run_id}"
|
||||||
|
)
|
||||||
|
if run.state is not RunState.ADMITTED:
|
||||||
|
raise QualificationRunTransitionError(
|
||||||
|
"only an admitted qualification run can start"
|
||||||
|
)
|
||||||
|
event = self.store.append_event(
|
||||||
|
run_id,
|
||||||
|
event_type="orchestrator.start-requested",
|
||||||
|
observed_at_utc=observed_at_utc,
|
||||||
|
host_monotonic_ns=host_monotonic_ns,
|
||||||
|
payload={"idempotency_key": key},
|
||||||
|
expected_revision=run.revision,
|
||||||
|
)
|
||||||
|
run = self.store.load(run_id)
|
||||||
|
if run.revision != event.sequence:
|
||||||
|
raise QualificationRunConflictError("start request journal changed")
|
||||||
|
return self._continue_start(
|
||||||
|
run,
|
||||||
|
observed_at_utc=observed_at_utc,
|
||||||
|
host_monotonic_ns=host_monotonic_ns,
|
||||||
|
)
|
||||||
|
|
||||||
|
def pause(
|
||||||
|
self,
|
||||||
|
run_id: str,
|
||||||
|
*,
|
||||||
|
observed_at_utc: str,
|
||||||
|
host_monotonic_ns: int,
|
||||||
|
sim_time_ns: int,
|
||||||
|
) -> QualificationRun:
|
||||||
|
with self._lock:
|
||||||
|
run = self.store.load(run_id)
|
||||||
|
if run.state is not RunState.RUNNING:
|
||||||
|
raise QualificationRunTransitionError("only a running run can pause")
|
||||||
|
try:
|
||||||
|
self.worker.pause(run)
|
||||||
|
except Exception as exc:
|
||||||
|
return self._fail_worker_operation(
|
||||||
|
run,
|
||||||
|
operation="pause",
|
||||||
|
exc=exc,
|
||||||
|
observed_at_utc=observed_at_utc,
|
||||||
|
host_monotonic_ns=host_monotonic_ns,
|
||||||
|
sim_time_ns=sim_time_ns,
|
||||||
|
)
|
||||||
|
return self.store.transition(
|
||||||
|
run_id,
|
||||||
|
RunState.PAUSED,
|
||||||
|
expected_revision=run.revision,
|
||||||
|
observed_at_utc=observed_at_utc,
|
||||||
|
host_monotonic_ns=host_monotonic_ns,
|
||||||
|
sim_time_ns=sim_time_ns,
|
||||||
|
)
|
||||||
|
|
||||||
|
def resume(
|
||||||
|
self,
|
||||||
|
run_id: str,
|
||||||
|
*,
|
||||||
|
observed_at_utc: str,
|
||||||
|
host_monotonic_ns: int,
|
||||||
|
sim_time_ns: int,
|
||||||
|
) -> QualificationRun:
|
||||||
|
with self._lock:
|
||||||
|
run = self.store.load(run_id)
|
||||||
|
if run.state is not RunState.PAUSED:
|
||||||
|
raise QualificationRunTransitionError("only a paused run can resume")
|
||||||
|
try:
|
||||||
|
self.worker.resume(run)
|
||||||
|
except Exception as exc:
|
||||||
|
return self._fail_worker_operation(
|
||||||
|
run,
|
||||||
|
operation="resume",
|
||||||
|
exc=exc,
|
||||||
|
observed_at_utc=observed_at_utc,
|
||||||
|
host_monotonic_ns=host_monotonic_ns,
|
||||||
|
sim_time_ns=sim_time_ns,
|
||||||
|
)
|
||||||
|
return self.store.transition(
|
||||||
|
run_id,
|
||||||
|
RunState.RUNNING,
|
||||||
|
expected_revision=run.revision,
|
||||||
|
observed_at_utc=observed_at_utc,
|
||||||
|
host_monotonic_ns=host_monotonic_ns,
|
||||||
|
sim_time_ns=sim_time_ns,
|
||||||
|
)
|
||||||
|
|
||||||
|
def step(
|
||||||
|
self,
|
||||||
|
run_id: str,
|
||||||
|
*,
|
||||||
|
step_count: int,
|
||||||
|
observed_at_utc: str,
|
||||||
|
host_monotonic_ns: int,
|
||||||
|
sim_time_ns: int,
|
||||||
|
) -> int:
|
||||||
|
if step_count < 1:
|
||||||
|
raise ValueError("step count must be positive")
|
||||||
|
with self._lock:
|
||||||
|
run = self.store.load(run_id)
|
||||||
|
if run.state is not RunState.PAUSED:
|
||||||
|
raise QualificationRunTransitionError("single-step requires a paused run")
|
||||||
|
try:
|
||||||
|
advanced_ns = self.worker.step(run, step_count)
|
||||||
|
except Exception as exc:
|
||||||
|
self._fail_worker_operation(
|
||||||
|
run,
|
||||||
|
operation="step",
|
||||||
|
exc=exc,
|
||||||
|
observed_at_utc=observed_at_utc,
|
||||||
|
host_monotonic_ns=host_monotonic_ns,
|
||||||
|
sim_time_ns=sim_time_ns,
|
||||||
|
)
|
||||||
|
raise SimulationOrchestratorError("worker step failed") from exc
|
||||||
|
if advanced_ns < 1:
|
||||||
|
self._fail_worker_operation(
|
||||||
|
run,
|
||||||
|
operation="step-invalid-advance",
|
||||||
|
exc=ValueError("worker returned a non-positive clock advance"),
|
||||||
|
observed_at_utc=observed_at_utc,
|
||||||
|
host_monotonic_ns=host_monotonic_ns,
|
||||||
|
sim_time_ns=sim_time_ns,
|
||||||
|
)
|
||||||
|
raise SimulationOrchestratorError(
|
||||||
|
"worker step returned a non-positive clock advance"
|
||||||
|
)
|
||||||
|
self.store.append_event(
|
||||||
|
run_id,
|
||||||
|
event_type="clock.single-step",
|
||||||
|
observed_at_utc=observed_at_utc,
|
||||||
|
host_monotonic_ns=host_monotonic_ns,
|
||||||
|
sim_time_ns=sim_time_ns + advanced_ns,
|
||||||
|
payload={"step_count": step_count, "advanced_ns": advanced_ns},
|
||||||
|
expected_revision=run.revision,
|
||||||
|
)
|
||||||
|
return advanced_ns
|
||||||
|
|
||||||
|
def stop(
|
||||||
|
self,
|
||||||
|
run_id: str,
|
||||||
|
*,
|
||||||
|
idempotency_key: str,
|
||||||
|
observed_at_utc: str,
|
||||||
|
host_monotonic_ns: int,
|
||||||
|
sim_time_ns: int | None,
|
||||||
|
terminal_state: RunState = RunState.COMPLETED,
|
||||||
|
reason: str = "operator-stop-clean",
|
||||||
|
) -> QualificationRun:
|
||||||
|
key = _idempotency_key(idempotency_key)
|
||||||
|
if terminal_state not in {RunState.COMPLETED, RunState.ABORTED}:
|
||||||
|
raise ValueError("operator stop terminal state must be completed or aborted")
|
||||||
|
with self._lock:
|
||||||
|
run = self.store.load(run_id)
|
||||||
|
existing = self._stop_request(run_id, key)
|
||||||
|
if existing is not None and (
|
||||||
|
existing.payload.get("terminal_state") != terminal_state.value
|
||||||
|
or existing.payload.get("reason") != reason
|
||||||
|
):
|
||||||
|
raise QualificationRunConflictError(
|
||||||
|
"idempotency key is bound to a different stop request"
|
||||||
|
)
|
||||||
|
if existing is not None and run.state.terminal:
|
||||||
|
return run
|
||||||
|
if existing is None:
|
||||||
|
if self._has_other_stop_request(run_id):
|
||||||
|
raise QualificationRunConflictError(
|
||||||
|
"qualification run is already bound to another stop request"
|
||||||
|
)
|
||||||
|
event = self.store.append_event(
|
||||||
|
run_id,
|
||||||
|
event_type="orchestrator.stop-requested",
|
||||||
|
observed_at_utc=observed_at_utc,
|
||||||
|
host_monotonic_ns=host_monotonic_ns,
|
||||||
|
sim_time_ns=sim_time_ns,
|
||||||
|
payload={
|
||||||
|
"idempotency_key": key,
|
||||||
|
"terminal_state": terminal_state.value,
|
||||||
|
"reason": reason,
|
||||||
|
},
|
||||||
|
expected_revision=run.revision,
|
||||||
|
)
|
||||||
|
run = self.store.load(run_id)
|
||||||
|
if run.revision != event.sequence:
|
||||||
|
raise QualificationRunConflictError("stop request journal changed")
|
||||||
|
if run.state not in {RunState.RUNNING, RunState.PAUSED, RunState.STOPPING}:
|
||||||
|
raise QualificationRunTransitionError(
|
||||||
|
"only a running, paused or stopping run can stop"
|
||||||
|
)
|
||||||
|
if run.state is not RunState.STOPPING:
|
||||||
|
run = self.store.transition(
|
||||||
|
run_id,
|
||||||
|
RunState.STOPPING,
|
||||||
|
expected_revision=run.revision,
|
||||||
|
observed_at_utc=observed_at_utc,
|
||||||
|
host_monotonic_ns=host_monotonic_ns,
|
||||||
|
sim_time_ns=sim_time_ns,
|
||||||
|
)
|
||||||
|
try:
|
||||||
|
outcome = self.worker.stop(run)
|
||||||
|
except Exception as exc:
|
||||||
|
return self._terminal_failure(
|
||||||
|
run,
|
||||||
|
operation="stop",
|
||||||
|
detail=type(exc).__name__,
|
||||||
|
observed_at_utc=observed_at_utc,
|
||||||
|
host_monotonic_ns=host_monotonic_ns,
|
||||||
|
sim_time_ns=sim_time_ns,
|
||||||
|
)
|
||||||
|
if not outcome.clean:
|
||||||
|
return self._terminal_failure(
|
||||||
|
run,
|
||||||
|
operation="stop-residue",
|
||||||
|
detail=",".join(outcome.residue_provider_ids),
|
||||||
|
observed_at_utc=observed_at_utc,
|
||||||
|
host_monotonic_ns=host_monotonic_ns,
|
||||||
|
sim_time_ns=sim_time_ns,
|
||||||
|
)
|
||||||
|
event = self.store.append_event(
|
||||||
|
run_id,
|
||||||
|
event_type="orchestrator.providers-stopped",
|
||||||
|
observed_at_utc=observed_at_utc,
|
||||||
|
host_monotonic_ns=host_monotonic_ns,
|
||||||
|
sim_time_ns=sim_time_ns,
|
||||||
|
payload={"provider_ids": list(outcome.stopped_provider_ids)},
|
||||||
|
expected_revision=run.revision,
|
||||||
|
)
|
||||||
|
return self.store.transition(
|
||||||
|
run_id,
|
||||||
|
terminal_state,
|
||||||
|
expected_revision=event.sequence,
|
||||||
|
observed_at_utc=observed_at_utc,
|
||||||
|
host_monotonic_ns=host_monotonic_ns,
|
||||||
|
sim_time_ns=sim_time_ns,
|
||||||
|
reason=reason,
|
||||||
|
)
|
||||||
|
|
||||||
|
def reset(
|
||||||
|
self,
|
||||||
|
run_id: str,
|
||||||
|
*,
|
||||||
|
idempotency_key: str,
|
||||||
|
new_run_id: str,
|
||||||
|
new_episode_id: str,
|
||||||
|
observed_at_utc: str,
|
||||||
|
host_monotonic_ns: int,
|
||||||
|
sim_time_ns: int | None,
|
||||||
|
) -> QualificationRun:
|
||||||
|
with self._lock:
|
||||||
|
key = _idempotency_key(idempotency_key)
|
||||||
|
if len(key) > 155:
|
||||||
|
raise ValueError("reset idempotency key must contain at most 155 characters")
|
||||||
|
terminal = self.stop(
|
||||||
|
run_id,
|
||||||
|
idempotency_key=f"{key}:stop",
|
||||||
|
observed_at_utc=observed_at_utc,
|
||||||
|
host_monotonic_ns=host_monotonic_ns,
|
||||||
|
sim_time_ns=sim_time_ns,
|
||||||
|
terminal_state=RunState.ABORTED,
|
||||||
|
reason="operator-reset",
|
||||||
|
)
|
||||||
|
return self.store.create_reset_episode(
|
||||||
|
terminal.run_id,
|
||||||
|
run_id=new_run_id,
|
||||||
|
episode_id=new_episode_id,
|
||||||
|
created_at_utc=observed_at_utc,
|
||||||
|
)
|
||||||
|
|
||||||
|
def reconcile(
|
||||||
|
self,
|
||||||
|
*,
|
||||||
|
observed_at_utc: str,
|
||||||
|
host_monotonic_ns: int,
|
||||||
|
) -> tuple[QualificationRun, ...]:
|
||||||
|
reconciled: list[QualificationRun] = []
|
||||||
|
with self._lock:
|
||||||
|
for run in self.store.list_runs():
|
||||||
|
if run.state not in ACTIVE_RECOVERY_STATES:
|
||||||
|
continue
|
||||||
|
try:
|
||||||
|
outcome = self.worker.reconcile(run)
|
||||||
|
residue = outcome.residue_provider_ids
|
||||||
|
except Exception as exc:
|
||||||
|
residue = (f"reconcile-error:{type(exc).__name__}",)
|
||||||
|
if residue:
|
||||||
|
event = self.store.append_event(
|
||||||
|
run.run_id,
|
||||||
|
event_type="orchestrator.recovery-residue",
|
||||||
|
observed_at_utc=observed_at_utc,
|
||||||
|
host_monotonic_ns=host_monotonic_ns,
|
||||||
|
payload={"provider_ids": list(residue)},
|
||||||
|
expected_revision=run.revision,
|
||||||
|
)
|
||||||
|
run = self.store.load(run.run_id)
|
||||||
|
if run.revision != event.sequence:
|
||||||
|
raise QualificationRunConflictError("recovery event journal changed")
|
||||||
|
reconciled.append(
|
||||||
|
self.store.reconcile_interrupted(
|
||||||
|
run.run_id,
|
||||||
|
observed_at_utc=observed_at_utc,
|
||||||
|
host_monotonic_ns=host_monotonic_ns,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
return tuple(reconciled)
|
||||||
|
|
||||||
|
def _continue_start(
|
||||||
|
self,
|
||||||
|
run: QualificationRun,
|
||||||
|
*,
|
||||||
|
observed_at_utc: str,
|
||||||
|
host_monotonic_ns: int,
|
||||||
|
) -> QualificationRun:
|
||||||
|
starting = self.store.transition(
|
||||||
|
run.run_id,
|
||||||
|
RunState.STARTING,
|
||||||
|
expected_revision=run.revision,
|
||||||
|
observed_at_utc=observed_at_utc,
|
||||||
|
host_monotonic_ns=host_monotonic_ns,
|
||||||
|
)
|
||||||
|
try:
|
||||||
|
outcome = self.worker.start(starting)
|
||||||
|
except Exception as exc:
|
||||||
|
with suppress(Exception):
|
||||||
|
self.worker.stop(starting)
|
||||||
|
return self._terminal_failure(
|
||||||
|
starting,
|
||||||
|
operation="start",
|
||||||
|
detail=type(exc).__name__,
|
||||||
|
observed_at_utc=observed_at_utc,
|
||||||
|
host_monotonic_ns=host_monotonic_ns,
|
||||||
|
sim_time_ns=0,
|
||||||
|
)
|
||||||
|
if outcome.profile_sha256 != starting.host_profile_sha256:
|
||||||
|
with suppress(Exception):
|
||||||
|
self.worker.stop(starting)
|
||||||
|
return self._terminal_failure(
|
||||||
|
starting,
|
||||||
|
operation="profile-drift",
|
||||||
|
detail=outcome.profile_sha256,
|
||||||
|
observed_at_utc=observed_at_utc,
|
||||||
|
host_monotonic_ns=host_monotonic_ns,
|
||||||
|
sim_time_ns=0,
|
||||||
|
)
|
||||||
|
event = self.store.append_event(
|
||||||
|
run.run_id,
|
||||||
|
event_type="orchestrator.providers-started",
|
||||||
|
observed_at_utc=observed_at_utc,
|
||||||
|
host_monotonic_ns=host_monotonic_ns,
|
||||||
|
sim_time_ns=0,
|
||||||
|
payload={
|
||||||
|
"provider_ids": list(outcome.provider_ids),
|
||||||
|
"profile_sha256": outcome.profile_sha256,
|
||||||
|
},
|
||||||
|
expected_revision=starting.revision,
|
||||||
|
)
|
||||||
|
return self.store.transition(
|
||||||
|
run.run_id,
|
||||||
|
RunState.RUNNING,
|
||||||
|
expected_revision=event.sequence,
|
||||||
|
observed_at_utc=observed_at_utc,
|
||||||
|
host_monotonic_ns=host_monotonic_ns,
|
||||||
|
sim_time_ns=0,
|
||||||
|
)
|
||||||
|
|
||||||
|
def _fail_worker_operation(
|
||||||
|
self,
|
||||||
|
run: QualificationRun,
|
||||||
|
*,
|
||||||
|
operation: str,
|
||||||
|
exc: Exception,
|
||||||
|
observed_at_utc: str,
|
||||||
|
host_monotonic_ns: int,
|
||||||
|
sim_time_ns: int | None,
|
||||||
|
) -> QualificationRun:
|
||||||
|
with suppress(Exception):
|
||||||
|
self.worker.stop(run)
|
||||||
|
return self._terminal_failure(
|
||||||
|
run,
|
||||||
|
operation=operation,
|
||||||
|
detail=type(exc).__name__,
|
||||||
|
observed_at_utc=observed_at_utc,
|
||||||
|
host_monotonic_ns=host_monotonic_ns,
|
||||||
|
sim_time_ns=sim_time_ns,
|
||||||
|
)
|
||||||
|
|
||||||
|
def _terminal_failure(
|
||||||
|
self,
|
||||||
|
run: QualificationRun,
|
||||||
|
*,
|
||||||
|
operation: str,
|
||||||
|
detail: str,
|
||||||
|
observed_at_utc: str,
|
||||||
|
host_monotonic_ns: int,
|
||||||
|
sim_time_ns: int | None,
|
||||||
|
) -> QualificationRun:
|
||||||
|
event = self.store.append_event(
|
||||||
|
run.run_id,
|
||||||
|
event_type="orchestrator.operation-failed",
|
||||||
|
observed_at_utc=observed_at_utc,
|
||||||
|
host_monotonic_ns=host_monotonic_ns,
|
||||||
|
sim_time_ns=sim_time_ns,
|
||||||
|
payload={"operation": operation, "detail": detail[:256]},
|
||||||
|
expected_revision=run.revision,
|
||||||
|
)
|
||||||
|
return self.store.transition(
|
||||||
|
run.run_id,
|
||||||
|
RunState.FAILED,
|
||||||
|
expected_revision=event.sequence,
|
||||||
|
observed_at_utc=observed_at_utc,
|
||||||
|
host_monotonic_ns=host_monotonic_ns,
|
||||||
|
sim_time_ns=sim_time_ns,
|
||||||
|
reason=f"orchestrator-{operation}-failed",
|
||||||
|
)
|
||||||
|
|
||||||
|
def _start_request(self, run_id: str, key: str) -> bool:
|
||||||
|
return any(
|
||||||
|
event.event_type == "orchestrator.start-requested"
|
||||||
|
and event.payload.get("idempotency_key") == key
|
||||||
|
for event in self.store.list_events(run_id)
|
||||||
|
)
|
||||||
|
|
||||||
|
def _has_other_start_request(self, run_id: str) -> bool:
|
||||||
|
return any(
|
||||||
|
event.event_type == "orchestrator.start-requested"
|
||||||
|
for event in self.store.list_events(run_id)
|
||||||
|
)
|
||||||
|
|
||||||
|
def _stop_request(self, run_id: str, key: str) -> QualificationEvent | None:
|
||||||
|
return next(
|
||||||
|
(
|
||||||
|
event
|
||||||
|
for event in self.store.list_events(run_id)
|
||||||
|
if event.event_type == "orchestrator.stop-requested"
|
||||||
|
and event.payload.get("idempotency_key") == key
|
||||||
|
),
|
||||||
|
None,
|
||||||
|
)
|
||||||
|
|
||||||
|
def _has_other_stop_request(self, run_id: str) -> bool:
|
||||||
|
return any(
|
||||||
|
event.event_type == "orchestrator.stop-requested"
|
||||||
|
for event in self.store.list_events(run_id)
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _idempotency_key(value: str) -> str:
|
||||||
|
normalized = value.strip()
|
||||||
|
if not 1 <= len(normalized) <= 160:
|
||||||
|
raise ValueError("idempotency key must contain 1..160 characters")
|
||||||
|
return normalized
|
||||||
|
|
@ -0,0 +1,302 @@
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import os
|
||||||
|
import re
|
||||||
|
import signal
|
||||||
|
import subprocess
|
||||||
|
import threading
|
||||||
|
import time
|
||||||
|
from dataclasses import dataclass
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Final
|
||||||
|
|
||||||
|
from k1link.artifacts import write_json_atomic
|
||||||
|
|
||||||
|
PROCESS_ID_PATTERN: Final = re.compile(r"^[a-z0-9][a-z0-9-]{0,63}$")
|
||||||
|
|
||||||
|
|
||||||
|
class ProcessSupervisorError(RuntimeError):
|
||||||
|
"""A provider process graph could not be owned or stopped safely."""
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True, slots=True)
|
||||||
|
class ProcessSpec:
|
||||||
|
process_id: str
|
||||||
|
argv: tuple[str, ...]
|
||||||
|
start_order: int
|
||||||
|
shutdown_order: int
|
||||||
|
environment: tuple[tuple[str, str], ...] = ()
|
||||||
|
startup_grace_seconds: float = 0.05
|
||||||
|
interrupt_timeout_seconds: float = 2.0
|
||||||
|
terminate_timeout_seconds: float = 1.0
|
||||||
|
|
||||||
|
def __post_init__(self) -> None:
|
||||||
|
if not PROCESS_ID_PATTERN.fullmatch(self.process_id):
|
||||||
|
raise ValueError("process id must be a safe lowercase identifier")
|
||||||
|
if not self.argv or any(not item or "\x00" in item for item in self.argv):
|
||||||
|
raise ValueError("process argv must contain nonempty NUL-free values")
|
||||||
|
if self.start_order < 0 or self.shutdown_order < 0:
|
||||||
|
raise ValueError("process order must not be negative")
|
||||||
|
if (
|
||||||
|
self.startup_grace_seconds < 0
|
||||||
|
or self.interrupt_timeout_seconds < 0
|
||||||
|
or self.terminate_timeout_seconds < 0
|
||||||
|
):
|
||||||
|
raise ValueError("process timeouts must not be negative")
|
||||||
|
keys = [key for key, _ in self.environment]
|
||||||
|
if len(keys) != len(set(keys)) or any(
|
||||||
|
not key or "=" in key or "\x00" in key or "\x00" in value
|
||||||
|
for key, value in self.environment
|
||||||
|
):
|
||||||
|
raise ValueError("process environment must contain unique safe entries")
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True, slots=True)
|
||||||
|
class OwnedProcess:
|
||||||
|
process_id: str
|
||||||
|
pid: int
|
||||||
|
pgid: int
|
||||||
|
start_order: int
|
||||||
|
shutdown_order: int
|
||||||
|
stdout_path: Path
|
||||||
|
stderr_path: Path
|
||||||
|
|
||||||
|
def to_dict(self) -> dict[str, object]:
|
||||||
|
return {
|
||||||
|
"process_id": self.process_id,
|
||||||
|
"pid": self.pid,
|
||||||
|
"pgid": self.pgid,
|
||||||
|
"start_order": self.start_order,
|
||||||
|
"shutdown_order": self.shutdown_order,
|
||||||
|
"stdout_path": self.stdout_path.name,
|
||||||
|
"stderr_path": self.stderr_path.name,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True, slots=True)
|
||||||
|
class ProcessStopResult:
|
||||||
|
stopped_process_ids: tuple[str, ...]
|
||||||
|
residue_process_ids: tuple[str, ...]
|
||||||
|
|
||||||
|
|
||||||
|
class PosixProcessSupervisor:
|
||||||
|
"""Own provider PGIDs and stop them in explicit reverse dependency order."""
|
||||||
|
|
||||||
|
def __init__(
|
||||||
|
self,
|
||||||
|
runtime_root: Path,
|
||||||
|
*,
|
||||||
|
base_environment: dict[str, str] | None = None,
|
||||||
|
) -> None:
|
||||||
|
if os.name != "posix":
|
||||||
|
raise ProcessSupervisorError("S1 process supervisor requires a POSIX worker")
|
||||||
|
self.runtime_root = runtime_root.expanduser().resolve()
|
||||||
|
self.runtime_root.mkdir(mode=0o700, parents=True, exist_ok=True)
|
||||||
|
self.base_environment = dict(base_environment or {})
|
||||||
|
self._lock = threading.RLock()
|
||||||
|
self._run_id: str | None = None
|
||||||
|
self._processes: dict[str, tuple[ProcessSpec, subprocess.Popen[bytes]]] = {}
|
||||||
|
|
||||||
|
def start(self, run_id: str, specs: tuple[ProcessSpec, ...]) -> tuple[OwnedProcess, ...]:
|
||||||
|
if not PROCESS_ID_PATTERN.fullmatch(run_id):
|
||||||
|
raise ProcessSupervisorError("run id must be a safe lowercase identifier")
|
||||||
|
if not specs:
|
||||||
|
raise ProcessSupervisorError("provider graph must not be empty")
|
||||||
|
identifiers = [spec.process_id for spec in specs]
|
||||||
|
if len(identifiers) != len(set(identifiers)):
|
||||||
|
raise ProcessSupervisorError("provider process ids must be unique")
|
||||||
|
start_orders = [spec.start_order for spec in specs]
|
||||||
|
shutdown_orders = [spec.shutdown_order for spec in specs]
|
||||||
|
if len(start_orders) != len(set(start_orders)):
|
||||||
|
raise ProcessSupervisorError("provider start orders must be unique")
|
||||||
|
if len(shutdown_orders) != len(set(shutdown_orders)):
|
||||||
|
raise ProcessSupervisorError("provider shutdown orders must be unique")
|
||||||
|
with self._lock:
|
||||||
|
if self._processes:
|
||||||
|
raise ProcessSupervisorError(
|
||||||
|
f"process authority is already owned by {self._run_id}"
|
||||||
|
)
|
||||||
|
run_root = self.runtime_root / run_id
|
||||||
|
if run_root.exists():
|
||||||
|
raise ProcessSupervisorError("run-scoped process directory already exists")
|
||||||
|
run_root.mkdir(mode=0o700)
|
||||||
|
self._run_id = run_id
|
||||||
|
try:
|
||||||
|
for spec in sorted(specs, key=lambda item: item.start_order):
|
||||||
|
self._start_one(run_root, spec)
|
||||||
|
records = self.snapshot()
|
||||||
|
self._persist_registry(run_root, records, terminal=False)
|
||||||
|
return records
|
||||||
|
except BaseException:
|
||||||
|
self.stop()
|
||||||
|
raise
|
||||||
|
|
||||||
|
def snapshot(self) -> tuple[OwnedProcess, ...]:
|
||||||
|
with self._lock:
|
||||||
|
records: list[OwnedProcess] = []
|
||||||
|
for process_id, (spec, process) in self._processes.items():
|
||||||
|
return_code = process.poll()
|
||||||
|
if return_code is not None:
|
||||||
|
raise ProcessSupervisorError(f"provider {process_id} exited with {return_code}")
|
||||||
|
records.append(
|
||||||
|
OwnedProcess(
|
||||||
|
process_id=process_id,
|
||||||
|
pid=process.pid,
|
||||||
|
pgid=os.getpgid(process.pid),
|
||||||
|
start_order=spec.start_order,
|
||||||
|
shutdown_order=spec.shutdown_order,
|
||||||
|
stdout_path=self.runtime_root
|
||||||
|
/ str(self._run_id)
|
||||||
|
/ f"{process_id}.stdout.log",
|
||||||
|
stderr_path=self.runtime_root
|
||||||
|
/ str(self._run_id)
|
||||||
|
/ f"{process_id}.stderr.log",
|
||||||
|
)
|
||||||
|
)
|
||||||
|
return tuple(sorted(records, key=lambda item: item.start_order))
|
||||||
|
|
||||||
|
def stop(self) -> ProcessStopResult:
|
||||||
|
with self._lock:
|
||||||
|
stopped: list[str] = []
|
||||||
|
residue: list[str] = []
|
||||||
|
remaining: dict[str, tuple[ProcessSpec, subprocess.Popen[bytes]]] = {}
|
||||||
|
run_id = self._run_id
|
||||||
|
for process_id, (spec, process) in sorted(
|
||||||
|
self._processes.items(),
|
||||||
|
key=lambda item: item[1][0].shutdown_order,
|
||||||
|
):
|
||||||
|
pgid = process.pid
|
||||||
|
if self._stop_group(process, pgid, spec):
|
||||||
|
stopped.append(process_id)
|
||||||
|
else:
|
||||||
|
residue.append(process_id)
|
||||||
|
remaining[process_id] = (spec, process)
|
||||||
|
if run_id is not None:
|
||||||
|
run_root = self.runtime_root / run_id
|
||||||
|
records = tuple(
|
||||||
|
OwnedProcess(
|
||||||
|
process_id=process_id,
|
||||||
|
pid=process.pid,
|
||||||
|
pgid=process.pid,
|
||||||
|
start_order=spec.start_order,
|
||||||
|
shutdown_order=spec.shutdown_order,
|
||||||
|
stdout_path=run_root / f"{process_id}.stdout.log",
|
||||||
|
stderr_path=run_root / f"{process_id}.stderr.log",
|
||||||
|
)
|
||||||
|
for process_id, (spec, process) in self._processes.items()
|
||||||
|
)
|
||||||
|
self._persist_registry(
|
||||||
|
run_root,
|
||||||
|
records,
|
||||||
|
terminal=True,
|
||||||
|
residue=tuple(residue),
|
||||||
|
)
|
||||||
|
self._processes = remaining
|
||||||
|
if not remaining:
|
||||||
|
self._run_id = None
|
||||||
|
return ProcessStopResult(tuple(stopped), tuple(residue))
|
||||||
|
|
||||||
|
def residue(self) -> tuple[str, ...]:
|
||||||
|
with self._lock:
|
||||||
|
return tuple(
|
||||||
|
process_id
|
||||||
|
for process_id, (_, process) in self._processes.items()
|
||||||
|
if _group_exists(process.pid)
|
||||||
|
)
|
||||||
|
|
||||||
|
def _start_one(self, run_root: Path, spec: ProcessSpec) -> None:
|
||||||
|
stdout_path = run_root / f"{spec.process_id}.stdout.log"
|
||||||
|
stderr_path = run_root / f"{spec.process_id}.stderr.log"
|
||||||
|
environment = dict(self.base_environment)
|
||||||
|
environment.update(spec.environment)
|
||||||
|
with stdout_path.open("xb") as stdout, stderr_path.open("xb") as stderr:
|
||||||
|
process = subprocess.Popen(
|
||||||
|
spec.argv,
|
||||||
|
cwd=run_root,
|
||||||
|
env=environment,
|
||||||
|
stdin=subprocess.DEVNULL,
|
||||||
|
stdout=stdout,
|
||||||
|
stderr=stderr,
|
||||||
|
start_new_session=True,
|
||||||
|
close_fds=True,
|
||||||
|
)
|
||||||
|
self._processes[spec.process_id] = (spec, process)
|
||||||
|
if spec.startup_grace_seconds:
|
||||||
|
time.sleep(spec.startup_grace_seconds)
|
||||||
|
return_code = process.poll()
|
||||||
|
if return_code is not None:
|
||||||
|
raise ProcessSupervisorError(
|
||||||
|
f"provider {spec.process_id} exited during startup with {return_code}"
|
||||||
|
)
|
||||||
|
if os.getpgid(process.pid) != process.pid:
|
||||||
|
raise ProcessSupervisorError(
|
||||||
|
f"provider {spec.process_id} did not acquire a dedicated process group"
|
||||||
|
)
|
||||||
|
|
||||||
|
def _stop_group(
|
||||||
|
self,
|
||||||
|
process: subprocess.Popen[bytes],
|
||||||
|
pgid: int,
|
||||||
|
spec: ProcessSpec,
|
||||||
|
) -> bool:
|
||||||
|
group_signal_supported = True
|
||||||
|
for sent_signal, timeout in (
|
||||||
|
(signal.SIGINT, spec.interrupt_timeout_seconds),
|
||||||
|
(signal.SIGTERM, spec.terminate_timeout_seconds),
|
||||||
|
(signal.SIGKILL, 1.0),
|
||||||
|
):
|
||||||
|
if process.poll() is not None and not _group_exists(pgid):
|
||||||
|
process.wait(timeout=0)
|
||||||
|
return True
|
||||||
|
try:
|
||||||
|
if group_signal_supported:
|
||||||
|
os.killpg(pgid, sent_signal)
|
||||||
|
else:
|
||||||
|
process.send_signal(sent_signal)
|
||||||
|
except ProcessLookupError:
|
||||||
|
return True
|
||||||
|
except PermissionError:
|
||||||
|
# Some macOS application sandboxes deny killpg for an owned
|
||||||
|
# child session. The reviewed WSL target retains group signals;
|
||||||
|
# local development falls back to the group leader and does
|
||||||
|
# not claim target residue acceptance from this path.
|
||||||
|
group_signal_supported = False
|
||||||
|
process.send_signal(sent_signal)
|
||||||
|
deadline = time.monotonic() + timeout
|
||||||
|
while time.monotonic() < deadline:
|
||||||
|
if process.poll() is not None:
|
||||||
|
process.wait(timeout=0)
|
||||||
|
return not group_signal_supported or not _group_exists(pgid)
|
||||||
|
time.sleep(0.01)
|
||||||
|
return process.poll() is not None and (
|
||||||
|
not group_signal_supported or not _group_exists(pgid)
|
||||||
|
)
|
||||||
|
|
||||||
|
def _persist_registry(
|
||||||
|
self,
|
||||||
|
run_root: Path,
|
||||||
|
records: tuple[OwnedProcess, ...],
|
||||||
|
*,
|
||||||
|
terminal: bool,
|
||||||
|
residue: tuple[str, ...] = (),
|
||||||
|
) -> None:
|
||||||
|
write_json_atomic(
|
||||||
|
run_root / "process-registry.json",
|
||||||
|
{
|
||||||
|
"schema_version": "missioncore.simulation-process-registry/v1",
|
||||||
|
"run_id": run_root.name,
|
||||||
|
"terminal": terminal,
|
||||||
|
"processes": [record.to_dict() for record in records],
|
||||||
|
"residue_process_ids": list(residue),
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _group_exists(pgid: int) -> bool:
|
||||||
|
try:
|
||||||
|
os.killpg(pgid, 0)
|
||||||
|
except ProcessLookupError:
|
||||||
|
return False
|
||||||
|
except PermissionError:
|
||||||
|
return True
|
||||||
|
return True
|
||||||
|
|
@ -257,6 +257,11 @@ class QualificationRunStore:
|
||||||
commands.append(command)
|
commands.append(command)
|
||||||
return tuple(commands)
|
return tuple(commands)
|
||||||
|
|
||||||
|
def list_events(self, run_id: str) -> tuple[QualificationEvent, ...]:
|
||||||
|
path = self._existing_run_path(run_id)
|
||||||
|
with self._lock:
|
||||||
|
return self._read_events(path, run_id)
|
||||||
|
|
||||||
def register_artifact(
|
def register_artifact(
|
||||||
self,
|
self,
|
||||||
run_id: str,
|
run_id: str,
|
||||||
|
|
|
||||||
|
|
@ -0,0 +1,134 @@
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import hashlib
|
||||||
|
from dataclasses import dataclass
|
||||||
|
from pathlib import Path, PurePosixPath
|
||||||
|
from typing import Protocol
|
||||||
|
|
||||||
|
from k1link.simulation.contracts import QualificationRun
|
||||||
|
from k1link.simulation.orchestrator import WorkerStartResult, WorkerStopResult
|
||||||
|
from k1link.simulation.process_supervisor import (
|
||||||
|
PosixProcessSupervisor,
|
||||||
|
ProcessSpec,
|
||||||
|
)
|
||||||
|
from k1link.simulation.s0 import S0Profile, load_s0_profile
|
||||||
|
|
||||||
|
|
||||||
|
class WorkerAdmissionError(RuntimeError):
|
||||||
|
"""The worker no longer matches the accepted S0 generation or D boundary."""
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True, slots=True)
|
||||||
|
class WorkerAdmission:
|
||||||
|
profile: S0Profile
|
||||||
|
profile_sha256: str
|
||||||
|
artifact_run_root: PurePosixPath
|
||||||
|
process_run_root: PurePosixPath
|
||||||
|
|
||||||
|
|
||||||
|
class SimulationWorldControl(Protocol):
|
||||||
|
def pause(self, run_id: str) -> None: ...
|
||||||
|
|
||||||
|
def resume(self, run_id: str) -> None: ...
|
||||||
|
|
||||||
|
def step(self, run_id: str, step_count: int) -> int: ...
|
||||||
|
|
||||||
|
|
||||||
|
class S0WorkerGuard:
|
||||||
|
"""Bind every S1 start to the exact accepted S0 profile and D artifact root."""
|
||||||
|
|
||||||
|
def __init__(self, profile_path: Path) -> None:
|
||||||
|
self.profile_path = profile_path.expanduser().resolve()
|
||||||
|
self.profile = load_s0_profile(self.profile_path)
|
||||||
|
self.profile_sha256 = hashlib.sha256(self.profile_path.read_bytes()).hexdigest()
|
||||||
|
|
||||||
|
def admit(
|
||||||
|
self,
|
||||||
|
run: QualificationRun,
|
||||||
|
artifact_run_root: PurePosixPath,
|
||||||
|
process_run_root: PurePosixPath,
|
||||||
|
) -> WorkerAdmission:
|
||||||
|
if run.host_profile_id != self.profile.profile_id:
|
||||||
|
raise WorkerAdmissionError("run host profile id does not match accepted S0")
|
||||||
|
if run.host_profile_sha256 != self.profile_sha256:
|
||||||
|
raise WorkerAdmissionError("run host profile digest does not match accepted S0")
|
||||||
|
artifact_roots = tuple(
|
||||||
|
item.wsl_path
|
||||||
|
for item in self.profile.storage.mutable_paths
|
||||||
|
if item.identifier == "artifacts"
|
||||||
|
)
|
||||||
|
if len(artifact_roots) != 1:
|
||||||
|
raise WorkerAdmissionError("accepted S0 profile has no unique artifact root")
|
||||||
|
runtime_roots = tuple(
|
||||||
|
item.wsl_path
|
||||||
|
for item in self.profile.storage.mutable_paths
|
||||||
|
if item.identifier == "runtime"
|
||||||
|
)
|
||||||
|
if len(runtime_roots) != 1:
|
||||||
|
raise WorkerAdmissionError("accepted S0 profile has no unique runtime root")
|
||||||
|
expected_artifact_parent = artifact_roots[0] / "s1" / "runs"
|
||||||
|
expected_runtime_parent = runtime_roots[0] / "s1" / "runs"
|
||||||
|
if artifact_run_root != expected_artifact_parent / run.run_id:
|
||||||
|
raise WorkerAdmissionError("S1 run root escapes the accepted D-only layout")
|
||||||
|
if process_run_root != expected_runtime_parent / run.run_id:
|
||||||
|
raise WorkerAdmissionError("S1 process root escapes the accepted D-only layout")
|
||||||
|
if not all(
|
||||||
|
component.qualification_state == "accepted" for component in self.profile.components
|
||||||
|
):
|
||||||
|
raise WorkerAdmissionError("S0 component generation is no longer accepted")
|
||||||
|
return WorkerAdmission(
|
||||||
|
profile=self.profile,
|
||||||
|
profile_sha256=self.profile_sha256,
|
||||||
|
artifact_run_root=artifact_run_root,
|
||||||
|
process_run_root=process_run_root,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
class LocalProcessWorkerAdapter:
|
||||||
|
"""Worker-local S1B adapter; transport and PX4 mapping remain separate ports."""
|
||||||
|
|
||||||
|
def __init__(
|
||||||
|
self,
|
||||||
|
guard: S0WorkerGuard,
|
||||||
|
supervisor: PosixProcessSupervisor,
|
||||||
|
specs: tuple[ProcessSpec, ...],
|
||||||
|
world_control: SimulationWorldControl,
|
||||||
|
) -> None:
|
||||||
|
self.guard = guard
|
||||||
|
self.supervisor = supervisor
|
||||||
|
self.specs = specs
|
||||||
|
self.world_control = world_control
|
||||||
|
|
||||||
|
def start(self, run: QualificationRun) -> WorkerStartResult:
|
||||||
|
artifact_root = next(
|
||||||
|
item.wsl_path
|
||||||
|
for item in self.guard.profile.storage.mutable_paths
|
||||||
|
if item.identifier == "artifacts"
|
||||||
|
)
|
||||||
|
artifact_run_root = artifact_root / "s1" / "runs" / run.run_id
|
||||||
|
process_run_root = PurePosixPath(self.supervisor.runtime_root.as_posix()) / run.run_id
|
||||||
|
admission = self.guard.admit(run, artifact_run_root, process_run_root)
|
||||||
|
records = self.supervisor.start(run.run_id, self.specs)
|
||||||
|
return WorkerStartResult(
|
||||||
|
provider_ids=tuple(record.process_id for record in records),
|
||||||
|
profile_sha256=admission.profile_sha256,
|
||||||
|
)
|
||||||
|
|
||||||
|
def pause(self, run: QualificationRun) -> None:
|
||||||
|
self.world_control.pause(run.run_id)
|
||||||
|
|
||||||
|
def resume(self, run: QualificationRun) -> None:
|
||||||
|
self.world_control.resume(run.run_id)
|
||||||
|
|
||||||
|
def step(self, run: QualificationRun, step_count: int) -> int:
|
||||||
|
return self.world_control.step(run.run_id, step_count)
|
||||||
|
|
||||||
|
def stop(self, run: QualificationRun) -> WorkerStopResult:
|
||||||
|
result = self.supervisor.stop()
|
||||||
|
return WorkerStopResult(
|
||||||
|
stopped_provider_ids=result.stopped_process_ids,
|
||||||
|
residue_provider_ids=result.residue_process_ids,
|
||||||
|
)
|
||||||
|
|
||||||
|
def reconcile(self, run: QualificationRun) -> WorkerStopResult:
|
||||||
|
return self.stop(run)
|
||||||
|
|
@ -0,0 +1,354 @@
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import hashlib
|
||||||
|
import os
|
||||||
|
import sys
|
||||||
|
from pathlib import Path, PurePosixPath
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
from k1link.simulation import (
|
||||||
|
ActiveQualificationRunError,
|
||||||
|
AuthorityProfile,
|
||||||
|
PosixProcessSupervisor,
|
||||||
|
ProcessSpec,
|
||||||
|
ProcessSupervisorError,
|
||||||
|
ProviderPin,
|
||||||
|
QualificationRun,
|
||||||
|
QualificationRunConflictError,
|
||||||
|
QualificationRunStore,
|
||||||
|
ReproducibilityTier,
|
||||||
|
RunKind,
|
||||||
|
RunState,
|
||||||
|
S0WorkerGuard,
|
||||||
|
SimulationApplicationService,
|
||||||
|
WorkerAdmissionError,
|
||||||
|
WorkerStartResult,
|
||||||
|
WorkerStopResult,
|
||||||
|
)
|
||||||
|
|
||||||
|
PROFILE = Path(__file__).resolve().parents[1] / "simulation/s0/qualification-profile.yaml"
|
||||||
|
SHA_A = "a" * 64
|
||||||
|
SHA_B = "b" * 64
|
||||||
|
|
||||||
|
|
||||||
|
def _run(
|
||||||
|
run_id: str = "run-s1b-001",
|
||||||
|
episode_id: str = "episode-s1b-001",
|
||||||
|
*,
|
||||||
|
host_profile_sha256: str | None = None,
|
||||||
|
) -> QualificationRun:
|
||||||
|
return QualificationRun(
|
||||||
|
run_id=run_id,
|
||||||
|
episode_id=episode_id,
|
||||||
|
kind=RunKind.SIMULATION_CLOSED_LOOP,
|
||||||
|
state=RunState.ADMITTED,
|
||||||
|
scenario_generation="stock-rover-v1",
|
||||||
|
scenario_sha256=SHA_A,
|
||||||
|
profile_generation="s1-ackermann-v1",
|
||||||
|
profile_sha256=SHA_B,
|
||||||
|
mission_core_commit="12fba7126d247f809f8d135c5f68922eb08d4683",
|
||||||
|
providers=(
|
||||||
|
ProviderPin("px4-autopilot", "v1.17.0", "v1.17.0"),
|
||||||
|
ProviderPin("gazebo", "harmonic", "8.14.0"),
|
||||||
|
),
|
||||||
|
host_profile_id="ai-worker-px4-gazebo-v1",
|
||||||
|
host_profile_sha256=host_profile_sha256 or ("c" * 64),
|
||||||
|
seed=42,
|
||||||
|
reproducibility_tier=ReproducibilityTier.R1,
|
||||||
|
authority=AuthorityProfile(
|
||||||
|
generation=1,
|
||||||
|
command_ttl_max_ns=250_000_000,
|
||||||
|
heartbeat_timeout_monotonic_ns=500_000_000,
|
||||||
|
),
|
||||||
|
clock_domain="gazebo:/clock",
|
||||||
|
created_at_utc="2026-07-24T18:00:00Z",
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
class _FakeWorker:
|
||||||
|
def __init__(self, profile_sha256: str = "c" * 64) -> None:
|
||||||
|
self.profile_sha256 = profile_sha256
|
||||||
|
self.calls: list[str] = []
|
||||||
|
self.fail_start = False
|
||||||
|
self.residue: tuple[str, ...] = ()
|
||||||
|
|
||||||
|
def start(self, run: QualificationRun) -> WorkerStartResult:
|
||||||
|
self.calls.append(f"start:{run.run_id}")
|
||||||
|
if self.fail_start:
|
||||||
|
raise RuntimeError("synthetic start failure")
|
||||||
|
return WorkerStartResult(("gazebo-server", "px4-sitl"), self.profile_sha256)
|
||||||
|
|
||||||
|
def pause(self, run: QualificationRun) -> None:
|
||||||
|
self.calls.append(f"pause:{run.run_id}")
|
||||||
|
|
||||||
|
def resume(self, run: QualificationRun) -> None:
|
||||||
|
self.calls.append(f"resume:{run.run_id}")
|
||||||
|
|
||||||
|
def step(self, run: QualificationRun, step_count: int) -> int:
|
||||||
|
self.calls.append(f"step:{run.run_id}:{step_count}")
|
||||||
|
return step_count * 2_000_000
|
||||||
|
|
||||||
|
def stop(self, run: QualificationRun) -> WorkerStopResult:
|
||||||
|
self.calls.append(f"stop:{run.run_id}")
|
||||||
|
return WorkerStopResult(("px4-sitl", "gazebo-server"), self.residue)
|
||||||
|
|
||||||
|
def reconcile(self, run: QualificationRun) -> WorkerStopResult:
|
||||||
|
self.calls.append(f"reconcile:{run.run_id}")
|
||||||
|
return WorkerStopResult(("px4-sitl", "gazebo-server"), self.residue)
|
||||||
|
|
||||||
|
|
||||||
|
def _service(tmp_path: Path) -> tuple[SimulationApplicationService, _FakeWorker]:
|
||||||
|
store = QualificationRunStore(tmp_path)
|
||||||
|
store.create(_run())
|
||||||
|
worker = _FakeWorker()
|
||||||
|
return SimulationApplicationService(store, worker), worker
|
||||||
|
|
||||||
|
|
||||||
|
def test_application_service_owns_idempotent_start_pause_step_resume_stop(
|
||||||
|
tmp_path: Path,
|
||||||
|
) -> None:
|
||||||
|
service, worker = _service(tmp_path)
|
||||||
|
running = service.start(
|
||||||
|
"run-s1b-001",
|
||||||
|
idempotency_key="start-001",
|
||||||
|
observed_at_utc="2026-07-24T18:00:01Z",
|
||||||
|
host_monotonic_ns=1,
|
||||||
|
)
|
||||||
|
repeated = service.start(
|
||||||
|
"run-s1b-001",
|
||||||
|
idempotency_key="start-001",
|
||||||
|
observed_at_utc="2026-07-24T18:00:02Z",
|
||||||
|
host_monotonic_ns=2,
|
||||||
|
)
|
||||||
|
paused = service.pause(
|
||||||
|
running.run_id,
|
||||||
|
observed_at_utc="2026-07-24T18:00:03Z",
|
||||||
|
host_monotonic_ns=3,
|
||||||
|
sim_time_ns=2_000_000,
|
||||||
|
)
|
||||||
|
advanced = service.step(
|
||||||
|
running.run_id,
|
||||||
|
step_count=2,
|
||||||
|
observed_at_utc="2026-07-24T18:00:04Z",
|
||||||
|
host_monotonic_ns=4,
|
||||||
|
sim_time_ns=2_000_000,
|
||||||
|
)
|
||||||
|
resumed = service.resume(
|
||||||
|
running.run_id,
|
||||||
|
observed_at_utc="2026-07-24T18:00:05Z",
|
||||||
|
host_monotonic_ns=5,
|
||||||
|
sim_time_ns=6_000_000,
|
||||||
|
)
|
||||||
|
completed = service.stop(
|
||||||
|
running.run_id,
|
||||||
|
idempotency_key="stop-001",
|
||||||
|
observed_at_utc="2026-07-24T18:00:06Z",
|
||||||
|
host_monotonic_ns=6,
|
||||||
|
sim_time_ns=6_000_000,
|
||||||
|
)
|
||||||
|
repeated_stop = service.stop(
|
||||||
|
running.run_id,
|
||||||
|
idempotency_key="stop-001",
|
||||||
|
observed_at_utc="2026-07-24T18:00:07Z",
|
||||||
|
host_monotonic_ns=7,
|
||||||
|
sim_time_ns=6_000_000,
|
||||||
|
)
|
||||||
|
|
||||||
|
assert running.state is RunState.RUNNING
|
||||||
|
assert repeated.state is RunState.RUNNING
|
||||||
|
assert paused.state is RunState.PAUSED
|
||||||
|
assert advanced == 4_000_000
|
||||||
|
assert resumed.state is RunState.RUNNING
|
||||||
|
assert completed.state is RunState.COMPLETED
|
||||||
|
assert repeated_stop == completed
|
||||||
|
assert worker.calls.count("start:run-s1b-001") == 1
|
||||||
|
assert worker.calls.count("stop:run-s1b-001") == 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_application_service_rejects_second_active_owner_and_rebound_keys(
|
||||||
|
tmp_path: Path,
|
||||||
|
) -> None:
|
||||||
|
store = QualificationRunStore(tmp_path)
|
||||||
|
store.create(_run())
|
||||||
|
store.create(_run("run-s1b-002", "episode-s1b-002"))
|
||||||
|
worker = _FakeWorker()
|
||||||
|
service = SimulationApplicationService(store, worker)
|
||||||
|
service.start(
|
||||||
|
"run-s1b-001",
|
||||||
|
idempotency_key="start-001",
|
||||||
|
observed_at_utc="2026-07-24T18:00:01Z",
|
||||||
|
host_monotonic_ns=1,
|
||||||
|
)
|
||||||
|
|
||||||
|
with pytest.raises(ActiveQualificationRunError, match="owned"):
|
||||||
|
service.start(
|
||||||
|
"run-s1b-002",
|
||||||
|
idempotency_key="start-002",
|
||||||
|
observed_at_utc="2026-07-24T18:00:02Z",
|
||||||
|
host_monotonic_ns=2,
|
||||||
|
)
|
||||||
|
with pytest.raises(QualificationRunConflictError, match="another start"):
|
||||||
|
service.start(
|
||||||
|
"run-s1b-001",
|
||||||
|
idempotency_key="different",
|
||||||
|
observed_at_utc="2026-07-24T18:00:02Z",
|
||||||
|
host_monotonic_ns=2,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def test_start_failure_and_profile_drift_fail_terminal(tmp_path: Path) -> None:
|
||||||
|
service, worker = _service(tmp_path)
|
||||||
|
worker.fail_start = True
|
||||||
|
failed = service.start(
|
||||||
|
"run-s1b-001",
|
||||||
|
idempotency_key="start-fail",
|
||||||
|
observed_at_utc="2026-07-24T18:00:01Z",
|
||||||
|
host_monotonic_ns=1,
|
||||||
|
)
|
||||||
|
assert failed.state is RunState.FAILED
|
||||||
|
assert failed.terminal_reason == "orchestrator-start-failed"
|
||||||
|
|
||||||
|
store = QualificationRunStore(tmp_path / "drift")
|
||||||
|
store.create(_run())
|
||||||
|
drift_worker = _FakeWorker("d" * 64)
|
||||||
|
drift_service = SimulationApplicationService(store, drift_worker)
|
||||||
|
drifted = drift_service.start(
|
||||||
|
"run-s1b-001",
|
||||||
|
idempotency_key="start-drift",
|
||||||
|
observed_at_utc="2026-07-24T18:00:01Z",
|
||||||
|
host_monotonic_ns=1,
|
||||||
|
)
|
||||||
|
assert drifted.state is RunState.FAILED
|
||||||
|
assert drifted.terminal_reason == "orchestrator-profile-drift-failed"
|
||||||
|
assert "stop:run-s1b-001" in drift_worker.calls
|
||||||
|
|
||||||
|
|
||||||
|
def test_reset_stops_parent_and_creates_new_identity(tmp_path: Path) -> None:
|
||||||
|
service, _ = _service(tmp_path)
|
||||||
|
service.start(
|
||||||
|
"run-s1b-001",
|
||||||
|
idempotency_key="start-001",
|
||||||
|
observed_at_utc="2026-07-24T18:00:01Z",
|
||||||
|
host_monotonic_ns=1,
|
||||||
|
)
|
||||||
|
reset = service.reset(
|
||||||
|
"run-s1b-001",
|
||||||
|
idempotency_key="reset-001",
|
||||||
|
new_run_id="run-s1b-002",
|
||||||
|
new_episode_id="episode-s1b-002",
|
||||||
|
observed_at_utc="2026-07-24T18:00:02Z",
|
||||||
|
host_monotonic_ns=2,
|
||||||
|
sim_time_ns=2_000_000,
|
||||||
|
)
|
||||||
|
|
||||||
|
assert service.store.load("run-s1b-001").state is RunState.ABORTED
|
||||||
|
assert reset.state is RunState.ADMITTED
|
||||||
|
assert reset.parent_run_id == "run-s1b-001"
|
||||||
|
|
||||||
|
|
||||||
|
def test_restart_reconcile_never_resumes_and_records_residue(tmp_path: Path) -> None:
|
||||||
|
service, worker = _service(tmp_path)
|
||||||
|
service.start(
|
||||||
|
"run-s1b-001",
|
||||||
|
idempotency_key="start-001",
|
||||||
|
observed_at_utc="2026-07-24T18:00:01Z",
|
||||||
|
host_monotonic_ns=1,
|
||||||
|
)
|
||||||
|
worker.residue = ("px4-sitl",)
|
||||||
|
|
||||||
|
reconciled = service.reconcile(
|
||||||
|
observed_at_utc="2026-07-24T18:01:00Z",
|
||||||
|
host_monotonic_ns=60,
|
||||||
|
)
|
||||||
|
|
||||||
|
assert reconciled[0].state is RunState.FAILED
|
||||||
|
assert reconciled[0].terminal_reason == "orchestrator-recovery-interrupted"
|
||||||
|
assert any(
|
||||||
|
event.event_type == "orchestrator.recovery-residue"
|
||||||
|
for event in service.store.list_events("run-s1b-001")
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.skipif(os.name != "posix", reason="target supervisor is POSIX-only")
|
||||||
|
def test_process_supervisor_owns_groups_and_stops_in_declared_order(
|
||||||
|
tmp_path: Path,
|
||||||
|
) -> None:
|
||||||
|
order_path = tmp_path / "order.log"
|
||||||
|
child = (
|
||||||
|
"import os,signal,time;"
|
||||||
|
"p=os.environ['ORDER_FILE'];n=os.environ['NAME'];"
|
||||||
|
"signal.signal(signal.SIGINT,lambda *_:(open(p,'a').write(n+'\\n'),exit(0)));"
|
||||||
|
"time.sleep(60)"
|
||||||
|
)
|
||||||
|
supervisor = PosixProcessSupervisor(tmp_path / "runtime")
|
||||||
|
records = supervisor.start(
|
||||||
|
"run-process-001",
|
||||||
|
(
|
||||||
|
ProcessSpec(
|
||||||
|
process_id="gazebo-server",
|
||||||
|
argv=(sys.executable, "-c", child),
|
||||||
|
start_order=1,
|
||||||
|
shutdown_order=20,
|
||||||
|
environment=(("ORDER_FILE", str(order_path)), ("NAME", "gazebo")),
|
||||||
|
),
|
||||||
|
ProcessSpec(
|
||||||
|
process_id="px4-sitl",
|
||||||
|
argv=(sys.executable, "-c", child),
|
||||||
|
start_order=2,
|
||||||
|
shutdown_order=10,
|
||||||
|
environment=(("ORDER_FILE", str(order_path)), ("NAME", "px4")),
|
||||||
|
),
|
||||||
|
),
|
||||||
|
)
|
||||||
|
|
||||||
|
assert [record.process_id for record in records] == ["gazebo-server", "px4-sitl"]
|
||||||
|
assert all(record.pid == record.pgid for record in records)
|
||||||
|
result = supervisor.stop()
|
||||||
|
|
||||||
|
assert result.residue_process_ids == ()
|
||||||
|
assert result.stopped_process_ids == ("px4-sitl", "gazebo-server")
|
||||||
|
assert order_path.read_text(encoding="utf-8").splitlines() == ["px4", "gazebo"]
|
||||||
|
assert supervisor.residue() == ()
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.skipif(os.name != "posix", reason="target supervisor is POSIX-only")
|
||||||
|
def test_process_supervisor_rolls_back_partial_start(tmp_path: Path) -> None:
|
||||||
|
supervisor = PosixProcessSupervisor(tmp_path / "runtime")
|
||||||
|
with pytest.raises(ProcessSupervisorError, match="exited during startup"):
|
||||||
|
supervisor.start(
|
||||||
|
"run-process-001",
|
||||||
|
(
|
||||||
|
ProcessSpec(
|
||||||
|
process_id="long-lived",
|
||||||
|
argv=(sys.executable, "-c", "import time;time.sleep(60)"),
|
||||||
|
start_order=1,
|
||||||
|
shutdown_order=10,
|
||||||
|
),
|
||||||
|
ProcessSpec(
|
||||||
|
process_id="failed",
|
||||||
|
argv=(sys.executable, "-c", "raise SystemExit(7)"),
|
||||||
|
start_order=2,
|
||||||
|
shutdown_order=20,
|
||||||
|
),
|
||||||
|
),
|
||||||
|
)
|
||||||
|
assert supervisor.residue() == ()
|
||||||
|
|
||||||
|
|
||||||
|
def test_s0_worker_guard_binds_exact_digest_and_d_only_run_path() -> None:
|
||||||
|
guard = S0WorkerGuard(PROFILE)
|
||||||
|
run = _run(host_profile_sha256=hashlib.sha256(PROFILE.read_bytes()).hexdigest())
|
||||||
|
root = PurePosixPath("/mnt/d/NDC_MISSIONCORE/simulation/artifacts/s1/runs/run-s1b-001")
|
||||||
|
process_root = PurePosixPath("/mnt/d/NDC_MISSIONCORE/simulation/runtime/s1/runs/run-s1b-001")
|
||||||
|
|
||||||
|
admission = guard.admit(run, root, process_root)
|
||||||
|
|
||||||
|
assert admission.profile.profile_id == "ai-worker-px4-gazebo-v1"
|
||||||
|
assert admission.profile_sha256 == run.host_profile_sha256
|
||||||
|
with pytest.raises(WorkerAdmissionError, match="escapes"):
|
||||||
|
guard.admit(run, PurePosixPath("/tmp/run-s1b-001"), process_root)
|
||||||
|
with pytest.raises(WorkerAdmissionError, match="process root"):
|
||||||
|
guard.admit(run, root, PurePosixPath("/tmp/run-s1b-001"))
|
||||||
|
with pytest.raises(WorkerAdmissionError, match="digest"):
|
||||||
|
guard.admit(_run(), root, process_root)
|
||||||
Loading…
Reference in New Issue