Safety supervisor

The SafetySupervisor is the single source of truth for whether it's currently safe to expose the observatory. Nothing else gets to decide. The Sequencer, Scheduler, Scripting (blockscript) engine, and the Response Engine all subscribe to its verdict and act on it.

States

UNKNOWN is treated identically to UNSAFE for action purposes — the SafetyVerdict.is_actionable_unsafe property returns true for both. There is no "probably safe" path. WARNING is also not safe: only a clean SAFE verdict releases the system. DISABLED is the one non-blocking non-safe state — it is deliberately not "actionable unsafe", so with monitoring off nothing is gated.

Sources

Today there's one primary source: the AAG CloudWatcher (Solo) driver. More can be added (all-sky in advisory mode, secondary weather station, manual e-stop / lockout). The supervisor subscribes to the Solo driver's events when armed:

The Solo poll runs on its own background thread (default every 30 s, set by [solo].poll_interval_s) and reports six sub-sensors — clouds, wind, rain, humidity, light/daylight, pressure — each of which the Solo itself rates SAFE, WARNING, or UNSAFE. The Solo's headline safe flag is true only when every sub-sensor is safe, and when the supervisor forces UNSAFE off the Solo it names exactly which sensors tripped (e.g. Solo: rain=UNSAFE, Solo: wind=WARNING). A Solo reading is also flagged stale once it is older than 90 s by the Solo's own clock.

Each event triggers a re-evaluation, which recomputes the verdict and, if the state or reasons changed, publishes safety.state_changed with {old, new}. Every evaluation also publishes a safety.tick so other modules (notably the health monitor) can watch the supervisor's heartbeat.

Arbitration rules

The supervisor's design invariants (which the code is explicit about not relaxing):

  1. Fail-closed. UNKNOWN is treated as UNSAFE. Lost contact with a primary source means UNSAFE, not "no data, keep going."
  2. Veto, not vote. A primary source declaring UNSAFE forces the system UNSAFE regardless of other sources.
  3. Asymmetric hysteresis. SAFE → UNSAFE is instant. UNSAFE → SAFE requires N seconds of continuous SAFE first.
  4. The Solo's own headline verdict is taken as authoritative for its own sensors; the supervisor layers its own meta-rules (staleness, source loss) on top.

_compute_verdict() applies these in order:

  1. No reading ever? → UNKNOWN ("Solo: no reading yet").
  2. Contact lost? If it's been more than the configured loss timeout (primary_loss_timeout_s, default 60 s) since the last reading → UNSAFE (fail closed).
  3. Reading stale? If the Solo's own reading reports is_stale → UNSAFE.
  4. Solo unsafe? If reading.overall_safe is false → UNSAFE, carrying the Solo's own unsafe_reasons.
  5. Solo safe, but coming out of UNSAFE/UNKNOWN? Apply the dwell (see below) → WARNING until the dwell elapses.
  6. Clean safe → SAFE.

What the supervisor blocks

The supervisor's job is to gate commands that open the observatory — the roof. It does not block:

The dome is closed during all of these — equipment isn't exposed. Only roof-opening commands check the supervisor, plus the opt-in unattended runners (the scheduler's wait-for-safe and the blockscript). So a present operator can image regardless of the verdict — the gate exists to keep an unattended rig from opening the roof under bad skies.

The Response Engine — reacting to state changes

The supervisor only decides the state; a separate, opt-in Response Engine is what lets you react to it. Turn it on with [response].enabled = true and give it a list of plans. Each plan has a trigger — safety.safe, safety.warning, safety.unsafe, or safety.unknown — and an ordered list of actions to run when the supervisor transitions into that state. (Plans fire on the actual state change, not on every re-evaluation, so a plan won't spam while the state holds.)

The built-in actions are deliberately safe — they inform, they don't move hardware:

Equipment-touching responses (park, close, cut power) are not the Response Engine's job — those live in the blockscript's per-state handlers and the scheduler's wait-for-safe roof close, so anything that moves the observatory stays on the audited, opt-in unattended paths.

Manual mode — no observatory, no weather sensor

If you image from the backyard or a portable setup — no roof, no AAG CloudWatcher / Solo — turn the master switch off: Settings → Safety → "Safety monitoring enabled" (uncheck), or [safety].enabled = false. With monitoring off:

Everything else works normally without an enclosure or sensor: Console, the Sequencer, autofocus, plate-solving, guiding, cooling, flat panels / cover calibrators and Auto Flat (a flat device is its own connected accessory, unrelated to the observatory), Sky Atlas and Skyhunter. The only things you simply don't use are the roof/dome actions and the fully-unattended roof-cycling campaign. Leave it on for a real observatory — that's where the gate earns its keep.

The UNSAFE → SAFE dwell (hysteresis)

SAFE → UNSAFE is instant. The reverse is deliberately slow to avoid bouncing the observatory open on a transient cloud break.

When the Solo first reads safe after a period of UNSAFE/UNKNOWN, the supervisor records the moment and reports WARNING, not SAFE. It keeps reporting WARNING — with a reason line counting the dwell elapsed and remaining, e.g. Solo: SAFE — dwell 40s/600s (560s to release) — until SAFE has held continuously for unsafe_to_safe_dwell_s (default 600 s / 10 minutes). Only then does it release to SAFE.

If any non-safe reading arrives during the dwell, the clock is reset and starts over the next time SAFE appears.

Sustained loss → EMERGENCY escalation

A reported-unsafe from the Solo (rain, cloud, wind) auto-recovers: when the sky clears, the dwell runs and the system releases to SAFE on its own. A loss of the weather source is different — never heard from it, no contact past the timeout, or a stale reading — and left alone it would leave an unattended rig flapping the roof under skies nobody can see. Sensor blindness is a higher emergency class than bad weather — a readable "unsafe" is recoverable, but flying blind is not — so a sustained loss latches to EMERGENCY instead of quietly auto-recovering.

[safety].weather_loss_emergency_s is the number of seconds a continuous loss must persist before escalating. It now defaults to 1800 (30 minutes) — the Voyager posture (ExitOnLossForTimeMin=30). Once the primary source has been in a continuous loss condition for that long, the supervisor fires the emergency path once: it hands off to a running blockscript's EMERGENCY handler, or — if no blockscript is active (a plain-sequence or idle rig) — commands a direct roof-close + mount-park itself, and always sends a critical operator alert (critical alerts are never rate-limited or suppressed). It fires only on genuine loss/silence, never on an ordinary rain/cloud UNSAFE, and it latches for that loss episode: the moment contact and a fresh reading return, the timer and the latch reset. The health monitor's Solo-silence watchdog reads this latch and stands down if the supervisor has already escalated, so the two never double-fire.

0 (or unset) means use the default 30 min — a legacy persisted 0 will not silently leave the rig flying blind. To deliberately turn the escalation off, set a negative value (or run MANUAL MODE, which skips all monitoring).

Config knobs

These live in the [safety] section of your config and are applied to the supervisor at startup:

The Solo poll itself is tuned separately under [solo] — host, poll_interval_s (default 30 s), and timeout_s.

Lifecycle

Every state change is written to the log with the old/new states and the reasons that triggered it.

Tie-in: the health monitor's safety escalation

The supervisor watches the sky/environment. The separate HealthMonitor (starship.core.health) watches Starship itself — the host, process, threads, and devices. They meet at exactly one point: if the controller becomes unhealthy past the most-severe threshold, that is itself a safety problem, and the health monitor escalates onto the safe-state path.

This matters because the supervisor cannot see, for example, that the controller is about to be killed by the operating system. The health monitor can, and it acts:

The hardwired CloudWatcher interlock and L0 boot-safety remain the real guarantee. The supervisor and the health monitor are the software layers on top: one decides whether it's safe to open, the other makes sure the thing deciding is still trustworthy.

Astroworx Starship v22.2.10 · this page is the in-app help of that buildDownload · HTTP API