Safety supervisor
The SafetySupervisor is the single source of truth for whether it's currently safe to expose the observatory. Nothing else gets to decide. The Sequencer, Scheduler, Scripting (blockscript) engine, and the Response Engine all subscribe to its verdict and act on it.
States
- SAFE — all reporting safety sources agree it's safe and the readings are fresh
- WARNING — the sources currently read safe, but the supervisor is not yet willing to release the system (it is serving out the
UNSAFE → SAFEdwell, see below) - UNSAFE — at least one source reports unsafe, a source has gone stale, or contact with a primary source has been lost
- UNKNOWN — the supervisor has never heard from a required source (no reading yet)
- DISABLED — monitoring is switched off (MANUAL MODE, see below). This state never blocks anything.
UNKNOWN is treated identically to UNSAFE for action purposes — the SafetyVerdict.is_actionable_unsafe property returns true for both. There is no "probably safe" path. WARNING is also not safe: only a clean SAFE verdict releases the system. DISABLED is the one non-blocking non-safe state — it is deliberately not "actionable unsafe", so with monitoring off nothing is gated.
Sources
Today there's one primary source: the AAG CloudWatcher (Solo) driver. More can be added (all-sky in advisory mode, secondary weather station, manual e-stop / lockout). The supervisor subscribes to the Solo driver's events when armed:
solo.reading— a fresh reading arrived; the supervisor records it and re-evaluatessolo.error— the driver reported an error; the supervisor re-evaluates but deliberately does not refresh the "last heard from" timestamp, so repeated errors eventually trip the staleness checksolo.stale— the driver flagged its own reading as stale
The Solo poll runs on its own background thread (default every 30 s, set by [solo].poll_interval_s) and reports six sub-sensors — clouds, wind, rain, humidity, light/daylight, pressure — each of which the Solo itself rates SAFE, WARNING, or UNSAFE. The Solo's headline safe flag is true only when every sub-sensor is safe, and when the supervisor forces UNSAFE off the Solo it names exactly which sensors tripped (e.g. Solo: rain=UNSAFE, Solo: wind=WARNING). A Solo reading is also flagged stale once it is older than 90 s by the Solo's own clock.
Each event triggers a re-evaluation, which recomputes the verdict and, if the state or reasons changed, publishes safety.state_changed with {old, new}. Every evaluation also publishes a safety.tick so other modules (notably the health monitor) can watch the supervisor's heartbeat.
Arbitration rules
The supervisor's design invariants (which the code is explicit about not relaxing):
- Fail-closed. UNKNOWN is treated as UNSAFE. Lost contact with a primary source means UNSAFE, not "no data, keep going."
- Veto, not vote. A primary source declaring UNSAFE forces the system UNSAFE regardless of other sources.
- Asymmetric hysteresis.
SAFE → UNSAFEis instant.UNSAFE → SAFErequires N seconds of continuous SAFE first. - The Solo's own headline verdict is taken as authoritative for its own sensors; the supervisor layers its own meta-rules (staleness, source loss) on top.
_compute_verdict() applies these in order:
- No reading ever? →
UNKNOWN("Solo: no reading yet"). - Contact lost? If it's been more than the configured loss timeout (
primary_loss_timeout_s, default 60 s) since the last reading →UNSAFE(fail closed). - Reading stale? If the Solo's own reading reports
is_stale→UNSAFE. - Solo unsafe? If
reading.overall_safeis false →UNSAFE, carrying the Solo's ownunsafe_reasons. - Solo safe, but coming out of UNSAFE/UNKNOWN? Apply the dwell (see below) →
WARNINGuntil the dwell elapses. - Clean safe →
SAFE.
What the supervisor blocks
The supervisor's job is to gate commands that open the observatory — the roof. It does not block:
- Capturing a test shot
- Slewing the mount
- Moving the focuser
- Changing filters
- Changing cooler temperature
The dome is closed during all of these — equipment isn't exposed. Only roof-opening commands check the supervisor, plus the opt-in unattended runners (the scheduler's wait-for-safe and the blockscript). So a present operator can image regardless of the verdict — the gate exists to keep an unattended rig from opening the roof under bad skies.
The Response Engine — reacting to state changes
The supervisor only decides the state; a separate, opt-in Response Engine is what lets you react to it. Turn it on with [response].enabled = true and give it a list of plans. Each plan has a trigger — safety.safe, safety.warning, safety.unsafe, or safety.unknown — and an ordered list of actions to run when the supervisor transitions into that state. (Plans fire on the actual state change, not on every re-evaluation, so a plan won't spam while the state holds.)
The built-in actions are deliberately safe — they inform, they don't move hardware:
- log — write a line to the log at a chosen level
- notify_console — print a banner (state, source, reasons) to the console
- write_state_file — write the current state and reasons to a small text file (handy for external scripts to poll)
- broadcast — publish a custom event on the bus for other modules to pick up
Equipment-touching responses (park, close, cut power) are not the Response Engine's job — those live in the blockscript's per-state handlers and the scheduler's wait-for-safe roof close, so anything that moves the observatory stays on the audited, opt-in unattended paths.
Manual mode — no observatory, no weather sensor
If you image from the backyard or a portable setup — no roof, no AAG CloudWatcher / Solo — turn the master switch off: Settings → Safety → "Safety monitoring enabled" (uncheck), or [safety].enabled = false. With monitoring off:
- the verdict reads
DISABLED(the UI shows MANUAL MODE, not a red UNSAFE), and it is never "actionable unsafe", so nothing is gated — roof-opening included, if you happen to have a roll-off roof you control yourself; - the weather poll, the loss-escalation, and the response engine don't run, so there's no spurious "no contact" fault;
- you are responsible for the weather and for keeping the rig safe.
Everything else works normally without an enclosure or sensor: Console, the Sequencer, autofocus, plate-solving, guiding, cooling, flat panels / cover calibrators and Auto Flat (a flat device is its own connected accessory, unrelated to the observatory), Sky Atlas and Skyhunter. The only things you simply don't use are the roof/dome actions and the fully-unattended roof-cycling campaign. Leave it on for a real observatory — that's where the gate earns its keep.
The UNSAFE → SAFE dwell (hysteresis)
SAFE → UNSAFE is instant. The reverse is deliberately slow to avoid bouncing the observatory open on a transient cloud break.
When the Solo first reads safe after a period of UNSAFE/UNKNOWN, the supervisor records the moment and reports WARNING, not SAFE. It keeps reporting WARNING — with a reason line counting the dwell elapsed and remaining, e.g. Solo: SAFE — dwell 40s/600s (560s to release) — until SAFE has held continuously for unsafe_to_safe_dwell_s (default 600 s / 10 minutes). Only then does it release to SAFE.
If any non-safe reading arrives during the dwell, the clock is reset and starts over the next time SAFE appears.
Sustained loss → EMERGENCY escalation
A reported-unsafe from the Solo (rain, cloud, wind) auto-recovers: when the sky clears, the dwell runs and the system releases to SAFE on its own. A loss of the weather source is different — never heard from it, no contact past the timeout, or a stale reading — and left alone it would leave an unattended rig flapping the roof under skies nobody can see. Sensor blindness is a higher emergency class than bad weather — a readable "unsafe" is recoverable, but flying blind is not — so a sustained loss latches to EMERGENCY instead of quietly auto-recovering.
[safety].weather_loss_emergency_s is the number of seconds a continuous loss must persist before escalating. It now defaults to 1800 (30 minutes) — the Voyager posture (ExitOnLossForTimeMin=30). Once the primary source has been in a continuous loss condition for that long, the supervisor fires the emergency path once: it hands off to a running blockscript's EMERGENCY handler, or — if no blockscript is active (a plain-sequence or idle rig) — commands a direct roof-close + mount-park itself, and always sends a critical operator alert (critical alerts are never rate-limited or suppressed). It fires only on genuine loss/silence, never on an ordinary rain/cloud UNSAFE, and it latches for that loss episode: the moment contact and a fresh reading return, the timer and the latch reset. The health monitor's Solo-silence watchdog reads this latch and stands down if the supervisor has already escalated, so the two never double-fire.
0 (or unset) means use the default 30 min — a legacy persisted 0 will not silently leave the rig flying blind. To deliberately turn the escalation off, set a negative value (or run MANUAL MODE, which skips all monitoring).
Config knobs
These live in the [safety] section of your config and are applied to the supervisor at startup:
enabled(default true) — the master switch.falseis MANUAL MODE (see above).primary_loss_timeout_s(default 60 s) — how long the supervisor will tolerate silence from the primary source before failing closed to UNSAFE.unsafe_to_safe_dwell_s(default 120 s / 2 minutes) — how long the raw reading stream must be continuously safe before the supervisor releases an UNSAFE/UNKNOWN back to SAFE (asymmetric debounce: suspend is instant, resume is slow, and any single unsafe/stale sample resets the clock). During this window the state is WARNING. Set to0to release immediately (no dwell).weather_loss_emergency_s(default 1800 s / 30 minutes) — how long a continuous primary-source loss must persist before the supervisor latches EMERGENCY instead of auto-recovering (see above).0/unset uses the 30-min default; a negative value disables it.
The Solo poll itself is tuned separately under [solo] — host, poll_interval_s (default 30 s), and timeout_s.
Lifecycle
arm()subscribes to the Solo events and starts reacting. Idempotent — calling it twice is a no-op. Until armed, the supervisor sits at its initialUNKNOWNverdict ("supervisor not yet armed").disarm()unsubscribes and stops reacting; used at shutdown.currentreturns the latestSafetyVerdict(state, reasons, primary source, timestamp) under lock.
Every state change is written to the log with the old/new states and the reasons that triggered it.
Tie-in: the health monitor's safety escalation
The supervisor watches the sky/environment. The separate HealthMonitor (starship.core.health) watches Starship itself — the host, process, threads, and devices. They meet at exactly one point: if the controller becomes unhealthy past the most-severe threshold, that is itself a safety problem, and the health monitor escalates onto the safe-state path.
This matters because the supervisor cannot see, for example, that the controller is about to be killed by the operating system. The health monitor can, and it acts:
- Memory escalation. The health monitor's memory check treats genuine system memory exhaustion (not just a large process) as the real OOM signal. When
psutil.virtual_memory().availablefalls belowsys_mem_critical_pct/sys_mem_critical_mb, the memory check is markedCRITICALandsevere=True— "imminent OOM." A large process RSS over the leak-guard ceiling (mem_critical_mb) is only escalated as severe when the system is also tight; on a roomy box a big-but-stable process is at mostDEGRADEDfor visibility. If the system-memory reading is unavailable, RSS alone is never marked severe, precisely so a false reading can't trip a safety escalation. - What escalation does. A severe-CRITICAL check fires
_escalate_to_safety()once per episode (a latch prevents it re-firing every poll). If a blockscript session is running, it forces that session's EMERGENCY path; otherwise it commands a direct safe-state via the console (roof_closethenmount_park). The trigger is also published toconsole.commandso the reason is visible in the UI event feed next to the resulting roof-close/park entries. - Before escalating, tight memory first triggers cheap reclamation (gc, drop the FITS preview cache) gated by the
[memory]config, so the safe-state path is a last resort, not a first response. - Supervisor heartbeat. The health monitor subscribes to the supervisor's
safety.tickand treats a lost heartbeat asCRITICAL/severe — but only when the Solo is actually enabled and the supervisor is armed (the supervisor is event-driven, so no ticks is normal when Solo is idle). This is the authoritative control-liveness signal; a missing web/UI thread, by contrast, is never severe and never escalates.
The hardwired CloudWatcher interlock and L0 boot-safety remain the real guarantee. The supervisor and the health monitor are the software layers on top: one decides whether it's safe to open, the other makes sure the thing deciding is still trustworthy.