Health monitor

The health monitor watches Starship itself — the software, the host machine, and the process — and keeps it running, always giving priority to safety. It is deliberately distinct from the Safety supervisor:

The two meet at exactly one point: if the controller is unhealthy past the most-severe threshold (imminent out-of-memory, a dead control thread, a lost supervisor heartbeat), that is itself a safety problem — so the health monitor escalates to the safe-state path (close roof + park mount).

It runs as a daemon thread (health-monitor) that polls every poll_interval_s (default 5s). Each tick runs a set of cheap, defensive checks, computes an overall verdict, publishes a health.state event when the level changes, and — if warranted — takes graduated action. Any single check that raises is recorded as UNKNOWN for that check and never crashes the monitor.

Health levels

Each poll produces a verdict, the worst of all the individual checks:

DEGRADED vs severe — the important distinction

CRITICAL alone does not close the roof. Only checks flagged severe trigger the safe-state escalation. A check is severe only when it represents a genuine danger to the controller or observatory:

Things that are merely DEGRADED — a stale device poll, a missing web UI thread, a big-but-stable process on a machine with plenty of free RAM — are reported and may be self-healed, but never escalate to closing the roof. This separation is intentional: it stops noisy-but-harmless conditions from needlessly shutting the observatory.

The checks

Each tick gathers these (defined in starship/core/health/__init__.py):

Memory: system RAM vs process RSS

This is the most nuanced check, and the one most worth understanding. There are two memory signals, and they play different roles.

System available RAM is the PRIMARY, real OOM signal. What matters for "are we about to be killed?" is how much RAM the whole machine has free, not how big this one process is. A 1.5 GB Starship process on a box with plenty free is fine; the same process on a box with 100 MB free is about to die. The monitor reads psutil.virtual_memory().available (which accounts for reclaimable cache — the right "can we allocate more?" number) and applies:

Process RSS is a SECONDARY leak guard. The RSS thresholds catch a leaking process, but on their own a large RSS is not an emergency:

So a high RSS escalates to the safe-state only in combination with low system memory. This is what stops a false "imminent OOM" on a machine that simply has a large but stable footprint.

Two-tier memory safe-state — hard floor vs soft, dismissible trip. Not every severe memory state should kill an attended test instantly. The monitor distinguishes two cases:

Fallback when system memory can't be read: if psutil.virtual_memory() is unavailable, the check falls back to RSS-only thresholds — but it will not mark the result severe on RSS alone, because without the system number it can't confirm a real OOM, and a false severe could needlessly trip the safety escalation.

If psutil is not installed at all, the memory check returns UNKNOWN (pip install psutil to enable it). The other checks still run.

The monitor also keeps a ~600s window of RSS samples to compute a growth_mb_per_min leak-rate indicator, surfaced in the status/diagnostics snapshot (it does not by itself change the verdict).

The stale-config trap — why mem_critical_mb must match the code default

mem_critical_mb was raised to 3000 in a recent build (from an old value of 1400). The older default sat below the normal ~1.5 GB working set of a busy imaging session, which means an out-of-date config.toml that still pins mem_critical_mb = 1400 would trip the RSS leak-guard constantly during completely normal operation.

On its own that old value only produces DEGRADED on a roomy machine — but the moment system memory also gets a little tight, the combination flips to CRITICAL/severe and escalates to close roof + park mount. The result is an observatory that safe-states itself on a normal night because of a stale number in the config file.

The guidance: if you have ever written a full config.toml, delete the mem_critical_mb line (and the other [health] memory lines) so the code default applies, or update it to match the current default (3000). Leaving an old, lower value pinned is the trap. The same logic applies to mem_warn_mb — an old low value just spams DEGRADED, which is annoying rather than dangerous, but still worth fixing.

Disk

_check_disk measures free space on the FITS images path ([fits].images_dir, falling back to its parent or the current directory if the exact path doesn't exist):

It also keeps a ~30 min window of free-space samples and estimates minutes-to-full at the recent write rate. If the disk is still above disk_warn_mb but is filling fast enough to run out within disk_eta_warn_min (default 90.0, 0 disables the forecast), it reports DEGRADED early — a "filling fast" warning before the night actually dies. The ETA is surfaced in the check detail.

CPU

_check_cpu reads system CPU load (and CPU temperature where a sensor is exposed — frequently unavailable on Windows, in which case only load is reported). This is purely a visibility/throttle warning and is never severe: it never closes the roof.

Wedged worker

_check_worker_stall catches a hung single-threaded driver worker — most often the ASCOM STA worker hanging on an un-timeout-able COM call. The signature is every connected device on one transport going stale at once: if two or more connected devices on a transport are all stale past stale_device_critical_s, the worker itself looks wedged. That is a control-liveness fault, so worker:<transport> is CRITICAL and severe. Unlike a single stale device (self-healed by reconnect), a wedged worker drives the restart path — the direct safe-state runs through the same wedged worker, so a clean process restart is the real recovery (see the restart level below).

Stale devices

_check_devices walks the connected ASCOM and Alpaca drivers and, for each device that reports a last_reading_age_s, flags a device that is connected but no longer returning fresh readings — a hung device, distinct from a cleanly disconnected one. Drivers that don't expose a reading age are skipped rather than guessed at.

Supervisor heartbeat

The safety supervisor is event-driven — it only ticks on Solo events, not on a timer. So a "lost heartbeat" is only a real fault when Solo is actually enabled ([solo].host set) and the supervisor is armed. Otherwise there is simply nothing to tick, which is reported as HEALTHY ("supervisor idle").

When Solo is active and armed, the threshold scales with the Solo poll cadence so a deliberately slow poll doesn't read as a fault: it uses max(supervisor_tick_max_age_s, poll_interval_s × 3).

Solo driver heartbeat (poll-thread liveness watchdog)

Because the supervisor is event-driven, its weather-loss → EMERGENCY timer ([safety].weather_loss_emergency_s) only advances while the Solo driver keeps publishing solo.reading / solo.error events. In the normal comms-loss case the driver thread keeps erroring and publishing, so the timer runs and escalates as designed. But if the Solo poll thread itself dies silently, no events fire at all, the supervisor stops ticking, and that loss timer freezes — a dead driver would otherwise never escalate.

This watchdog closes that gap. It taps the raw Solo event stream directly (not the derived safety.tick), so it catches a dead/wedged poll thread even if the supervisor's own tick chain is the thing that broke. It is the event-stream twin of the wedged-ASCOM-worker detector: on multi-cycle silence it returns CRITICAL and severe, driving the same force_emergency safe-state through the escalate-to-safety path.

Safeguards (it fires only on a genuine thread death, and never twice):

Graduated authority — what it does about a problem

The monitor has four escalating levels of authority, all configurable and all capped:

  1. Report — every change of level publishes a health.state event, appends to the in-memory history, logs it, and fires a notifier message (info / warn / critical) if a notifier is wired.
  1. Cheap memory-pressure response — when the memory check is DEGRADED/CRITICAL, before any escalation, the monitor does best-effort reclamation (gated by [memory], see below): a gc.collect() and/or dropping the FITS preview cache. This never raises into the health loop.
  1. Bounded self-heal — for clearly-safe faults (stale devices), the monitor reconnects the affected driver. This is capped: at most heal_attempts_cap (default 3) reconnects per rolling heal_window_s (default 300.0) seconds. Hit the cap and it stops trying and escalates instead, rather than thrashing a reconnect loop. Gated by self_heal_enabled (default true).
  1. Escalate to safety — for the most-severe class only (a CRITICAL+severe check). This fires the safe-state path:

- If a blockscript session is active, it forces that session's EMERGENCY path (force_emergency()). - Otherwise it commands a direct safe-state through the console: roof_close(confirm_unparked=True) then mount_park(). - It also publishes a console.command event naming which check forced the safe-state, so the trigger is visible in the UI event feed next to the roof-close/park entries rather than looking unexplained. - Latched once per episode: the escalation fires once when the severe condition appears and will not re-fire every poll while it persists (without the latch, the safe-state re-commanded the hardware and spammed the log on every cycle). The latch resets when the severe condition clears. Gated by escalate_to_safety (default true).

- A running sequence is stopped first (v22.2.3): before the roof/park commands the monitor calls the sequencer's safety stop (journal safety.stop, the reason on the run), so the sequence never keeps exposing a mount that was just parked.

  1. Restart (last resort) — only for a CRITICAL+severe condition that escalation can't fix: imminent OOM, a dead control thread, or a wedged driver worker (the memory, threads, and worker:<transport> checks). The wedged-worker case is the key one — the direct safe-state runs through the same hung worker, so only a clean process restart actually recovers it. Disabled by default (restart_enabled = false). When enabled, it is heavily bounded:

- max_restarts_per_night (default 2) — hit the cap and it holds for a human instead of looping. - restart_backoff_s (default 30.0) — minimum gap between restarts. - restart_good_health_reset_s (default 1800.0) — stay HEALTHY this long and the restart counter resets. - restart_mode (default "external") — external exits with restart_exit_code (default 42) for a watchdog to relaunch; self spawns a detached child and exits; both prefers external and only self-relaunches if no live watchdog heartbeat is detected. - watchdog_heartbeat_path (default "starship_watchdog.heartbeat") — the stamp an external watchdog touches; recent freshness tells both mode a watchdog is supervising.

The site is hardwired to CloudWatcher and has L0 boot-safety, which is why the restart path can exit immediately without an elaborate safe-state dance — the hardware interlock is the real guarantee.

Configuration

[health] — the monitor and its thresholds

[memory] — the behavioural responses

This section holds only the responses to memory pressure; the thresholds deliberately live in [health] so there is one source of truth for the limits.

Where the data shows up

Safety and gotchas

How it ties into the rest of Starship

Current build: v20.84.

Astroworx Starship v22.2.10 · this page is the in-app help of that buildDownload · HTTP API