Health monitor
The health monitor watches Starship itself — the software, the host machine, and the process — and keeps it running, always giving priority to safety. It is deliberately distinct from the Safety supervisor:
- Safety supervisor answers "is the sky/environment safe to observe?" (cloud, wind, rain, etc. via Solo/CloudWatcher).
- Health monitor answers "is Starship itself healthy enough to be trusted?" (RAM, disk, threads, hung devices, lost heartbeat).
The two meet at exactly one point: if the controller is unhealthy past the most-severe threshold (imminent out-of-memory, a dead control thread, a lost supervisor heartbeat), that is itself a safety problem — so the health monitor escalates to the safe-state path (close roof + park mount).
It runs as a daemon thread (health-monitor) that polls every poll_interval_s (default 5s). Each tick runs a set of cheap, defensive checks, computes an overall verdict, publishes a health.state event when the level changes, and — if warranted — takes graduated action. Any single check that raises is recorded as UNKNOWN for that check and never crashes the monitor.
Health levels
Each poll produces a verdict, the worst of all the individual checks:
- HEALTHY — everything nominal.
- DEGRADED — something is off but not dangerous: disk getting low, a device poll is stale, the web UI thread is missing, RSS climbing on a roomy machine. The monitor reports it and may self-heal, but it does not touch the hardware.
- CRITICAL — a check crossed a hard line. Whether this escalates depends on whether the check is also marked severe (see below).
- UNKNOWN — a check couldn't be evaluated (e.g.
psutilnot installed so memory can't be read). A single UNKNOWN check never drags the whole verdict to UNKNOWN — the overall level is the worst of the checks that produced a real reading; UNKNOWN is reported only if every check is UNKNOWN. - NUCLEAR — defined in the level ladder but reserved for the safety supervisor's last-resort escalation; the health monitor itself does not emit it.
DEGRADED vs severe — the important distinction
CRITICAL alone does not close the roof. Only checks flagged severe trigger the safe-state escalation. A check is severe only when it represents a genuine danger to the controller or observatory:
- genuine system-memory exhaustion (imminent OOM)
- disk completely full on the images path
- the supervisor heartbeat lost while Solo is active and armed
Things that are merely DEGRADED — a stale device poll, a missing web UI thread, a big-but-stable process on a machine with plenty of free RAM — are reported and may be self-healed, but never escalate to closing the roof. This separation is intentional: it stops noisy-but-harmless conditions from needlessly shutting the observatory.
The checks
Each tick gathers these (defined in starship/core/health/__init__.py):
- memory — system RAM and process RSS (detailed below).
- disk — free space on the FITS
images_dir, plus a fill-rate forecast (detailed below). - cpu — system CPU load, and CPU temperature where a sensor exists (often not on Windows, then load only). A visibility/throttle warning only — high CPU is DEGRADED and is never severe.
- threads — expected worker threads are alive. The only one checked today is the web server thread (
starship-web). Losing the UI is not control-critical (the observatory is governed by the supervisor + blockscript engine, not the UI), so a missing web thread is DEGRADED and never severe. - supervisor_heartbeat — the safety supervisor's
safety.tickis still arriving. - solo_heartbeat — the Solo poll thread is still publishing (detailed below).
- device:<role> — per-connected-driver stale-poll age, one check per connected ASCOM/Alpaca device.
- worker:<transport> — a wedged single-threaded driver worker (detailed below).
Memory: system RAM vs process RSS
This is the most nuanced check, and the one most worth understanding. There are two memory signals, and they play different roles.
System available RAM is the PRIMARY, real OOM signal. What matters for "are we about to be killed?" is how much RAM the whole machine has free, not how big this one process is. A 1.5 GB Starship process on a box with plenty free is fine; the same process on a box with 100 MB free is about to die. The monitor reads psutil.virtual_memory().available (which accounts for reclaimable cache — the right "can we allocate more?" number) and applies:
sys_mem_critical_pct(default 8.0) — below this % of system RAM available -> CRITICAL and severe (real OOM risk).sys_mem_critical_mb(default 512.0) — OR below this many MB available -> CRITICAL and severe (whichever trips first).sys_mem_critical_pct_max_mb(default 2048.0) — v22.2.9: the percentage rule only counts while the amount available is also below this many MB. 8 % of a 32 GB PC is 2.6 GB, which is not an emergency; the absolute floor above is unchanged.sys_mem_warn_pct(default 15.0) — below this % available -> DEGRADED.sys_mem_critical_polls(default 3) — the floor above must hold for this many consecutive polls (15 s at the default poll interval) before it counts as imminent OOM; a single low reading is DEGRADED (v22.2.3).- A low reading taken while a camera frame is downloading never counts at all — the download is the dip. Before v22.2.3 a large sensor's download alone could trip the floor and park the mount mid-run.
Process RSS is a SECONDARY leak guard. The RSS thresholds catch a leaking process, but on their own a large RSS is not an emergency:
mem_warn_mb(default 2200.0) — RSS at/above this -> DEGRADED (leak watch). Set above the normal ~1.5 GB working plateau so stable use still reads HEALTHY.mem_critical_mb(default 3000.0) — the RSS leak-guard ceiling. Crossing it only escalates to CRITICAL/severe if system memory is ALSO tight (available <sys_mem_warn_pct). On a roomy machine, a big-but-stable RSS over the ceiling is reported as DEGRADED ("system has headroom"), not as an emergency.
So a high RSS escalates to the safe-state only in combination with low system memory. This is what stops a false "imminent OOM" on a machine that simply has a large but stable footprint.
Two-tier memory safe-state — hard floor vs soft, dismissible trip. Not every severe memory state should kill an attended test instantly. The monitor distinguishes two cases:
- Hard floor — genuine system-memory exhaustion (available below
sys_mem_critical_pct/sys_mem_critical_mb). This is imminent OOM, never dismissible, and escalates to the safe-state immediately. It must persist forsys_mem_critical_pollspolls and is never counted during a frame download (v22.2.3). - Soft trip — a high RSS (over the leak-guard ceiling) while system memory is only tightish (available below
sys_mem_warn_pctbut above the hard floor). This is severe but dismissible: instead of acting at once it opens a short countdown (mem_soft_countdown_s, default 30s) and publishes ahealth.memory_countdownevent so an attended UI can show it. If the operator dismisses it (dismiss_memory_countdown()), the soft trip is suppressed formem_soft_dismiss_cooldown_s(default 300s). If the countdown elapses undismissed, it escalates. Disable the dismiss path entirely withmem_soft_dismiss_enabled = false, and the soft trip then escalates like the hard floor. The hard floor (and disk / heartbeat / wedged worker) is never dismissible.
Fallback when system memory can't be read: if psutil.virtual_memory() is unavailable, the check falls back to RSS-only thresholds — but it will not mark the result severe on RSS alone, because without the system number it can't confirm a real OOM, and a false severe could needlessly trip the safety escalation.
If psutil is not installed at all, the memory check returns UNKNOWN (pip install psutil to enable it). The other checks still run.
The monitor also keeps a ~600s window of RSS samples to compute a growth_mb_per_min leak-rate indicator, surfaced in the status/diagnostics snapshot (it does not by itself change the verdict).
The stale-config trap — why mem_critical_mb must match the code default
mem_critical_mb was raised to 3000 in a recent build (from an old value of 1400). The older default sat below the normal ~1.5 GB working set of a busy imaging session, which means an out-of-date config.toml that still pins mem_critical_mb = 1400 would trip the RSS leak-guard constantly during completely normal operation.
On its own that old value only produces DEGRADED on a roomy machine — but the moment system memory also gets a little tight, the combination flips to CRITICAL/severe and escalates to close roof + park mount. The result is an observatory that safe-states itself on a normal night because of a stale number in the config file.
The guidance: if you have ever written a full config.toml, delete the mem_critical_mb line (and the other [health] memory lines) so the code default applies, or update it to match the current default (3000). Leaving an old, lower value pinned is the trap. The same logic applies to mem_warn_mb — an old low value just spams DEGRADED, which is annoying rather than dangerous, but still worth fixing.
Disk
_check_disk measures free space on the FITS images path ([fits].images_dir, falling back to its parent or the current directory if the exact path doesn't exist):
disk_warn_mb(default 2000.0) — at/below this -> DEGRADED.disk_critical_mb(default 500.0) — at/below this -> CRITICAL and severe (a full disk means captures fail and the night is lost).
It also keeps a ~30 min window of free-space samples and estimates minutes-to-full at the recent write rate. If the disk is still above disk_warn_mb but is filling fast enough to run out within disk_eta_warn_min (default 90.0, 0 disables the forecast), it reports DEGRADED early — a "filling fast" warning before the night actually dies. The ETA is surfaced in the check detail.
CPU
_check_cpu reads system CPU load (and CPU temperature where a sensor is exposed — frequently unavailable on Windows, in which case only load is reported). This is purely a visibility/throttle warning and is never severe: it never closes the roof.
cpu_warn_pct(default 92.0) — sustained load at/above this -> DEGRADED.cpu_temp_warn_c(default 80.0) — CPU temperature at/above this -> DEGRADED ("warm").cpu_temp_critical_c(default 92.0) — at/above this -> still DEGRADED, flagged "hot (throttling likely)".
Wedged worker
_check_worker_stall catches a hung single-threaded driver worker — most often the ASCOM STA worker hanging on an un-timeout-able COM call. The signature is every connected device on one transport going stale at once: if two or more connected devices on a transport are all stale past stale_device_critical_s, the worker itself looks wedged. That is a control-liveness fault, so worker:<transport> is CRITICAL and severe. Unlike a single stale device (self-healed by reconnect), a wedged worker drives the restart path — the direct safe-state runs through the same wedged worker, so a clean process restart is the real recovery (see the restart level below).
Stale devices
_check_devices walks the connected ASCOM and Alpaca drivers and, for each device that reports a last_reading_age_s, flags a device that is connected but no longer returning fresh readings — a hung device, distinct from a cleanly disconnected one. Drivers that don't expose a reading age are skipped rather than guessed at.
stale_device_warn_s(default 30.0) — poll age at/above this -> DEGRADED.stale_device_critical_s(default 120.0) — at/above this -> still DEGRADED, but flagged "will retry" (these feed the bounded self-heal, below; a stale device is never itself severe).
Supervisor heartbeat
The safety supervisor is event-driven — it only ticks on Solo events, not on a timer. So a "lost heartbeat" is only a real fault when Solo is actually enabled ([solo].host set) and the supervisor is armed. Otherwise there is simply nothing to tick, which is reported as HEALTHY ("supervisor idle").
When Solo is active and armed, the threshold scales with the Solo poll cadence so a deliberately slow poll doesn't read as a fault: it uses max(supervisor_tick_max_age_s, poll_interval_s × 3).
supervisor_tick_max_age_s(default 30.0) — base age beyond which a missing tick -> CRITICAL and severe. A genuinely lost safety heartbeat is a control-liveness failure, so it forces the safe-state.
Solo driver heartbeat (poll-thread liveness watchdog)
Because the supervisor is event-driven, its weather-loss → EMERGENCY timer ([safety].weather_loss_emergency_s) only advances while the Solo driver keeps publishing solo.reading / solo.error events. In the normal comms-loss case the driver thread keeps erroring and publishing, so the timer runs and escalates as designed. But if the Solo poll thread itself dies silently, no events fire at all, the supervisor stops ticking, and that loss timer freezes — a dead driver would otherwise never escalate.
This watchdog closes that gap. It taps the raw Solo event stream directly (not the derived safety.tick), so it catches a dead/wedged poll thread even if the supervisor's own tick chain is the thing that broke. It is the event-stream twin of the wedged-ASCOM-worker detector: on multi-cycle silence it returns CRITICAL and severe, driving the same force_emergency safe-state through the escalate-to-safety path.
Safeguards (it fires only on a genuine thread death, and never twice):
- Armed only after life is seen — it requires Solo to be active (
[solo].hostset), the supervisor armed, and at least one Solo event already observed. A driver that never started is a startup fault (covered by the supervisor heartbeat), not a thread death, so this watchdog stays quiet for it. - No double-fire — if the supervisor's own sustained-loss escalation has already latched EMERGENCY for this episode, the watchdog stands down (reports DEGRADED, non-severe). The two escalations are otherwise naturally exclusive: this one fires only on event silence, while the loss timer needs events to keep flowing.
- Safe, not restart — a dead Solo thread safes the rig (force_emergency); it does not trigger a process restart, because the safe-state path doesn't run through the Solo thread.
solo_heartbeat_enabled(default true) — master switch for the watchdog. Default on: it never fires spuriously (only on multi-cycle silence) and closing the gap is fail-closed.solo_heartbeat_poll_multiple(default 3.0) — silence threshold as a multiple of the Solo poll interval (2–3× is sensible). Threshold =max(solo_heartbeat_min_age_s, poll_interval_s × this).solo_heartbeat_min_age_s(default 0.0) — absolute floor for the threshold (0 = pure multiple).
Graduated authority — what it does about a problem
The monitor has four escalating levels of authority, all configurable and all capped:
- Report — every change of level publishes a
health.stateevent, appends to the in-memory history, logs it, and fires a notifier message (info / warn / critical) if a notifier is wired.
- Cheap memory-pressure response — when the memory check is DEGRADED/CRITICAL, before any escalation, the monitor does best-effort reclamation (gated by
[memory], see below): agc.collect()and/or dropping the FITS preview cache. This never raises into the health loop.
- Bounded self-heal — for clearly-safe faults (stale devices), the monitor reconnects the affected driver. This is capped: at most
heal_attempts_cap(default 3) reconnects per rollingheal_window_s(default 300.0) seconds. Hit the cap and it stops trying and escalates instead, rather than thrashing a reconnect loop. Gated byself_heal_enabled(default true).
- Escalate to safety — for the most-severe class only (a CRITICAL+severe check). This fires the safe-state path:
- If a blockscript session is active, it forces that session's EMERGENCY path (force_emergency()). - Otherwise it commands a direct safe-state through the console: roof_close(confirm_unparked=True) then mount_park(). - It also publishes a console.command event naming which check forced the safe-state, so the trigger is visible in the UI event feed next to the roof-close/park entries rather than looking unexplained. - Latched once per episode: the escalation fires once when the severe condition appears and will not re-fire every poll while it persists (without the latch, the safe-state re-commanded the hardware and spammed the log on every cycle). The latch resets when the severe condition clears. Gated by escalate_to_safety (default true).
- A running sequence is stopped first (v22.2.3): before the roof/park commands the monitor calls the sequencer's safety stop (journal safety.stop, the reason on the run), so the sequence never keeps exposing a mount that was just parked.
- Restart (last resort) — only for a CRITICAL+severe condition that escalation can't fix: imminent OOM, a dead control thread, or a wedged driver worker (the
memory,threads, andworker:<transport>checks). The wedged-worker case is the key one — the direct safe-state runs through the same hung worker, so only a clean process restart actually recovers it. Disabled by default (restart_enabled = false). When enabled, it is heavily bounded:
- max_restarts_per_night (default 2) — hit the cap and it holds for a human instead of looping. - restart_backoff_s (default 30.0) — minimum gap between restarts. - restart_good_health_reset_s (default 1800.0) — stay HEALTHY this long and the restart counter resets. - restart_mode (default "external") — external exits with restart_exit_code (default 42) for a watchdog to relaunch; self spawns a detached child and exits; both prefers external and only self-relaunches if no live watchdog heartbeat is detected. - watchdog_heartbeat_path (default "starship_watchdog.heartbeat") — the stamp an external watchdog touches; recent freshness tells both mode a watchdog is supervising.
The site is hardwired to CloudWatcher and has L0 boot-safety, which is why the restart path can exit immediately without an elaborate safe-state dance — the hardware interlock is the real guarantee.
Configuration
[health] — the monitor and its thresholds
enabled(default true) — master switch. When off, the monitor thread never starts.poll_interval_s(default 5.0) — seconds between checks.- Memory thresholds —
sys_mem_critical_pct(8.0),sys_mem_critical_mb(512.0),sys_mem_warn_pct(15.0),sys_mem_critical_polls(3),mem_warn_mb(2200.0),mem_critical_mb(3000.0). See the memory section above, and heed the stale-config trap. - Memory soft-dismiss —
mem_soft_dismiss_enabled(true),mem_soft_countdown_s(30.0),mem_soft_dismiss_cooldown_s(300.0). See the two-tier memory safe-state above. - Disk —
disk_warn_mb(2000.0),disk_critical_mb(500.0),disk_eta_warn_min(90.0,0disables the fill-rate forecast). - CPU —
cpu_warn_pct(92.0),cpu_temp_warn_c(80.0),cpu_temp_critical_c(92.0). - Stale device —
stale_device_warn_s(30.0),stale_device_critical_s(120.0). These also drive theworker:<transport>wedged-worker detector (all connected devices stale past the critical age). - Supervisor —
supervisor_tick_max_age_s(30.0). - Solo driver heartbeat —
solo_heartbeat_enabled(true),solo_heartbeat_poll_multiple(3.0),solo_heartbeat_min_age_s(0.0). See the Solo driver heartbeat section above. - Self-heal —
self_heal_enabled(true),heal_attempts_cap(3),heal_window_s(300.0). - Escalation —
escalate_to_safety(true). - Restart —
restart_enabled(false),restart_mode("external"),restart_exit_code(42),max_restarts_per_night(2),restart_backoff_s(30.0),restart_good_health_reset_s(1800.0),watchdog_heartbeat_path("starship_watchdog.heartbeat").
[memory] — the behavioural responses
This section holds only the responses to memory pressure; the thresholds deliberately live in [health] so there is one source of truth for the limits.
enabled(default true) — master switch for memory-pressure response.gc_on_warning(default true) — rungc.collect()when the memory check is DEGRADED/CRITICAL. NumPy arrays aren't cyclic so this rarely reclaims much, but it's a cheap safety valve.drop_preview_cache_on_warning(default true) — also drop the FITS preview LRU cache when memory is tight; cheap to rebuild, frees a handful of rendered PNGs and their stats.
Where the data shows up
health.stateevent on the bus on every level change, carrying{level, reason, checks}.console.commandevent when an escalation fires, naming the triggering checks.health.memory_countdownevent during a soft-memory countdown ({seconds_remaining, dismissible, reason}) and on dismiss ({seconds_remaining: 0, dismissed: true}), so an attended UI can show and dismiss the countdown.status()— the structured snapshot behind the health/diagnostics view: current level, every check (name/level/detail/severe/dismissible),restarts_tonight, heal attempts in window, the last 20 history entries, and the memory snapshot (rss_mb,vms_mb,percent,sys_available_mb,sys_available_pct,growth_mb_per_min).- Notifier messages (if a notifier is configured) on each level change and on critical events.
Safety and gotchas
- The hardware interlock is the real guarantee. The escalation's roof-close/park is best-effort command-and-report; the roof is hardwired to CloudWatcher and physically won't open against an unsafe verdict. Don't treat the health monitor as the safety system — it's a backstop that invokes the safe-state when the controller itself is failing.
- Mind the stale-config trap. An old
mem_critical_mb(e.g. 1400) pinned inconfig.tomlwill, in combination with even mildly tight system memory, escalate to safe-state during a normal night. Match the code default (3000) or delete the line. - CRITICAL is not always an escalation. Only CRITICAL+severe checks close the roof. A CRITICAL that isn't severe (it doesn't currently happen for the built-in checks, but the model allows it) is reported, not acted on at the hardware level.
- Restart is off by default and needs a watchdog. With
restart_mode = "external"(the default mode) and no watchdog watching for exit code 42, an external restart just exits the process. Only enablerestart_enabledonce your relaunch mechanism is in place. psutilis required for the memory check. Without it the memory check is UNKNOWN and the leak guard / OOM detection are inactive (the other checks still run).- Self-heal is bounded on purpose. If a device keeps going stale and the heal cap is reached, the monitor stops reconnecting and escalates rather than masking a persistent hardware fault.
How it ties into the rest of Starship
- Safety supervisor / Solo — the health monitor consumes
safety.tickto verify the supervisor is alive, and (only when Solo is active and armed) treats a lost tick as severe. See the Safety and Solo topics. - Blockscript engine — a severe escalation prefers the active blockscript's EMERGENCY path; failing that it commands the console directly.
- Console — the direct safe-state uses
console.roof_close+console.mount_park. See the Console topic. - FITS — the disk check watches
[fits].images_dir, and the memory-pressure response can drop the FITS preview cache. See the FITS topic. - ASCOM / Alpaca — the stale-device check and self-heal reconnect operate on the connected drivers. See the ASCOM and Alpaca topics.
- Diagnostics / Audit —
health.stateandconsole.commandevents flow through the bus into the ops log and diagnostics bundle. See the Audit topic.
Current build: v20.84.