Blockscript
The safety-automation scripting engine. A blockscript defines a whole unattended observing session as a state machine: it brings the observatory up, runs the night, reacts to weather and device failures, and tears everything down to a safe state — all without a human present. It is the layer that turns "the sky went unsafe" into "park, close the roof, warm the cooler, notify, and stand by."
A blockscript is built from short, named action lists — one per state. The engine moves between states in response to events (safety verdicts, device health, body completion, manual overrides) and runs the matching action list on entry to each state.
What it is
A blockscript runs as a state machine on its own thread (BlockscriptEngine). Each state has an associated action list that runs once on entry to that state. Only one blockscript runs at a time — the observatory is a single resource, and BlockscriptManager enforces this.
States
- IDLE — nothing running (initial / no engine).
- STARTUP — runs the
startuplist. If all actions succeed, flows to RUNNING; if any required action fails, flows to EMERGENCY. - RUNNING — the session body. Reached after startup and after every successful recovery/resume.
- WARNING — tier 1. A soft problem (safety WARNING/UNKNOWN, or a recoverable device loss). The engine waits and tries to recover; it does not tear down.
- UNSAFE — tier 2. Safety verdict is UNSAFE. The engine pre-empts the session now and puts the observatory in a safe state. Can later RESUME.
- RESUME — transient. Entered only when it is safe again (the Supervisor's own sustained-safe dwell already served), it is dark enough to re-open, and resume is not held. Re-entry is unlimited (bounded by the absolute dawn deadline, not a cap). Immediately flows back to RUNNING (incrementing the resume counter).
- SHUTDOWN — runs the
shutdownlist, then flows to DONE. - EMERGENCY — reachable from any state. Runs the
emergency(safe-state) list, then either auto-returns to RUNNING (if allowed and the safe-state actions succeeded) or halts. - DONE — terminal, clean. The session finished normally.
- HALTED — terminal, safed, awaiting a human. The session hit an unrecoverable condition but the observatory was put in a safe state first.
Happy path: STARTUP -> RUNNING -> SHUTDOWN -> DONE. From RUNNING the engine can divert to WARNING, UNSAFE (then back via RESUME), or EMERGENCY at any time.
How to use it
You author a blockscript object with a name and one action list per state name (startup, running, warning, unsafe, resume, shutdown, emergency). Each list is a list of action dicts. A minimal action is {"type": "<action>", ...params}.
The engine is event-driven. Once started it:
- Enters STARTUP and runs the
startuplist. - On success, enters RUNNING and runs the
runninglist. - Reacts to events posted from the rest of Starship (safety ticks, device-lost/recovered, body-done) and to manual overrides (force emergency / force shutdown / hold resume).
You normally don't drive the engine by hand — BlockscriptManager.start(blockscript) wires it to the safety bus and runs it. The manager exposes force_emergency(), force_shutdown(), hold_resume(held), status(), and stop().
Example block
name: "deep_sky_night"
startup:
- type: unpark
- type: open_roof
- type: cool_to
target_c: -10.0
policy:
retry_count: 2
wait_s: 30
on_error: emergency
- type: run_sequence
sequence: "M51_LRGB"
running:
- type: notify
severity: info
message: "Session running"
warning:
- type: notify
severity: warning
message: "Conditions degraded - holding"
unsafe:
- type: abort_exposure
- type: park
- type: close_roof
resume:
- type: unpark
- type: open_roof
shutdown:
- type: park
- type: close_roof
- type: warm_cooler
- type: notify
severity: info
message: "Session complete"
emergency:
- type: abort_exposure
- type: park
- type: close_roof
- type: warm_cooler
- type: notify
severity: critical
message: "EMERGENCY safe-state engaged"
Action types
Every action type maps to a _do_<type> handler on the runner, which calls a real Console method. Unknown types fail with unknown action '<type>'. An action that raises is caught and reported as a failure — it never crashes the engine.
Mount
- park —
mount_park. Best-effort readback confirms the mount reportsat_park = True; a mismatch is notified, never blocked. - unpark —
mount_unpark.
Roof
- open_roof —
roof_open. The Console refuses to open unless the verdict is SAFE, and the hardware refuses too (the roof is hardwired to CloudWatcher). - close_roof —
roof_close. Behaviour depends onobservatory.roof_motion_requires_safe_mount:
- False (default) — mounts clear the roof in any position, so it closes regardless of park state (confirm_unparked=True). - True — it parks first, then closes. If the park fails it still closes (safety beats optics) but sends a critical notification. - On success it reads back shutter_status == 1 (Closed); on failure it sends a critical notify noting the hardware interlock still applies. The close is never blocked.
Guiding (not wired)
- start_guiding / stop_guiding — PHD2 owns guiding; these are no-ops in Starship. They return success with the note
guiding is managed by PHD2. Present so blockscripts can express intent, but they do nothing yet.
Cooler
- warm_cooler —
cooler_warm_up. Use in shutdown/emergency before power-down. - cool_to —
cooler_ramp_to(target_c). Paramtarget_c(float, default 0.0) is the setpoint in degrees C. Optionalminutes(float) ramps down over that long (e.g. cool to −10 over 5 minutes to spare the sensor from thermal shock); omit it (or leave it blank) to use the cooler's configured default ramp rate.
Capture
- abort_exposure —
capture_abort. Stops the in-progress exposure; use it first in unsafe/emergency lists so the mount/roof can move.
Sequences
- run_sequence — runs a named sequence via
Console.run_sequence. Paramsequence(required; empty name fails). As of v20.9 this is really wired: the Console loads + compiles the named plan from the plan library and starts it as an injected (supervised) run — the sequencer skips observatory-domain phase actions (roof/mount/park) because the blockscript owns them. A missing/unknown plan, a disabled sequencer, or a busy sequencer returns a clear failure (not a silent success), so anon_errorpolicy can retry orgoto. On success the injected run signalsbody_donewhen imaging finishes, moving RUNNING -> SHUTDOWN. See the Sequencer topic.
Scheduler
- run_scheduler — launches the multi-night dispatch Scheduler as the session body. Put it in the running list to turn a blockscript into an unattended campaign runner: the dispatcher images the best eligible target every clear, dark, safe night — swapping targets through your per-target recipes — until every active project's goals are filled, then it signals
body_done(RUNNING -> SHUTDOWN). It is non-blocking and idempotent-ish: if the dispatcher is already running the action reports success rather than erroring. Optional paramschedule(a saved-schedule name) loads that saved schedule into the target pool first, so one auto-starting blockscript can pull up the right campaign after a restart; omit it (or leave it blank) to run whatever projects are already loaded. (If the dispatcher is already running when ascheduleis given, the current pool is kept and the named schedule is ignored — it won't yank targets out from under a live run.) While the blockscript supervises, each scheduled visit runs as an injected sequencer run, so the blockscript keeps the roof/mount/park/safety domain and the scheduler only chooses and images targets. See the Scheduler topic.
Device recovery
- retry_connect — reconnects a device. Param
role(required). Tries theascomthenalpacatransports'connect(role). Used automatically by the engine on device-lost (see below); you can also place it in an action list.
Control flow
- wait — pauses the action list, in one of several modes (param
mode, defaultduration). Every mode is cancellable: it polls a stop/abort event and returns early if the engine is stopping or a manual override is queued.
- duration (default) — sleeps seconds (float, default 0). Back-compat: an action with no mode and just seconds still means a plain sleep. - until_safe — holds until the safety supervisor reports SAFE. With no supervisor configured it returns immediately. - until_time — holds until until. Accepts "HH:MM" (the next local occurrence — so 05:30 waits to tomorrow morning if it's already past) or a full ISO datetime (2026-06-30T22:00). An unparseable value fails the action. - until_dark — holds until astronomical darkness (Sun at/below −18° at the configured site). - until_safe_dark — holds until it is both SAFE and astronomically dark — the natural "open the roof and start imaging" gate. - The gated modes (everything but duration) take an optional timeout_min (float) that caps the wait; on timeout the action fails (so an on_error policy can react). If the site location is unset, the darkness modes can't compute dusk — they skip the darkness half rather than block forever (and still honour the safe half of until_safe_dark).
- notify — sends a notification. Params
severity(defaultinfo) andmessage(default empty). If no notifier is configured it logs instead. Best-effort — always succeeds.
External scripts
- run_script — runs an external command via the shared, sandboxed
[scripts]executor (disabled by default; honoursallowed_dir, a hard timeout, and no-shell). Paramscommand(orscript; required),args,timeout(ortimeout_s),cwd,shell, andfail_ok(treat a non-zero exit as success with a note). The engine's abort event is handed to the executor, so a long script can never delay a manual emergency/shutdown. Escalation on failure is the engine's job — attach apolicyto the action. See the External scripts & plugins topic.
On-fail policy
Every action may carry an optional policy dict. With no policy, behaviour is one attempt; on failure the error is logged and the rest of the list keeps running (the safe-state lists must run to completion — you don't want a failed park to skip the close_roof after it). Policy is opt-in and changes nothing unless present. Fields:
- retry_count (0) — extra attempts after the first before declaring failure.
- wait_s (0.0) — seconds to wait between attempts. The wait is sliced so
stop()and events stay responsive. - notify (false) — on final failure, send a notification describing the failed action.
- on_error (
continue) — what to do after all attempts fail:
- continue (default) — record the failure, keep running the list. - raise / emergency — enter EMERGENCY (both behave identically). - goto — enter the state named in goto (case-insensitive: e.g. SHUTDOWN, unsafe). An unknown target logs a warning and falls through to continue.
- goto — target state name, used only when
on_errorisgoto.
This is a uniform per-action error contract, modelled on ROCK's per-block "On Error" options.
Emergency and force-emergency
EMERGENCY is reachable from any state and is where the observatory is forced to a safe state:
- How it's reached: startup failure; an action policy with
on_error: raise/emergency; a device lost beyond the retry cap; or a manual force_emergency override. - On entry the
emergencyaction list runs (your safe-state sequence: abort, park, close roof, warm cooler, notify). - Then
_do_emergencydecides: ifemergency_auto_returnis enabled, the recovery count is withinmax_recovery_attempts, and the safe-state actions all succeeded, it auto-returns to RUNNING. Otherwise it enters HALTED — safed, awaiting a human. - force_emergency (and force_shutdown) are manual overrides that win in any state.
BlockscriptManager.force_emergency()/force_shutdown()post these.
How safety ties in
BlockscriptManager.start() subscribes to the safety.state_changed bus event and bridges each verdict into the engine via on_safety(state). The engine reacts to SafetyState:
- UNSAFE — from RUNNING/WARNING/RESUME, enter UNSAFE and run the safe-state list. Clears the safe-since timer.
- WARNING or UNKNOWN — from RUNNING, enter WARNING (UNKNOWN fails toward caution, treated as at-least WARNING). Does not relax an existing UNSAFE.
- SAFE — from WARNING, return to RUNNING ("safety cleared"). From UNSAFE, attempt resume.
Resume guards (leaving UNSAFE)
_maybe_resume only promotes UNSAFE -> RESUME when all hold:
- It is safe again. The environmental anti-flap dwell is owned by the Safety Supervisor (asymmetric debounce: instant suspend, but it holds WARNING until the raw reading stream has been continuously clean-safe for its own dwell), so a SAFE verdict reaching the engine already means sustained-safe — the engine no longer imposes a second
resume_dwell_shold. - Resume is not held by an operator (
hold_resume(True)). - It is dark enough to re-open (the roof must not re-open onto a daytime sky; skipped when the site is unset).
Weather re-entry is UNLIMITED — an unstable night keeps recovering as long as it stays dark. The never-clears case is owned by the absolute next-dawn deadline (weather_suspend_dawn_alt_deg), not an attempt cap: if the sky stays UNSAFE until dawn, the engine escalates to a verified teardown (SHUTDOWN -> warm/park/close -> DONE) and notifies. max_resumes is now advisory/inert (kept for back-compat).
The engine also re-checks resume on its idle tick (every ~0.5 s) — and checks the dawn deadline there too — so neither depends on a fresh safety event arriving.
Device health
on_device_lost(role) / on_device_recovered(role) feed device health in. As of v20.9 the manager wires these for you: it subscribes to the equipment.device bus event and bridges each snapshot into the engine. A loss (connected False with an error from a failed poll) escalates; a deliberate disconnect (connected False, no error) is ignored so you can take a device offline without tripping EMERGENCY; healthy polls are de-duplicated so a steady device never spams recovery. On a loss the engine increments a per-role retry counter:
- Within
device_loss_retry_cap(default 3) — enter WARNING and fire aretry_connectfor that role. If it reconnects, clear the counter and return to RUNNING. - Beyond the cap — enter EMERGENCY (
device unrecoverable).
body_done (posted via signal_body_done()) moves RUNNING -> SHUTDOWN ("session body complete"). A supervised run_sequence posts this automatically when imaging finishes; if it arrives while transiently in WARNING/UNSAFE/RESUME it is latched and applied on the return to RUNNING, so a finished session never hangs past dawn.
The safety model for this site
These hardware facts shape why the engine is "command-and-report" rather than a hard gate:
- The roof is hardwired to CloudWatcher. It physically cannot open against an unsafe verdict. Software here commands and reports; the hardware interlock is the real guarantee.
- Mounts clear the roof in any position, so closing the roof does not require parking first — unless
observatory.roof_motion_requires_safe_mountis True. - Actuation readback is best-effort. Close/park read back the driver's reported state where available and surface a mismatch (notify / EMERGENCY), but the engine never blocks the close. Better to close against an unconfirmed readback than to leave the roof open.
Config knobs
On config.blockscript:
- enabled (False) — whether the L4 blockscript feature is on at all.
- device_loss_retry_cap (3) — device-lost retries (WARNING + retry_connect) before EMERGENCY.
- auto_start (empty) — name of a blockscript to relaunch automatically when Starship starts (set from the ⟳ Auto-start on boot selector on the Blockscripts page). On startup the service waits on a background thread for a first safety verdict, then starts that blockscript — so an interrupted unattended campaign resumes (the scheduler ledger persists across the restart). Note: this only fires once Starship's process is running; on a fresh power-up something at the OS level (a Windows startup task/service) still has to launch Starship first.
On the blockscript object itself:
- resume_dwell_s (120.0) — inert since the Safety Supervisor now owns the environmental resume dwell (kept for back-compat).
- max_resumes (3) — advisory/inert. Weather resume is UNLIMITED, bounded by the absolute dawn deadline, not this cap.
- weather_suspend_dawn_alt_deg (-6.0) — Sun altitude (civil dawn) at which a never-clearing weather suspend ends the session with a verified teardown + notify.
- startup_wait_for_safe (True) / startup_wait_safe_timeout_min (720.0) — events-masked boot: when the startup list opens the roof, wait for safe darkness first (a storm-time reboot holds the roof shut until safe instead of failing into EMERGENCY), capped by the timeout.
- max_recovery_attempts (2) — EMERGENCY auto-return budget.
- emergency_auto_return (False) — whether EMERGENCY may auto-return to RUNNING at all. If off, EMERGENCY always halts.
- resume_scope (
restart_step) — what a resumed session picks up at. Today onlyrestart_stepis live (checkpoint-resume is deferred).
On config.observatory:
- roof_motion_requires_safe_mount (False) — if True,
close_roofparks first (and still closes even if the park fails).
Safety and gotchas
- Safe-state lists must run to completion. Without a policy, a failed action is logged and the list continues — by design, so a failed
parknever prevents the followingclose_roof. Order safe-state lists so the critical close/warm steps come after the ones that merely help. on_error: raiseandemergencyare equivalent — both go to EMERGENCY. There is no "abort the list but stay" option; usegotofor a different destination state.- Guiding actions are no-ops.
start_guiding/stop_guidingdo nothing today (PHD2 owns guiding) — they report success but don't change anything. - run_sequence runs under supervision. A blockscript's
run_sequencestarts the plan as an injected run, so the sequencer leaves the roof/mount/safety to the blockscript and just images. A bad plan name, a disabled sequencer, or a sequencer already running fails the action cleanly — give it anon_errorpolicy if you don't want a failed start to be merely logged. Requires[sequencer].enabledand a plan in the library. - Readback never blocks. A
park/close_roofmismatch is only surfaced (notify/log). Trust the hardware interlock, and treat readback as a heads-up, not a gate. - Only one blockscript at a time.
start()refuses if one is already active (any state except DONE/HALTED/IDLE). - Terminal states need a human. HALTED means safed-but-stuck; the engine will not restart itself. DONE is the clean finish.
How it ties into the rest of Starship
- Safety supervisor — the manager bridges
safety.state_changedverdicts into the engine; the engine's WARNING/UNSAFE/RESUME logic is driven entirely bySafetyState. See the Safety topic. - Console — every action calls a Console method (
mount_park,roof_close,cooler_ramp_to,capture_abort, ...). The engine is the policy; the Console is the hardware bridge. - Sequencer —
run_sequenceinvokes a named sequence as an injected (supervised) run, so the sequencer skips its own roof/mount/park phase actions and leaves them to the blockscript. See the Sequencer topic. - Scheduler — when a blockscript is active the dispatch scheduler runs each chunk as an injected run too (the blockscript owns the observatory; the scheduler only chooses + images targets). Only a
run_sequencebody signalsbody_done; the scheduler's many chunks do not, so a many-target night runs to completion instead of shutting down after the first chunk. See the Scheduler topic. - Bus / Monitor — the engine publishes
blockscript.state({name, state, reason}) andblockscript.action({action, ok, error, attempt}) so the Monitor and audit trail can track the session. A 20-entry state history is kept forstatus(). - Notifications —
notifyactions and failure/readback notices go through the configured NotificationHandler.
What's deferred
- Guiding control (
start_guiding/stop_guiding) is not wired — PHD2 owns guiding. - Checkpoint resume. A blockscript carries a
resume_scope(defaultrestart_step); resume-from-checkpoint is a separate roadmap item, so a resumed session restarts the current step rather than picking up mid-frame.
See the Roadmap topic for the broader list.