Smart Agent Teams

Scheduler

How sat scheduler fires routines and agent heartbeats on their cron schedules and expires runs whose runner died.

The API stores cron expressions on routines (cron) and agents (heartbeat_cron) but never evaluates them. sat scheduler start does. It runs on any machine with the sat CLI, fires each schedule by calling the API, and sweeps runs whose runner stopped reporting. Without a scheduler, routines run only when someone triggers them and agents wake only when they get work.

What it schedules

SourceConditionOn each tick
Routineenabled is truePOST /routines/{routine_id}/trigger: creates a todo task from the routine, assigned to its agent, and queues a routine run
Agent heartbeatheartbeat_cron is set and the agent is not paused, pending_approval or terminatedPOST /agents/{agent_id}/wake with {"reason": "heartbeat"}: queues a manual run without a task, whose first transcript entry is heartbeat
Stale runsAlwaysPOST /runs/expire-stale every --expire-interval seconds (default 60), and once at start

A heartbeat run has no task, so the agent gets the heartbeat prompt: review open work and company goals, and do the most valuable next step.

If a heartbeat fires while an earlier heartbeat (or any other run for that agent without a task) is still queued, the API returns the queued run instead of adding another. Slow runners do not pile up heartbeats.

Start it

sat scheduler start
sat scheduler start --timezone Europe/Berlin --expire-interval 120

On start it logs each schedule and its next fire time:

scheduler serving NoteFlow
scheduler: 3 schedules active
  routine "Weekly report": next 2d from now
  Grace's heartbeat: next 12m from now
  Ken's heartbeat: next 12m from now
FlagDefaultMeaning
--timezone <tz>The machine's time zoneIANA time zone used to evaluate every cron expression, for example Europe/Berlin or America/New_York
--expire-interval <seconds>60How often to call POST /runs/expire-stale. Values under 15 are raised to 15

The scheduler needs only the sat CLI and a credential with member access to the company. It does not need an API key, but an API key is the practical choice for a long-running process, because a password session ends 7 days after sign-in. It stops on SIGINT or SIGTERM.

Cron syntax

The scheduler uses the croner library. Expressions have 5 fields, or 6 with a leading seconds field:

ExpressionFires
0 9 * * 1Mondays at 09:00
*/30 * * * *Every 30 minutes
0 */4 * * 1-5Every 4 hours on weekdays
30 0 9 * * *Every day at 09:00:30 (6 fields, seconds first)

The API stores cron and heartbeat_cron without validating them. The scheduler skips an invalid expression and logs skipping routine "Weekly report": invalid cron "..." (...); the other schedules keep running.

A job does not overlap itself: if a fire is still in progress when the next tick comes, that tick is skipped.

Keeping schedules current

The scheduler reloads routines and agents:

  • One second after any routine.* or agent.* event on the live stream, debounced.
  • Every 5 minutes, in case the stream was down.

A reload adds new schedules, stops removed ones and replaces any whose cron changed. Pausing an agent stops its heartbeat at the next reload; resuming it starts the heartbeat again.

Before a routine fires, the scheduler fetches it again and skips it if it was disabled in the meantime.

Duplicates and multiple schedulers

Run one scheduler per company. Nothing in the API enforces this.

A second scheduler is mostly harmless for routines. Before triggering, the scheduler reads the routine's last_run_at, and skips the tick if the routine ran within the duplicate window:

duplicate window = min(50 seconds, half the time between the schedule's next two fire times)

For 0 9 * * 1 the window is 50 seconds; for * * * * * * (every second) it is half a second. A routine someone triggered by hand within the window is skipped too, with the log line routine "Weekly report" already ran just now; skipping.

The duplicate window does not apply to heartbeats. Two schedulers both wake the agent; the API merges the second wake into the first run only while that run is still queued. Once a runner has claimed it, the second wake queues another run.

Missed ticks are not backfilled

The scheduler fires only while it runs. Ticks that fall while it is stopped, restarting or unable to reach the API are lost; it does not catch up when it comes back. If a trigger call fails (for example 400 INVALID_STATE because the routine's agent is paused), the scheduler logs it and waits for the next tick.

Stale-run expiry

Every --expire-interval seconds the scheduler calls POST /runs/expire-stale. The API fails each running run whose last heartbeat is older than RUN_LEASE_SECONDS (default 300), frees its agent and enforces its budget. The scheduler logs expired run <id>: its runner stopped reporting for each one.

Runners also trigger expiry, because every claim expires stale runs first. A scheduler matters when no runner is claiming, for example when the only runner crashed. See Leases and expiry.

On this page