Smart Agent Teams

Monitoring

The health endpoint, request logs, Cloud Run metrics, what to alert on, and the signals runners and leases give you.

SAT exposes one health endpoint, writes a log line for every request, and runs on Cloud Run, which supplies request and instance metrics. There is no metrics endpoint and no tracing. This page tells you what each signal means and what to alert on.

Health endpoint

GET /health is public and needs no token. It runs SELECT 1 against the database on every call and also reports whether startup migrations succeeded.

ConditionHTTPBody
Database reachable and migrations applied200{"status": "healthy", ...}
Database unreachable503{"status": "unhealthy", ..., "database": "unreachable"}
Database reachable but init_db() failed at startup503{"status": "unhealthy", ..., "database": "initialization failed"}
curl -s -w '\n%{http_code}\n' https://sat.example.com/health
{
  "status": "healthy",
  "service": "SAT API",
  "version": "2.1.0",
  "environment": "production",
  "timestamp": 1791178030.84
}
FieldMeaning
statushealthy or unhealthy
serviceAlways SAT API
versionA constant (2.1.0) in api/main_api.py. It does not change between releases.
environmentThe ENVIRONMENT variable
timestampServer time, Unix seconds
databaseOnly when unhealthy: unreachable or initialization failed

Who uses it:

  • Your deploy pipeline should probe /health on a new revision before it takes traffic, and roll back on failure.
  • The Docker image declares HEALTHCHECK every 30 seconds (curl -fsS http://localhost:8080/health). Cloud Run ignores Docker HEALTHCHECK; it applies only when you run the image with Docker.
  • sat doctor calls it first.

/health is limited to 100 requests per minute per client address. A monitor polling every few seconds is well within that.

A failed init_db() is not retried. The instance keeps answering 503 until it is replaced, by a new deploy or by Cloud Run recycling it. Fix the cause, then deploy again.

Logs

The API writes to stdout and stderr, which Cloud Run sends to Cloud Logging.

Request logs

A middleware in api/main_api.py logs every request through structlog with a JSON payload:

2026-10-05 05:27:10,839 - api.main_api - INFO - {"method": "GET", "url": "https://sat.example.com/api/companies/.../tasks", "status_code": 200, "client_ip": "...", "user_agent": "...", "process_time": 0.0151, "event": "HTTP request processed", "logger": "api.main_api", "level": "info", "timestamp": "2026-10-05T05:27:10.839672Z"}
FieldMeaning
eventHTTP request processed, or HTTP request failed when the handler raised
method, url, status_codeThe request and response. url includes the query string.
client_ip, user_agentAs seen by the API
process_timeSeconds spent in the API. Also returned as the X-Process-Time response header.
level, timestampstructlog level and ISO 8601 UTC time

Each line is the standard logging prefix (time - logger - LEVEL - ) followed by the JSON object. Cloud Logging stores such lines as text (textPayload), not as structured jsonPayload, so use substring or regular-expression filters rather than JSON field filters, and severity is not set from level.

Other log lines use the plain prefix format: authentication events (User logged in successfully: ..., Login failed: ..., Account locked due to failed login attempts: ...), password reset requests, and SMTP warnings.

Useful queries

On Cloud Run, in the Logs Explorer for your project, with resource cloud_run_revision and your API service:

resource.type="cloud_run_revision"
resource.labels.service_name="<service>"
textPayload:"\"status_code\": 5"
FindFilter on textPayload
Server errors"status_code": 5
Unhandled exceptionsUnhandled exception
Startup migration failureDatabase initialization failed
Database outageDatabase health check failed
Account lockoutsAccount locked due to failed login attempts
Rate limiting"status_code": 429
Email not sentSMTP is not configured or Sending email failed
Slow requestsprocess_time values above your threshold (extract with a log-based metric)

Startup writes Starting SAT API and either Database initialized successfully or Database initialization failed, followed by Alembic's Running upgrade 003 -> 004 style lines when migrations apply.

Error responses

Unhandled exceptions are logged with a stack trace as Unhandled exception and answered with 500:

{"error": {"code": "INTERNAL_ERROR", "message": "An unexpected error occurred", "details": {}}}

When ENVIRONMENT is not exactly production, details.error contains the exception text.

Cloud Run metrics

Cloud Run records these for the API service without any code in SAT. Find them in Cloud Monitoring under Cloud Run Revision.

MetricWhy it matters for SAT
Request count by response code class5xx rate is the main error signal. 2% is a reasonable alert and deploy-gate threshold.
Request latenciesNormal requests take milliseconds to a few hundred milliseconds. SSE requests stay open for as long as a client is connected, so exclude /events when you look at latency percentiles.
Container instance countShould be 1. More than one breaks live updates while EVENTS_BACKEND=memory.
Container CPU and memory utilisationOne instance serves every request and holds every SSE connection.
Max concurrent requestsEach connected browser tab and each runner holds one SSE request open.

For the database, watch CPU, memory, disk utilisation and active connections. The API opens a new connection per request (NullPool), so connection count tracks request concurrency.

What to alert on

AlertConditionSeverityFirst step
API downUptime check on https://<your-host>/health fails from 2 or more regions for 2 minutesPageCheck the latest revision and its logs for Database initialization failed or a crash loop
Database unhealthy/health returns 503, or Database health check failed appearsPageCheck the database status and the private network path to it
Error rate5xx above 2% of requests over 5 minutesPageSearch Unhandled exception
Failed deployA rollout fails its /health gate or is rolled backNotifyRead the new revision's startup logs
More than one instanceContainer instance count above 1NotifyRestore --max-instances 1; see Scaling
Database diskDisk utilisation above 80%NotifySee data growth
Lease expiriesRising count of runs failed with lease expired (below)NotifyCheck runner hosts
LockoutsSpike in Account locked due to failed login attemptsNotifyPossible credential stuffing; see Security
Email failingSending email failed in productionNotifyCheck SMTP settings and the app password

Runner and lease signals

The API cannot see runners directly; it sees their claims, heartbeats and reports. These are the signals that work is flowing.

Lease expiry

A runner heartbeats every 30 seconds and reports transcript output every 2 seconds. A running run whose last heartbeat (or start) is older than RUN_LEASE_SECONDS (default 300) is failed by the next POST /runs/claim, POST /runs/expire-stale or sat scheduler sweep, with:

  • runs.error set to Runner <runner_id> stopped reporting (lease expired) (or The process running it stopped reporting (lease expired) for runs with no runner id),
  • the agent set to error if it has no other running run,
  • an audit event with verb failed and actor type system.

Occasional expiries mean a runner host was shut down mid-run. A steady rate means runner hosts are crashing, losing network, or being killed (for example by an out-of-memory killer). Find them in the activity log or:

sat runs list --status failed --limit 100 --json | jq '.[] | select(.error // "" | test("lease expired"))'

Queue depth

Queued runs that nobody claims mean no runner serves those agents.

sat runner status          # runs in flight by runner, and the queue
sat runs list --status queued

A run stays queued when no runner is running for the company, no runner serves the agent's adapter, the agent is paused, awaiting approval or terminated, or the agent is busy with another run. See Troubleshooting.

Budget pauses

When an agent reaches its monthly budget the API pauses it and opens a budget_override approval. A paused agent's runs are not claimed. Watch pending_approvals on the dashboard (GET /dashboard) or the inbox.

sat doctor

sat doctor checks a profile end to end from wherever you run it. It exits 1 if any check fails.

sat doctor --profile mine
✓ profile              mine → https://sat.example.com (config ~/.config/sat/config.json, credentials in keychain)
✓ api                  healthy · production · v2.1.0
✓ auth                 you@example.com via api key (keychain)
✓ company              NoteFlow (NF), role owner
✓ live stream          connected via https://<your Cloud Run service URL>
✓ adapter claude_code  /usr/local/bin/claude 2.0.0 (Claude Code)
– adapter codex        `codex` not found on PATH (npm i -g @openai/codex)
✓ git                  /usr/bin/git
CheckFails when
api/health is unreachable within 10 seconds or not 200
authGET /api/me fails (not signed in, expired or revoked credential)
companyNo current company, or you are not a member
live streamThe SSE stream does not connect within --stream-timeout seconds (default 20). Runners then fall back to polling every 30 seconds.
adapter ..., gitShown as – (not a failure) when the tool is missing

Add --json for machine-readable output.

On this page