Monitoring
The health endpoint, request logs, Cloud Run metrics, what to alert on, and the signals runners and leases give you.
SAT exposes one health endpoint, writes a log line for every request, and runs on Cloud Run, which supplies request and instance metrics. There is no metrics endpoint and no tracing. This page tells you what each signal means and what to alert on.
Health endpoint
GET /health is public and needs no token. It runs SELECT 1 against the database on every call and also reports whether startup migrations succeeded.
| Condition | HTTP | Body |
|---|---|---|
| Database reachable and migrations applied | 200 | {"status": "healthy", ...} |
| Database unreachable | 503 | {"status": "unhealthy", ..., "database": "unreachable"} |
Database reachable but init_db() failed at startup | 503 | {"status": "unhealthy", ..., "database": "initialization failed"} |
curl -s -w '\n%{http_code}\n' https://sat.example.com/health{
"status": "healthy",
"service": "SAT API",
"version": "2.1.0",
"environment": "production",
"timestamp": 1791178030.84
}| Field | Meaning |
|---|---|
status | healthy or unhealthy |
service | Always SAT API |
version | A constant (2.1.0) in api/main_api.py. It does not change between releases. |
environment | The ENVIRONMENT variable |
timestamp | Server time, Unix seconds |
database | Only when unhealthy: unreachable or initialization failed |
Who uses it:
- Your deploy pipeline should probe
/healthon a new revision before it takes traffic, and roll back on failure. - The Docker image declares
HEALTHCHECKevery 30 seconds (curl -fsS http://localhost:8080/health). Cloud Run ignores DockerHEALTHCHECK; it applies only when you run the image with Docker. sat doctorcalls it first.
/health is limited to 100 requests per minute per client address. A monitor polling every few seconds is well within that.
A failed init_db() is not retried. The instance keeps answering 503 until it is replaced, by a new deploy or by Cloud Run recycling it. Fix the cause, then deploy again.
Logs
The API writes to stdout and stderr, which Cloud Run sends to Cloud Logging.
Request logs
A middleware in api/main_api.py logs every request through structlog with a JSON payload:
2026-10-05 05:27:10,839 - api.main_api - INFO - {"method": "GET", "url": "https://sat.example.com/api/companies/.../tasks", "status_code": 200, "client_ip": "...", "user_agent": "...", "process_time": 0.0151, "event": "HTTP request processed", "logger": "api.main_api", "level": "info", "timestamp": "2026-10-05T05:27:10.839672Z"}| Field | Meaning |
|---|---|
event | HTTP request processed, or HTTP request failed when the handler raised |
method, url, status_code | The request and response. url includes the query string. |
client_ip, user_agent | As seen by the API |
process_time | Seconds spent in the API. Also returned as the X-Process-Time response header. |
level, timestamp | structlog level and ISO 8601 UTC time |
Each line is the standard logging prefix (time - logger - LEVEL - ) followed by the JSON object. Cloud Logging stores such lines as text (textPayload), not as structured jsonPayload, so use substring or regular-expression filters rather than JSON field filters, and severity is not set from level.
Other log lines use the plain prefix format: authentication events (User logged in successfully: ..., Login failed: ..., Account locked due to failed login attempts: ...), password reset requests, and SMTP warnings.
Useful queries
On Cloud Run, in the Logs Explorer for your project, with resource cloud_run_revision and your API service:
resource.type="cloud_run_revision"
resource.labels.service_name="<service>"
textPayload:"\"status_code\": 5"| Find | Filter on textPayload |
|---|---|
| Server errors | "status_code": 5 |
| Unhandled exceptions | Unhandled exception |
| Startup migration failure | Database initialization failed |
| Database outage | Database health check failed |
| Account lockouts | Account locked due to failed login attempts |
| Rate limiting | "status_code": 429 |
| Email not sent | SMTP is not configured or Sending email failed |
| Slow requests | process_time values above your threshold (extract with a log-based metric) |
Startup writes Starting SAT API and either Database initialized successfully or Database initialization failed, followed by Alembic's Running upgrade 003 -> 004 style lines when migrations apply.
Error responses
Unhandled exceptions are logged with a stack trace as Unhandled exception and answered with 500:
{"error": {"code": "INTERNAL_ERROR", "message": "An unexpected error occurred", "details": {}}}When ENVIRONMENT is not exactly production, details.error contains the exception text.
Cloud Run metrics
Cloud Run records these for the API service without any code in SAT. Find them in Cloud Monitoring under Cloud Run Revision.
| Metric | Why it matters for SAT |
|---|---|
| Request count by response code class | 5xx rate is the main error signal. 2% is a reasonable alert and deploy-gate threshold. |
| Request latencies | Normal requests take milliseconds to a few hundred milliseconds. SSE requests stay open for as long as a client is connected, so exclude /events when you look at latency percentiles. |
| Container instance count | Should be 1. More than one breaks live updates while EVENTS_BACKEND=memory. |
| Container CPU and memory utilisation | One instance serves every request and holds every SSE connection. |
| Max concurrent requests | Each connected browser tab and each runner holds one SSE request open. |
For the database, watch CPU, memory, disk utilisation and active connections. The API opens a new connection per request (NullPool), so connection count tracks request concurrency.
What to alert on
| Alert | Condition | Severity | First step |
|---|---|---|---|
| API down | Uptime check on https://<your-host>/health fails from 2 or more regions for 2 minutes | Page | Check the latest revision and its logs for Database initialization failed or a crash loop |
| Database unhealthy | /health returns 503, or Database health check failed appears | Page | Check the database status and the private network path to it |
| Error rate | 5xx above 2% of requests over 5 minutes | Page | Search Unhandled exception |
| Failed deploy | A rollout fails its /health gate or is rolled back | Notify | Read the new revision's startup logs |
| More than one instance | Container instance count above 1 | Notify | Restore --max-instances 1; see Scaling |
| Database disk | Disk utilisation above 80% | Notify | See data growth |
| Lease expiries | Rising count of runs failed with lease expired (below) | Notify | Check runner hosts |
| Lockouts | Spike in Account locked due to failed login attempts | Notify | Possible credential stuffing; see Security |
| Email failing | Sending email failed in production | Notify | Check SMTP settings and the app password |
Runner and lease signals
The API cannot see runners directly; it sees their claims, heartbeats and reports. These are the signals that work is flowing.
Lease expiry
A runner heartbeats every 30 seconds and reports transcript output every 2 seconds. A running run whose last heartbeat (or start) is older than RUN_LEASE_SECONDS (default 300) is failed by the next POST /runs/claim, POST /runs/expire-stale or sat scheduler sweep, with:
runs.errorset toRunner <runner_id> stopped reporting (lease expired)(orThe process running it stopped reporting (lease expired)for runs with no runner id),- the agent set to
errorif it has no other running run, - an audit event with verb
failedand actor typesystem.
Occasional expiries mean a runner host was shut down mid-run. A steady rate means runner hosts are crashing, losing network, or being killed (for example by an out-of-memory killer). Find them in the activity log or:
sat runs list --status failed --limit 100 --json | jq '.[] | select(.error // "" | test("lease expired"))'Queue depth
Queued runs that nobody claims mean no runner serves those agents.
sat runner status # runs in flight by runner, and the queue
sat runs list --status queuedA run stays queued when no runner is running for the company, no runner serves the agent's adapter, the agent is paused, awaiting approval or terminated, or the agent is busy with another run. See Troubleshooting.
Budget pauses
When an agent reaches its monthly budget the API pauses it and opens a budget_override approval. A paused agent's runs are not claimed. Watch pending_approvals on the dashboard (GET /dashboard) or the inbox.
sat doctor
sat doctor checks a profile end to end from wherever you run it. It exits 1 if any check fails.
sat doctor --profile mine✓ profile mine → https://sat.example.com (config ~/.config/sat/config.json, credentials in keychain)
✓ api healthy · production · v2.1.0
✓ auth you@example.com via api key (keychain)
✓ company NoteFlow (NF), role owner
✓ live stream connected via https://<your Cloud Run service URL>
✓ adapter claude_code /usr/local/bin/claude 2.0.0 (Claude Code)
– adapter codex `codex` not found on PATH (npm i -g @openai/codex)
✓ git /usr/bin/git| Check | Fails when |
|---|---|
api | /health is unreachable within 10 seconds or not 200 |
auth | GET /api/me fails (not signed in, expired or revoked credential) |
company | No current company, or you are not a member |
live stream | The SSE stream does not connect within --stream-timeout seconds (default 20). Runners then fall back to polling every 30 seconds. |
adapter ..., git | Shown as – (not a failure) when the tool is missing |
Add --json for machine-readable output.