Troubleshooting
Known problems with local development, the API, the sat CLI and the runner, and how to fix them.
Problems are grouped by area. Each entry gives the symptom, the cause and the fix. If you hit something that is not here, add it to this page in the same pull request as the fix.
Local development
Symptom. uv venv succeeds, but .venv/bin/python is a dangling symlink, for example to ~/.local/share/uv/python/cpython-3.12-macos-aarch64-none/bin/python3 on a uv install that ships only python and python3.12.
Fix. Create the venv from the concrete interpreter path:
cd apps/api
rm -rf .venv
uv venv -p "$(uv python find 3.12)" .venv
uv pip install -p .venv/bin/python -r requirements-production.txt
.venv/bin/python --version # Python 3.12.xOn Windows the interpreter is .venv\Scripts\python.exe.
Cause. Migration 002_agent_company alters constraints, which SQLite cannot do without Alembic batch mode. The API needs PostgreSQL. Only the pytest suite uses SQLite, and it builds tables directly.
Effect. The API keeps running, but /health answers 503 with "database": "initialization failed".
Fix. Point DATABASE_URL at PostgreSQL 16, either from Docker or a local cluster:
initdb -D ~/.sat-pg -U sat -A trust
pg_ctl -D ~/.sat-pg -o "-p 5433" -l ~/.sat-pg.log start
createdb -h 127.0.0.1 -p 5433 -U sat sat_main
export DATABASE_URL=postgresql://sat@127.0.0.1:5433/sat_mainOn Windows, use Docker Desktop, or the EDB installer and createdb from its bin folder. Confirm the API log shows Running upgrade 001 -> 002 and no Database initialization failed.
Cause. Something else is already listening on 5432, typically another project's PostgreSQL container, or an old SAT volume created with a different password. POSTGRES_PASSWORD only applies when the volume is first created.
Fix. Run SAT's database on another port (see the previous entry, -p 5433), or recreate SAT's volume. Recreating deletes local data:
docker compose -f docker-compose.db.yml down -v
docker compose -f docker-compose.db.yml up -d postgres redisFind the owner of the port with lsof -nP -iTCP:5432 -sTCP:LISTEN (macOS, Linux) or netstat -ano | findstr :5432 (Windows).
Cause. python api/main_api.py uses API_PORT, which defaults to 8000. The Vite proxy, the e2e suite, bootstrap.sh and the CLI's local profile all expect 8080.
Fix. Start the API through its entry point, which defaults to 8080:
cd apps/api && .venv/bin/uvicorn main:app --port 8080
# or, from the repository root, with local defaults:
apps/api/.venv/bin/python start_api_local.pyIf the API must run elsewhere, point the dev server at it: SAT_API_URL=http://localhost:9000 bun run dev.
Cause. The web app holds an SSE connection open, so waitForLoadState('networkidle') never fires.
Fix. In tests and scripts, wait for a visible element, for example page.getByRole('heading', { level: 1 }), instead of networkidle.
Cause. A sandbox blocks Bun's default temp and cache directories.
Fix. Point both at a writable directory:
TMPDIR=$TMPDIR BUN_TMPDIR=$TMPDIR BUN_INSTALL_CACHE_DIR=$TMPDIR/bun-cache bun installCause. @yarlisai/ui is a restricted npm package.
Fix. Run npm login with an account that can read the @yarlisai scope, or write a read-only token to ~/.npmrc:
echo "//registry.npmjs.org/:_authToken=<read-only token>" > ~/.npmrc
bun installCI reads the same kind of read-only token from a CI secret.
API in production
Cause. init_db() failed when this instance started: a migration error, or the database was unreachable at that moment. The instance does not retry.
Fix. Read the startup logs for Database initialization failed and the Alembic error. Fix the migration or the connectivity, then deploy again so a new instance starts. A deploy that gates on /health rolls back a revision in this state. See Database.
Cause. SELECT 1 failed: the database is down, restarting or out of connections, or the private network path to it (for example VPC egress) is broken.
Fix. Check the database instance status and its active connections. Confirm the service still has its VPC egress settings (on Cloud Run, --network, --subnet and --vpc-egress private-ranges-only) and that the database URL secret points at the instance's current private IP.
Cause. JWT_SECRET_KEY is empty and APP_ENV (or, if unset, ENVIRONMENT) is anything other than development, dev, local or test, for example staging, production or prod. The error message names the value it saw.
Fix. Attach the secret, on Cloud Run: gcloud run services update <service> --update-secrets JWT_SECRET_KEY=<jwt-secret>:latest. The API's service account needs roles/secretmanager.secretAccessor on that secret.
Check, in order:
- Events origin. The web build must set
VITE_EVENTS_ORIGINto your Cloud Run service URL. Through the Firebase Hosting rewrite, events are buffered. Checkapps/web/.env.productionor.env.staging. - CORS. The web origin must be in
ALLOWED_ORIGINS, because the stream is cross-origin. - Instances. If more than one API instance is serving, events reach only clients on the instance that handled the write. Restore
--max-instances 1or switch toEVENTS_BACKEND=redis; see Scaling. - Unauthenticated invocation. The browser calls Cloud Run without Google credentials, so the API service must allow unauthenticated invocation.
Cause. SMTP is not configured (SMTP_HOST, SMTP_USER and SMTP_PASSWORD must all be set), or sending fails.
Fix. Search the logs for SMTP is not configured or Sending email failed. Set the variables as described in Configuration. The SMTP password secret must be readable by the API's service account. Also note that a second request within 60 seconds for the same account sends nothing.
Cause. Five wrong passwords in a row lock an account for 30 minutes.
Fix. Wait 30 minutes, or complete a password reset, which clears the lock.
CLI and agent runner
Cause. Completing a password reset ends every API key created before it, as well as every session.
Fix. Sign in again and create a new key (sat apikey create grace-laptop-runner --store), then give it to the runner or script. The reset revoked the old keys, so they no longer appear in sat apikey list and the name is free to reuse.
Cause. On PostgreSQL, refresh_tokens.expires_at is timezone-aware, and older API versions compared it with a naive datetime.utcnow(). The comparison raised, so every refresh failed. The lockout check had the same bug. SQLite returns naive datetimes, so the pytest suite never caught it.
Fix. Fixed in services/auth_service.py, which now compares aware UTC datetimes. The cli-e2e CI job runs the CLI against PostgreSQL with a forced refresh, so this cannot regress silently. To check by hand: sign in, replace the stored access token with garbage, and run sat whoami. It should succeed.
Cause. bun build --compile output has an invalid ad-hoc signature on recent macOS, so the kernel kills it on launch. Linux is not affected.
Fix. Build with bun run --cwd apps/cli build, which signs the binary, or sign it yourself: codesign --force --sign - apps/cli/dist/sat.
Cause. Runners refuse password sessions, because agents use the runner's credential through sat, and a session could mint keys or sign people out.
Fix. Create and store a key once per machine and profile:
sat apikey create laptop-runner --store
sat runner startOr pass one for a single process: SAT_API_KEY=sat_live_... sat runner start.
Check, in order:
- Is a runner running for that company? Run
sat runner status. - Can a runner execute that agent's adapter? By default a runner serves only the adapters installed on its machine (
sat runner adapters).cursor,gemini,opencodeandhttpare not implemented yet. - Is the agent dormant? Paused agents (including budget pauses), agents awaiting a hire approval and terminated agents are skipped. Look in
sat inbox. - Is the agent busy? An agent runs one run at a time. A run stuck in
runningwhose runner died is failed afterRUN_LEASE_SECONDS(default 300) by the next claim or bysat scheduler. You can also callPOST /runs/expire-staleyourself.
Cause. The runner stopped heartbeating for longer than RUN_LEASE_SECONDS: it was killed, its machine slept, or it lost network.
Fix. Keep runner machines awake and connected, run the runner under a process supervisor, and check its log for heartbeat failed. If runs legitimately go quiet for long periods, raise RUN_LEASE_SECONDS on the API. See Monitoring.
Cause. The agent CLI's own sandbox (for example Claude Code's Bash sandbox from ~/.claude/settings.json) blocks network access to the API host, often localhost:8080.
Fix. Pass settings that allow the host: sat runner start --claude-settings '<json or file>', or point the profile at a reachable API URL. The agent then posts BLOCKED: ..., and the runner moves the task to blocked with the agent's question. Nothing is lost.
Cause. A sandboxed shell, such as an agent's sandbox or the Claude Code sandbox, forbids local network access or port binding.
Fix. Allow local networking for that sandbox. In Claude Code, set sandbox.network.allowLocalBinding: true or use /sandbox. Or run the command outside the sandbox.
Cause. The profile's API URL is behind a proxy that buffers responses (Firebase Hosting), and the profile has no events URL. The built-in prod profile has one; the built-in staging profile does not.
Fix. Set the events origin for the profile:
sat profile set <profile> --events-url https://<your Cloud Run service URL>Without a stream, runners still work: they poll for claims every 30 seconds.