Smart Agent Teams

Architecture

How SAT is built, from system context down to tables, and the design decisions behind it.

SAT is a control plane for an AI agent company. The API is the system of record: it stores companies, agents, goals, tasks, runs, approvals, routines and the audit log, and it enforces budgets. It does not run agents and never calls a model. Starting work only queues a run. A runner, such as sat runner from apps/cli or any client that follows the runner protocol, claims queued runs, executes them with an agent CLI, and reports back.

The diagrams follow the C4 model: context, containers, components, deployment. They are drawn as flowcharts with C4-style boxes so they render everywhere.

System context

Who uses SAT and what it talks to.

Containers

The deployable pieces. The runner is part of the sat CLI and runs on your machines, not in Google Cloud.

ContainerCodeRuns where
Web control planeapps/webStatic files on Firebase Hosting
SAT APIapps/api (entry point main:app)A Cloud Run service (or any container host), one instance
DatabaseAlembic migrations in apps/api/alembicPostgreSQL 16 (for example Cloud SQL), private IP
Event busapps/api/company/events.pyIn the API process by default; Redis when EVENTS_BACKEND=redis
CLI, runner, schedulerapps/cliOperator laptops, CI, or any machine with the agent CLIs installed

Components of the API

The agent-company domain lives in apps/api/company. Authentication lives in apps/api/api/auth_endpoints.py, services/auth_service.py and services/api_keys.py.

ComponentFileResponsibility
Routerscompany/routers/{companies,agents,tasks,runs,insights}.pyHTTP endpoints. Mounted in this order by company/routers/__init__.py. Operation ids are the Python function names.
Dependenciescompany/deps.pyResolves the caller's membership (CompanyContext), require_admin, get_owned (loads a row only if it belongs to the company), error helpers.
Servicescompany/services.pyqueue_run (reuses a queued run for the same agent, task and routine, whatever the trigger), next_task_number (row lock on the company), enforce_agent_budget, finish_run, expire_stale_runs, run_actor.
Auditcompany/audit.pyrecord adds an AuditEvent to the session; publish sends the SSE event after commit; diff computes {field: [old, new]}.
Eventscompany/events.pyInMemoryBus (default) or RedisBus, chosen by EVENTS_BACKEND at import time.
Schemascompany/schemas.pyPydantic request and response models and the status literals.
Modelsdatabase/models.py, database/company_models.pySQLAlchemy tables.

Deployment

SAT ships as two deployable parts: a static web build (apps/web/dist) and an API container image (apps/api/Dockerfile). Anything that serves static files and runs one container next to a PostgreSQL 16 database can host it. The reference shape, which this documentation assumes where it names Google Cloud services, is Firebase Hosting for the web app, Cloud Run for the API and Cloud SQL for PostgreSQL for the database.

To run it yourself:

  • Database. Create a PostgreSQL 16 database and a user for the API. The API applies its own migrations at startup. See Database.
  • API. Deploy the image to Cloud Run (or any container host) with DATABASE_URL, JWT_SECRET_KEY, APP_ENV, ENVIRONMENT, ALLOWED_ORIGINS and WEB_APP_URL set, secrets coming from your secret store, and at most one instance while EVENTS_BACKEND=memory. See Configuration.
  • Web app. Build apps/web with VITE_EVENTS_ORIGIN set to the API's direct URL and serve apps/web/dist, rewriting /api/** and /health to the API.

Live updates bypass Firebase Hosting

Firebase Hosting buffers responses that it rewrites to Cloud Run, so server-sent events would arrive late or never. Set VITE_EVENTS_ORIGIN in the web build to the API's Cloud Run URL, and the browser opens the event stream there directly. Every other request goes through the Hosting rewrite on the same origin. The browser must therefore reach the Cloud Run service without Google credentials, so the service must allow unauthenticated invocation. The API's own authentication still applies to every route except the public ones: /, /health, registration, sign-in, token refresh and password reset.

The hosted Yarlis environments are deployed by the Yarlis release system; Yarlis staff: see the internal SAT runbook on the staff portal (portal.yarlis.com/docs).

Data model

All tables live in one PostgreSQL database. Every company-scoped table has a company_id foreign key, and every query filters on it. Money columns are integer cents; a budget of 0 means no limit. Primary keys are UUIDs.

Notes on the model:

  • Unique constraints. company_members (company_id, user_id) and tasks (company_id, number). A task key such as NF-12 is companies.task_prefix plus tasks.number; it is computed, not stored.
  • Shared tables. agents and projects predate the company model. Their older columns (agents.type, agents.session_id, projects.tech_stack, and others) are still present and agents.type is still NOT NULL, so the API copies role into it. company_id is nullable on both tables for the same reason.
  • Approvals also reference requested_by_agent_id, requested_by_user_id and decided_by_user_id.
  • Runs carry transcript as a JSON array of {ts, role, text} entries in the same row. Long runs grow that row; see Limits.
  • Legacy tables created by migration 001 and not used by the company domain: sessions, session_logs, agent_tasks, deployments, system_metrics, system_health_checks.

See Database for migrations.

Request flow

A typical write, assigning a task to Grace in the NoteFlow company:

  1. Authentication. get_current_user reads the Bearer token. A token starting with sat_live_ or sat_test_ is looked up as an API key by its SHA-256; anything else is decoded as an HS256 JWT.
  2. Authorization. get_company_context loads the caller's company_members row for the company in the path. No row means 404 NOT_FOUND.
  3. Write and audit. The router changes rows and calls audit.record, which adds an audit_events row in the same transaction.
  4. Commit, then publish. After COMMIT, audit.publish sends {"type": "<entity>.<verb>", "id": ...} on the company channel, so a subscriber that refetches never sees the old row.
  5. Invalidate. The web app maps each event type to the SWR caches it makes stale and refetches them (apps/web/src/lib/live.ts). The runner uses run.created and agent.updated to try a claim, and run.updated with status: cancelled to stop a run.

Design decisions

The API records; runners execute

The API never runs an agent, holds a model key or makes outbound calls to an LLM. It queues runs and accepts reports. Execution happens in sat runner on machines you control, with the agent CLIs and credentials already installed there.

  • Why. Agent work needs a filesystem, git, long-running processes and model credentials. Keeping those out of the API keeps the API stateless apart from the database, cheap to run on one small Cloud Run instance, and free of customer model keys.
  • Consequence. Nothing happens unless a runner is running. A run stays queued until a runner that serves its agent's adapter claims it.
  • Safety. Claims are atomic (FOR UPDATE SKIP LOCKED plus a status = 'queued' guard on the update), so two runners never execute the same run. Leases (runner_id, heartbeat_at, RUN_LEASE_SECONDS) fail runs whose runner disappeared, so an agent is never stuck behind a dead run.

Events are invalidations, not data

Each SSE message carries only the event type, the entity id and a few scalar hints (for example status and agent_id on run.updated). Clients refetch the data they display.

  • Why. Clients never apply partial state, so they cannot drift from the database. A missed event costs one stale view until the next event or refetch, not corrupted state.
  • Consequence. Events are not a durable log. There is no replay, no event id and no Last-Event-ID support. Use the activity log (GET /activity) for history.

In-memory event bus by default

InMemoryBus delivers events to subscribers in the same process only. Each subscriber gets an asyncio.Queue with room for 1,000 events.

  • Why. One Cloud Run instance needs no extra infrastructure, and SAT does not require Redis.
  • Consequence. The API must run as exactly one instance. A second instance would split SSE clients across two buses. Deploy with --max-instances 1, and do not split traffic between two revisions. To run more instances, set EVENTS_BACKEND=redis; see Scaling.

Membership is the only boundary

Every company path is resolved through the caller's membership. Non-members get 404, not 403, so company ids cannot be probed. See Security.

Crons are evaluated outside the API

Routines and agent heartbeats store a cron expression, but the API never evaluates it. sat scheduler does, and calls the same endpoints a person would. See Scheduler.

On this page