Smart Agent Teams

Scaling

The current capacity limits of SAT, why the API runs as one instance, and what to change to scale it out.

SAT is designed today to run as one API instance, one PostgreSQL database, and as many runners as you start. This page explains each limit, which parts are already safe to run in parallel, and the exact changes needed to run more than one API instance.

Default shape

PartDefaultSet by
API1 instance (on Cloud Run, --max-instances 1), one uvicorn process, no workers flagYour deployment, apps/api/Dockerfile
Event busIn memory, inside that processEVENTS_BACKEND unset (memory)
DatabasePostgreSQL 16; a small instance is enough for one API instanceYour deployment
ConnectionsNullPool: one new connection per requestdatabase/connection.py
Rate-limit countersIn memory, per processslowapi default storage
Traffic splitNone. Serve one revision at a timeYour deployment
RunnersAny number, on your machinessat runner start
SchedulerOne per company, on your machinesat scheduler start

Why one API instance

The SSE bus is in memory

InMemoryBus (company/events.py) keeps a list of subscribers per company in the process. An event published by a request is delivered only to SSE connections held by the same process. With two instances, a browser connected to instance A never hears about a change made through instance B, and a runner connected to A never hears run.created for a run queued through B.

Runners would still work, because they poll for claims every 30 seconds, but live updates in the web app would be unreliable. That is why the API should run with at most one instance, and why a rollout should switch all traffic to a new revision at once (after a health check) instead of splitting traffic gradually.

Migrations run on every start

init_db() runs alembic upgrade head in every instance's startup, without a lock. Several instances starting together race on the same migration; an instance that fails records the error and answers 503 on /health until it restarts. See Database.

Rate limits are per process

slowapi keeps counters in memory, so with N instances each limit is effectively N times higher.

What is already safe in parallel

These parts were built to tolerate several API instances and several runners:

OperationWhy it is safe
Run claims (POST /runs/claim)SELECT ... FOR UPDATE SKIP LOCKED on the oldest eligible queued run, then UPDATE ... WHERE status = 'queued'. If the update affects no row, the claim retries, up to 5 times, then answers 204. Two runners never get the same run.
One run per agentA claim skips agents that already have a running run.
Task numbersnext_task_number locks the company row (SELECT ... FOR UPDATE) and increments task_counter; tasks (company_id, number) is unique.
Duplicate runsqueue_run reuses an identical queued run (same agent, task and routine) instead of adding another.
Lease expiryAny claim, expire-stale call or scheduler sweep can expire stale runs; finishing a run is idempotent on status.
Sessions and authStateless JWTs and API keys; nothing is held in API memory between requests.

Scaling out the API

To run more than one API instance, make these changes in order.

Provide Redis

Create a Redis instance reachable from the Cloud Run service's VPC (for example Memorystore for Redis in the same VPC). SAT does not provision one today.

Switch the event bus to Redis

gcloud run services update <service> --project <project> --region <region> \
  --update-env-vars EVENTS_BACKEND=redis,REDIS_URL=redis://10.x.x.x:6379/0

RedisBus publishes each event to the channel sat:company:<company_id>:events and every instance's SSE handlers subscribe to it, so every client hears every event. Publishing is synchronous inside the request, so Redis latency adds to write latency, and a Redis outage makes writes fail after their database commit (the change is saved, but the request errors and the event is lost).

Run migrations once, before the new revision

Move alembic upgrade head out of instance startup, for example into a Cloud Run job that runs the same image with alembic upgrade head before the service deploys. init_db() would still run on each instance; once the schema is at head it does nothing. Until migrations are separated, scale out only between releases that contain no new migration.

Add a connection pool

Replace NullPool with a bounded pool (pool_size, max_overflow, pool_pre_ping=True) sized so that instances times pool size stays under the database's connection limit, and size the database instance for the extra connections. With NullPool, connections equal concurrent requests across all instances.

Move rate-limit storage

Point slowapi at shared storage (it supports Redis) if the per-endpoint limits must hold across instances.

Raise the instance limit

gcloud run services update <service> --project <project> --region <region> --max-instances 3

Only then can your rollout split traffic between revisions. Each instance holds its SSE connections open, so set Cloud Run concurrency to allow for one open request per connected browser tab and runner, plus normal traffic.

Scaling runners

Runners scale horizontally with no API changes.

ControlDefaultEffect
--concurrency <n>2Runs one runner process executes at once. Each run spawns one agent CLI process.
Number of runnersanyStart runners on several machines. Claims are atomic, so they never collide.
--agents, --adaptersall agents whose adapter is installedPartition work: dedicate a machine to some agents or one adapter.
--poll <seconds>30 (minimum 5)Safety-net claim interval when the event stream is down.

Limits that do not change with more runners:

  • One run per agent at a time. An agent with a long queue is processed serially no matter how many runners you start. Spread work across agents to parallelise it.
  • The API's single instance handles every claim, heartbeat (every 30 seconds per run) and transcript upload (every 2 seconds or 4 KB per run). Many concurrent runs mean many small writes against one instance and one database.
  • Transcripts are one JSON column per run. Each upload rewrites the run's growing transcript array. Very long runs make each update larger.

Scaling the scheduler

Run one sat scheduler per company. It is a single process with in-memory cron jobs.

  • A second scheduler is mostly harmless: before triggering a routine it re-reads last_run_at and skips the routine if it ran within the duplicate window (50 seconds, capped at half the schedule's interval). Heartbeat wakes reuse an identical queued run. Two schedulers can still double-fire if their clocks or the API are slow enough to miss that window.
  • Ticks missed while no scheduler is running are not backfilled.
  • The scheduler reloads definitions when it sees routine.* or agent.* events and every 5 minutes, and expires stale runs every 60 seconds (--expire-interval).

See Scheduler.

Data volume

ConcernTodayWhen it matters
GET /tasksReturns up to 1,000 tasks (default 500), no cursorCompanies with more open tasks than that need filters
GET /runsUp to 500 (default 100), newest first, no cursorUse agent_id, task_id or status filters
GET /activityUp to 200 (default 50), cursor beforePaginate with before
Dashboard and costsAggregate queries over the month's runsGrows with run count per month
audit_events, runsNever deletedPlan disk and archiving; see Database

All limits are listed in Limits.

On this page