Deployment¶
Maestro deploys as a single Docker Compose stack on one host: Postgres, MongoDB, Qdrant, an embedding server, the API, the web app, and Caddy for TLS. Caddy is the only service that opens a port. Everything reachable from the internet is served from one origin, which is why the frontend image carries no domain and there is no CORS to configure.
Any VM with 4 GB of RAM will do. Oracle Cloud's Always Free Ampere instance
(4 ARM cores, 24 GB) runs it comfortably at no cost; the published images are
built for arm64 as well as amd64 on every release tag.
Why not Vercel, Railway, or a serverless platform¶
The API cannot run as a serverless function:
- WebSockets.
/api/v1/tasks/{id}/streamand/api/v1/architect/liveare long-lived upgrades. Vercel Functions do not serve them. - Task duration.
TASK_TIMEOUT_SECONDSdefaults to 1800. Vercel's ceiling is 300 seconds. - Background work. A task keeps running after the HTTP response is returned. A serverless runtime freezes or kills the process at that point.
The frontend alone would sit happily on Vercel. See Running the frontend elsewhere if you want that split.
First deploy¶
1. Prepare the host¶
Install Docker Engine with the Compose plugin (v2.17 or newer — the stack
relies on service_completed_successfully):
Point a DNS A record at the host and open ports 80 and 443.
Oracle Cloud (Always Free Ampere) specifics¶
The reference free host is a VM.Standard.A1.Flex instance (4 OCPU / 24 GB,
arm64). Three things bite on Oracle that don't elsewhere:
- Ingress is blocked at two layers. Add an ingress rule for TCP 80 and 443
in the VCN Security List (or NSG) and open the host firewall — Oracle's
Ubuntu images ship iptables rules that REJECT everything but SSH, and
ufwdoes not remove them:
sudo iptables -I INPUT 6 -m state --state NEW -p tcp --dport 80 -j ACCEPT
sudo iptables -I INPUT 6 -m state --state NEW -p tcp --dport 443 -j ACCEPT
sudo netfilter-persistent save
- A1 capacity is scarce. "Out of capacity" on launch is normal for Always Free accounts: retry, try another availability domain, or upgrade the account to Pay As You Go — Always Free resources stay free, but provisioning gets priority.
- Use a release tag, not
latest. arm64 app images are only built onv*release tags; pushes tomainpublish amd64 only. On an Ampere host,IMAGE_TAGin.env.prodmust always be a release version or the pull will fail (or run nothing) for lack of an arm64 manifest.
2. Place the deployment files¶
Only four files live on the server, and one of them holds every secret you
have. None of them come from a git clone:
sudo mkdir -p /opt/maestro && cd /opt/maestro
# copy docker-compose.prod.yml, Caddyfile, scripts/backup.sh and
# .env.prod.example from the repo
cp .env.prod.example .env.prod
chmod 600 .env.prod
chmod +x backup.sh
After the first placement, deploy.yml keeps backup.sh in sync with the
repo on every tagged rollout; the other three files are host-managed only.
3. Fill in .env.prod¶
Generate real secrets — the defaults in .env.example are placeholders, and with
ENVIRONMENT=production the backend refuses to start until each one is replaced:
openssl rand -hex 32 # JWT_SECRET
openssl rand -base64 32 # API_KEY_MASTER_KEY (AES-256-GCM, encrypts BYOK keys)
openssl rand -base64 24 # POSTGRES_PASSWORD, MONGO_PASSWORD, QDRANT_API_KEY
Set DOMAIN to your hostname, then substitute the passwords you generated into
POSTGRES_URL and MONGODB_URL. Keep ?authSource=admin on the Mongo URL: the
root user cannot authenticate without it and index creation fails at startup.
The boot guard checks the substitution for you. With ENVIRONMENT=production the
backend refuses to start while CHANGE_ME is still in POSTGRES_URL,
MONGODB_URL or REDIS_URL, while either of the first two is still the
development default, or while a password is guessable — and the error in
docker compose logs backend lists every variable at fault in one line, without
echoing any value. It also refuses RATE_LIMIT_ENABLED=false and an unset
TRUST_PROXY_HEADERS. It does not demand a password where there is none: a
datastore reachable only on the compose network may run without auth. Note that
the compose-level ${POSTGRES_PASSWORD:?} guards cover the container's
variables, not the URLs the backend reads — the two can be wrong independently,
which is what this guard closes.
4. Bring it up¶
docker compose -f docker-compose.prod.yml --env-file .env.prod up -d
docker compose -f docker-compose.prod.yml --env-file .env.prod ps
migrate and ollama-pull run once and exit 0; the rest should report
healthy or running. backend does not start until migrate succeeds, so a
broken migration leaves the previous release serving rather than starting a new
one against an unmigrated schema.
Caddy requests a certificate on first request. Then:
5. Seed the marketplace (optional, once)¶
docker compose -f docker-compose.prod.yml --env-file .env.prod \
run --rm --no-deps backend python -m app.scripts.seed_marketplace
6. Schedule the account purge¶
Deletion requests are irreversible after a 30-day grace period, and something has to do the purging. Compose has no scheduler, so use the host's:
0 3 * * * cd /opt/maestro && docker compose -f docker-compose.prod.yml --env-file .env.prod run --rm --no-deps backend python -m app.scripts.purge_deleted_accounts
The script takes a Postgres advisory lock and is idempotent, so an overlapping run is harmless.
7. Schedule the email-token sweep¶
email_tokens rows are otherwise only removed by the account-deletion cascade,
so verification and password-reset rows accumulate forever. This sweep deletes
rows whose links died more than 30 days ago:
30 3 * * * cd /opt/maestro && docker compose -f docker-compose.prod.yml --env-file .env.prod run --rm --no-deps backend python -m app.scripts.purge_email_tokens
Offset half an hour from the account purge so the two never contend. It takes
its own advisory lock, deletes in batches, and accepts --dry-run to report
what a run would remove without touching anything.
8. Schedule the backups¶
See Backups below for what gets backed up and how restore works; the cron line lives there next to the rest of the backup documentation.
Continuous deployment¶
Pushing a v* tag triggers .github/workflows/deploy.yml, which SSHes to the
host and runs docker compose pull && up -d. Migrations are part of the stack,
not a separate step.
After up -d the workflow gates on health: it polls the backend's
/health/ready and the frontend's / from inside the containers for up to
~3 minutes. If the gate fails, it dumps the last container logs into the
Actions output, rolls back automatically to the image tag that was serving
before the rollout, and fails the job. A release that starts but cannot serve
never stays up just because docker compose up -d exited 0. (On a first
deploy there is nothing to roll back to; the job just fails loudly.)
Repository secrets:
| Secret | Purpose |
|---|---|
DEPLOY_HOST |
Server hostname or IP |
DEPLOY_USER |
SSH user with docker access |
DEPLOY_SSH_KEY |
Private key for that user |
Add required reviewers to the production environment in repository settings to
gate rollouts behind an approval.
Images are published to ghcr.io/yigtwxx/maestro-backend and
ghcr.io/yigtwxx/maestro-frontend on every push to main (amd64) and every
release tag (amd64 + arm64). If you keep the packages private, run
docker login ghcr.io once on the host with a read-only token.
Rolling back¶
A rollout whose health gate fails is rolled back automatically (see above). For a manual rollback — images are immutable per tag, so it is a re-deploy of the previous one:
cd /opt/maestro
IMAGE_TAG=1.2.2 docker compose -f docker-compose.prod.yml --env-file .env.prod \
up -d --no-deps backend frontend
--no-deps matters: a plain up -d re-runs the migrate one-shot with the
old image, whose alembic cannot resolve a head revision created by the newer
release — the migration fails and backend (which waits on it) never starts.
Skipping migrate leaves the schema where it is.
Note that this does not roll back the database. Migrations must be
backward-compatible with the release before them, or a rollback needs
alembic downgrade run by hand first.
Models¶
The ollama service exists for one reason: embeddings. RAG retrieval and
document upload call it no matter which chat provider a user picked, so without
it document upload fails and memory retrieval silently degrades. It pulls only
nomic-embed-text (~275 MB), which is fast on CPU.
No chat model is served. Users bring their own key (Gemini, OpenAI,
Anthropic). .env.prod sets OLLAMA_CHAT_ENABLED=false, so picking the free
local model rejects the task start with an explicit 400 telling the user to
connect a key or self-host — instead of spawning a task whose every subtask
fails with ollama chat failed and still reports completed. The UI shows the
same guidance next to the provider selector.
A user's own Ollama, running on their laptop, cannot be used by a hosted instance. Every LLM call is made by the backend, so
localhost:11434is the server, not the visitor's machine. The free tier only works when the whole stack runs on the user's own machine — or if you, the operator, opt in below.
To actually offer the free chat tier from the server, all three steps are required:
- Raise the
ollamaservice memory limit indocker-compose.prod.ymlwell above its default2g(a 9B model needs roughly 8 GB to load; the limit exists for the embedding-only default and the model will be OOM-killed inside it). - Pull a chat model:
- Set
OLLAMA_CHAT_ENABLED=truein.env.prodandup -d backend.
If you want to point the backend at an Ollama on the host rather than in a
container, set FREE_MODEL_ENDPOINT=http://host.docker.internal:11434/v1 and
start that Ollama with OLLAMA_HOST=0.0.0.0 — bound to loopback it will
refuse the connection.
Scaling to multiple workers¶
The backend defaults to a single uvicorn worker. To run more, set
WEB_CONCURRENCY (the Dockerfile passes it to uvicorn --workers):
Multi-worker (>1) requires REDIS_URL to be set — enforced at boot. The
backend refuses to start (clear config error in the logs) when
WEB_CONCURRENCY>1 and REDIS_URL is empty, instead of silently degrading to
process-local state. The production config guard runs first, so clear anything it
reports before expecting to see this one. With more than one worker,
task execution, the live event stream, human-in-the-loop answers, and task
cancellation must coordinate across processes; they do so over Redis:
- Event bus — an event emitted by the worker running a task is published on
maestro:events:{task_id}so a WebSocket subscribed on any worker receives it. With noREDIS_URL, the bus is in-process and a client connected to a different worker sees no live updates. - Control channel — cancel and HITL answers are routed on
maestro:ctrl:{task_id}to whichever worker owns the task. - Reconciliation — every worker sweeps for tasks orphaned by a crashed peer
and atomically re-claims them (a single conditional
UPDATEguarantees exactly one winner), so a mid-run crash never leaves a task stuck atrunning.
Durable state (Postgres task_runs/checkpoints, Mongo agent_logs) is always
authoritative, so a Redis outage degrades liveness (missed live ticks, slower
cancel) but never correctness — clients recover on reconnect via the
?after_seq= snapshot cursor. Leave WEB_CONCURRENCY unset (single worker) if
you are not running Redis.
Security notes¶
CODE_EXECUTION_ENABLEDshipsfalseand must stay that way in production. The tool shells out to thedockerCLI, so enabling it means mounting/var/run/docker.sockinto the backend — which hands agent-generated code the ability to start privileged containers on the host. Two independent things have to go wrong before that happens now (the socket mounted and the variable set), and the missing-daemon probe is a third, but it is an availability check rather than a security boundary — do not treat it as the gate. Leaving the tool off costs nothing but that one feature.- The refresh cookie. The three
REFRESH_COOKIE_*variables default to the secure setting and need no entry in.env.prod; the boot guard refuses production ifREFRESH_COOKIE_SECUREis turned off. Two things are worth knowing on the release that introduces it. Everyone is signed out once, because sessions predating it carried their refresh token inlocalStorageand nothing reads that any more — say so in the release note. And those tokens stay valid server-side for up toREFRESH_TOKEN_EXPIRE_DAYS; if you want them retired at the same moment, run once after the deploy:
docker compose -f docker-compose.prod.yml exec postgres \
psql -U maestro -d maestro \
-c "UPDATE refresh_tokens SET revoked_at = now() WHERE revoked_at IS NULL;"
This is a one-off operational choice, not a schema change, which is why it is
not an Alembic migration — the cookie change itself needs none.
- TRUST_PROXY_HEADERS has no safe default, which is why production will not
boot without it. .env.prod.example ships true, correct for this stack: Caddy
is the only service that opens a port, and it appends the peer it actually saw.
Set false only if you put the backend on a public port yourself — leaving
true there would let any client forge X-Forwarded-For and get a fresh
rate-limit bucket per request. Upgrading a deployment whose .env.prod predates
this guard means adding the one line before the next up -d.
- PAYMENT_PROVIDER=mock. No real money moves, and no real card should ever
be entered: payment_methods would fall under PCI scope. Ship a real
processor adapter before advertising billing.
- Transactional email. Five variables in .env.prod: EMAIL_PROVIDER,
RESEND_API_KEY, EMAIL_FROM, SITE_URL, EMAIL_VERIFICATION_REQUIRED.
A hosted instance sets EMAIL_PROVIDER=resend, a real RESEND_API_KEY, and
SITE_URL=https://<your-domain> (the base for verification/reset links in
emails). The default EMAIL_PROVIDER=console only logs messages, so
verification and password-reset links would never reach users. The backend
reads these via env_file: .env.prod, so no compose change is needed.
EMAIL_VERIFICATION_REQUIRED ships false for exactly that reason — never
turn it on before a real sender works end to end, or every account is locked
out of task start and API-key creation with no way to verify. When you do,
flip EMAIL_VERIFICATION_LIVE in frontend/src/lib/legal/config.ts in the
same change: it is a build-time constant the backend cannot set.
- FastAPI's Swagger UI, ReDoc and /openapi.json are disabled when
ENVIRONMENT=production.
- Security headers come from Caddy (header blocks in the Caddyfile),
on every response — app, API and Umami alike: HSTS (one year,
includeSubDomains, no preload), a Content-Security-Policy,
X-Content-Type-Options: nosniff, X-Frame-Options: DENY,
Referrer-Policy and Permissions-Policy; Server and X-Powered-By are
stripped. The CSP is pragmatic rather than strict: script-src and
style-src keep 'unsafe-inline' because Next.js hydration scripts, the
JSON-LD block in the root layout and React style attributes all break
without nonces. The only external origin allowed is https://*.sentry.io
in connect-src — if you point FRONTEND_SENTRY_DSN at a self-hosted
Sentry, edit that line to match its host. Header changes take effect with
the zero-downtime reload documented under "Enabling on an existing
deployment" (exec caddy caddy reload).
- Datastores publish no ports. They are reachable only from the compose network.
- .env.prod holds API_KEY_MASTER_KEY, which decrypts every user's stored
provider keys. Losing it means losing them; leaking it means leaking them.
Monitoring¶
The stack ships a lightweight observability setup: dependency health probes, self-contained operator alerting, Prometheus-format metrics, structured logs, and optional error tracking. No metrics scraper (Prometheus / Grafana) runs on the host — deliberately, to keep RAM free on a single small box. The app exposes the metrics; point your own scraper at them if you want a time series.
Set ALERT_WEBHOOK_URL or ALERT_EMAIL_TO. Without one of them nothing tells
you the site is down or erroring — the health probes are pull-based, and Sentry
alert rules are configured in Sentry's own UI, not here. The alerting below costs
nothing and needs no extra service.
Health probes¶
Two endpoints, both reachable through Caddy without authentication:
GET /health— liveness. Returns{"status":"ok"}without touching any dependency. The DockerHEALTHCHECKuses this; keep an uptime monitor on it.GET /health/ready— readiness. Pings Postgres, Mongo, Qdrant and Redis and returns200 {"status":"ready"}or, if any required service is down,503 {"status":"degraded"}.
Point a free uptime monitor (UptimeRobot, Better Stack, …) at both:
https://<your-domain>/health # expect 200 — process is up
https://<your-domain>/health/ready # expect 200 — dependencies are up
Set a 1–3 minute interval and an email/Slack alert on a non-200 response. The
/health/ready monitor catches "the API is running but Mongo/Qdrant is
unreachable" — a state /health alone would miss.
Which dependency is down is not in the public response: naming it hands an
anonymous caller a live map of your infrastructure, including the moment Redis
drops and rate-limit buckets fall back to process-local counters. Set
HEALTH_DETAIL_TOKEN in .env.prod (any random string — it grants nothing else)
and pass it to get the per-dependency breakdown:
curl -H "X-Health-Token: $HEALTH_DETAIL_TOKEN" https://<your-domain>/health/ready
# {"status":"degraded","checks":{"postgres":"ok","mongo":"error", …}}
Leave it unset and the map is withheld from everyone; the status code an uptime
monitor alerts on is unaffected either way. redis: "skipped" means no
REDIS_URL is configured, which is a supported topology, not a failure.
Operator alerts¶
The backend watches itself and pushes an alert when something breaks. Two channels, either or both:
ALERT_WEBHOOK_URL=https://hooks.slack.com/services/… # or a Discord webhook
ALERT_EMAIL_TO=ops@example.com # via EMAIL_PROVIDER
One payload serves Slack and Discord (Slack reads text, Discord reads
content), so there is no per-platform setting — paste an incoming-webhook URL
from either. Anything that accepts a JSON POST works, including an internal
notifier on the compose network. With both empty, alerting is a silent no-op
that makes no network call at all, exactly like an empty SENTRY_DSN.
Configuring a channel is the enable; there is no separate switch.
Two things fire:
| Alert | Fires when | Recovers |
|---|---|---|
| Readiness | check_readiness() fails ALERT_READINESS_FAILURES ticks in a row (2 × 60s by default) |
One good tick, with the downtime |
| Error rate | 5xx share of served requests exceeds ALERT_ERROR_RATE_THRESHOLD (5%) over ALERT_ERROR_RATE_WINDOW_SECONDS (300s), given at least ALERT_ERROR_RATE_MIN_REQUESTS (20) requests |
— |
Alerts fire on a state transition, never on a tick: a dependency that stays
down pages once and then goes quiet until it recovers. ALERT_COOLDOWN_SECONDS
(900) is a second line of defence, not the first. Degraded and recovered carry
different dedupe keys, so a recovery is never swallowed by the outage's cooldown.
Two defaults that look arbitrary and are not:
- Two failing ticks, one recovering tick. A restarting Postgres finishes well inside 120s, and a backend that boots before its dependencies pass their own healthchecks must not wake anyone. Recovery is immediate because there is no cost to being told early that things are fine.
- The 5xx threshold is a ratio, not a count. Each uvicorn worker serves
roughly 1/N of the traffic, so a raw count would silently become N times
stricter per worker; a ratio means the same thing at any
WEB_CONCURRENCY.
With REDIS_URL set, the "send this alert" right is claimed with SET NX EX, so
N workers page once rather than N times. If Redis is itself the dead dependency
the claim falls back to process-local — each worker then reports the Redis
outage, which is the right failure mode for an outage that would otherwise
silence its own alert.
The alert body names the failing dependency, which /health/ready
deliberately withholds from an anonymous caller. That is not an inconsistency:
the recipient here is the operator who has to go fix it. Everything on its way
out passes a redactor that rewrites URL credentials (postgres://user:pw@host)
and masks the exact value of every configured secret, so a body cannot carry a
key even if a future code path interpolates one.
Known limit: this runs inside the backend it watches. It reports a dead dependency, not a dead backend. See the uptime sidecar below.
Metrics¶
GET /metrics serves the Prometheus text exposition format from in-process
counters — no scraper, no extra dependency, no path label (unbounded
cardinality is how a hand-rolled registry becomes an OOM).
METRICS_TOKEN=$(openssl rand -hex 16) # in .env.prod, then restart the backend
curl -H "X-Metrics-Token: $METRICS_TOKEN" http://localhost:8000/metrics
| Metric | Type | Labels |
|---|---|---|
maestro_build_info |
gauge | version, environment |
maestro_process_start_time_seconds / _uptime_seconds |
gauge | — |
maestro_http_requests_total |
counter | status_class (1xx…5xx) |
maestro_http_request_duration_seconds |
histogram | le (+ _sum, _count) |
maestro_dependency_up |
gauge | dependency |
maestro_readiness_up |
gauge | — |
maestro_alerts_sent_total / _suppressed_total |
counter | kind |
Every series carries a worker label. Each worker keeps its own counters and
publishes a snapshot to Redis; /metrics returns its own plus every peer's, so
sum(rate(maestro_http_requests_total[5m])) is correct and a worker restart
resets one series instead of making a summed counter run backwards.
Health-probe and /metrics traffic is excluded from the counters. That is
load-bearing: /health/ready answers 503 while degraded, so counting it would
make every dependency outage also trip the 5xx alert.
Three deliberate choices worth not "fixing":
- Empty
METRICS_TOKEN→ 404, not 401. An install that never configured it is indistinguishable from one where the route does not exist. - A separate token from
HEALTH_DETAIL_TOKEN. This one is pasted into a long-lived scraper config and exposes far more; their rotations should not be coupled. - The
Caddyfiledoes not route/metrics. Its@api path /api/* /health*matcher does not match it, so the endpoint is reachable only from the compose network or an SSH tunnel — the same posture as the Umami dashboard. The token is defence in depth, not the boundary. Scrape it over a tunnel:
Uptime sidecar (optional)¶
The in-process watchdog cannot report a backend that is itself wedged,
OOM-killed or deadlocked — the failure you most need to hear about. The uptime
profile adds a container that probes from outside the process and POSTs to the
webhook itself:
# .env.prod
COMPOSE_PROFILES=uptime # or analytics,uptime
ALERT_WEBHOOK_URL=https://… # required by this profile
docker compose -f docker-compose.prod.yml --env-file .env.prod up -d
It polls backend:8000 directly rather than through Caddy, so a failure means
the app is down and not the proxy, and it alerts on a transition with the same
two-failure discipline. It is curlimages/curl running a shell loop: 3–6 MB RSS,
capped at 64 MB. Uptime Kuma would give you a dashboard for 150–250 MB plus a
volume and an interactive first run — install it separately if you want that and
have the memory; it is the wrong default for the single small box this compose
file targets.
It shares the host it watches, so it closes the "backend is wedged" blind
spot and not the "host is gone" one. Keep a free external monitor
(UptimeRobot, Better Stack) on https://<your-domain>/health — that is the only
thing that can tell you the machine itself disappeared.
Error tracking (Sentry)¶
Optional and off by default, on both sides of the stack.
Backend — to enable:
- Create a project at sentry.io (free tier is enough) and copy its DSN.
- Set
SENTRY_DSN=<dsn>in.env.prod(see alsoSENTRY_ENVIRONMENT,SENTRY_TRACES_SAMPLE_RATE) and restart the backend. - Configure an alert rule + email in the Sentry project.
Unhandled errors — including background task failures, which never reach the
request handler — are reported automatically via the logging integration. PII is
scrubbed before events leave the process: send_default_pii=False plus a
before_send hook that masks credential headers and drops request bodies, so
API keys, prompts and card data are never sent. With SENTRY_DSN empty, Sentry
is a no-op and the app makes no external calls.
Frontend — a second Sentry project (platform: Next.js), because backend and frontend events need separate DSNs and dashboards:
- Create the project and copy its DSN.
- Set
FRONTEND_SENTRY_DSN=<dsn>in.env.prodanddocker compose -f docker-compose.prod.yml --env-file .env.prod up -d frontend.
The DSN is read at runtime (server-side, like SITE_URL), so the image stays
domain-agnostic. Server-side render errors are captured via onRequestError;
browser errors are captured by the error boundaries once the lazily-loaded SDK
initializes. With the variable empty, no Sentry chunk is ever served to
browsers and the app makes zero requests to any ingest host.
Deliberate trade-offs (all cheap to revisit): no source-map upload (client
stack traces are minified; enabling it would need withSentryConfig plus a CI
auth token), no tunnel route (ad-blockers may drop browser events; server
capture is unaffected), no session replay (bundle size + PII surface).
Logs¶
LOG_FORMAT=json (the production default in .env.prod.example) emits one JSON
object per line — pipe docker compose logs into any aggregator. LOG_FORMAT=text
keeps the readable format for local debugging.
Every HTTP response carries an X-Request-ID header, and each request emits one
maestro.access log line with request_id, method, path, status and
duration_ms (health probes excluded to keep the noise down). The same id is
attached to error logs and Sentry events, so a user-reported failure correlates
directly:
Caddy writes JSON access logs to stdout (log directive in the Caddyfile),
giving request-level visibility in front of both apps — including uptime-monitor
hits on /health*, which are not filtered at the proxy.
Container logs are rotated by the json-file caps in docker-compose.prod.yml
(max-size: 10m, max-file: 5 per service — roughly 550 MB worst case for the
whole stack). Changing the caps requires recreating the containers; up -d
does that and drops the old log files.
Analytics (optional, self-hosted Umami)¶
Off by default. When enabled it is first-party (data never leaves the host),
cookieless, consent-gated (the notice becomes a real Accept/Reject, the script
loads only after an explicit yes), and counts the public marketing pages only —
never the signed-in app. The Umami dashboard is deliberately not exposed to the
internet; only /a/script.js and /a/api/send pass through Caddy.
Enabling¶
- In
.env.prod, set:
COMPOSE_PROFILES=analytics
UMAMI_DB_PASSWORD=<openssl rand -base64 24>
UMAMI_APP_SECRET=<openssl rand -hex 32>
-
docker compose -f docker-compose.prod.yml --env-file .env.prod up -d.umami-db-initruns once and creates theumamirole and database inside the existing Postgres (idempotent — safe on an already-initializedpgdatavolume, which is why an initdb script would not work);umamithen runs its own Prisma migrations and comes up healthy. -
Reach the dashboard through an SSH tunnel — it listens on the host's loopback only:
Open http://localhost:3001, log in as admin / umami, and change that
password immediately. Then Settings → Websites → Add website (name:
maestro, domain: your DOMAIN) and copy the Website ID.
- Put the id in
.env.prodasUMAMI_WEBSITE_ID=<id>and restart the frontend so it picks up the new env:
Visitors now get the Accept/Reject notice; views appear in the dashboard only
for those who accept, and only on marketing pages. Consent can be changed any
time at the bottom of /cookies.
Enabling on an existing deployment¶
Copy the updated docker-compose.prod.yml and Caddyfile to the host first,
then follow the steps above. up -d recreates Caddy with the /a/* route (or
reload in place: docker compose -f docker-compose.prod.yml --env-file .env.prod
exec caddy caddy reload --config /etc/caddy/Caddyfile). Nothing else in the
stack restarts.
Disabling¶
Set COMPOSE_PROFILES= and UMAMI_WEBSITE_ID= back to empty and up -d
--remove-orphans. The frontend reverts to the informational notice; the
umami database stays in pgdata (harmless) unless you drop it.
Backups¶
backup.sh (from scripts/backup.sh in the repo, kept in sync by deploy.yml)
dumps every durable store into /opt/maestro/backups, applies retention and
optionally pushes the files offsite. It is scheduled from the host crontab,
exactly like the purge job.
What is backed up¶
| Store | What | File |
|---|---|---|
| Postgres | maestro DB — accounts, subscriptions, quota ledger; the one that matters most |
maestro-pg-<ts>.sql.gz (pg_dump --clean --if-exists) |
| Postgres | umami DB — only when the analytics profile is active in .env.prod |
maestro-umami-<ts>.sql.gz |
| MongoDB | maestro DB — agent logs, task sessions, marketplace, agent configs |
maestro-mongo-<ts>.archive.gz (mongodump --archive --gzip) |
| Qdrant | every collection, enumerated dynamically — per-collection snapshots over the HTTP API | maestro-qdrant-<collection>-<ts>.snapshot |
Qdrant is distroless and publishes no ports, so the script talks to it through
a one-shot curlimages/curl container joined to the compose network; each
snapshot is downloaded over HTTP and then deleted server-side so nothing
accumulates inside the qdrantdata volume.
Qdrant snapshots require named-volume storage. On the production stack (
qdrantdatanamed volume) a snapshot of a near-empty collection is ~140 KB and restores cleanly. On a Windows Docker Desktop bind mount (the dev stack's./.data/qdrant) the same snapshot balloons to ~400 MB (sparse WAL files get materialized) and itswal/first-indexis zeroed, so the restore fails with a WAL deserialize error. This is a dev-only filesystem artifact — verified 2026-07-13 againstv1.18.2both ways — not a production risk.
Deliberately not backed up: redis (ephemeral rate-limit buckets,
persistence is disabled on purpose), ollamadata (models are re-pulled by
ollama-pull), caddydata (TLS certificates re-issue automatically on the
first request after a rebuild).
.env.prod is never touched by the script. Back it up manually,
encrypted (e.g. in a password manager) — losing API_KEY_MASTER_KEY makes
every BYOK key in a Postgres dump permanently undecryptable, and leaking it
decrypts all of them. The script pushes only the backups directory offsite,
so secrets structurally cannot leave the host through it.
Schedule and retention¶
Daily at 03:30 (offset from the 03:00 purge), guarded by flock so an
overlapping run exits instead of stacking:
30 3 * * * cd /opt/maestro && RCLONE_REMOTE=oci:maestro-backups flock -n /opt/maestro/backup.lock ./backup.sh >> /opt/maestro/backup.log 2>&1
Retention is applied by the script itself, locally and remotely: daily/
keeps 7 days, weekly/ (a copy made every Sunday) keeps 28 days. Pruning is
scoped to the maestro-* naming pattern and never deletes anything else.
Run the script once by hand over SSH before trusting the cron line, and note
backup.log grows about a line per day — truncate it ad hoc.
Offsite copy (Oracle Object Storage via rclone)¶
The offsite push is env-gated: leave RCLONE_REMOTE unset (drop it from the
cron line) and the script stays local-only. To enable it:
- In the OCI console create a private bucket, e.g.
maestro-backups(Always Free includes 20 GB of Object Storage). - Create a Customer Secret Key for a least-privilege IAM user (Identity → Users → Customer Secret Keys) — this is OCI's S3-compatible credential.
- Install rclone on the host (
sudo apt install rcloneor the arm64 static binary) and configure~/.config/rclone/rclone.conf:
[oci]
type = s3
provider = Other
access_key_id = <customer secret key id>
secret_access_key = <customer secret key>
endpoint = https://<namespace>.compat.objectstorage.<region>.oraclecloud.com
region = <region>
- Set
RCLONE_REMOTE=oci:maestro-backupsin the cron line.
The script uses rclone copy (additive), deliberately not sync: syncing
from a freshly rebuilt host with an empty backups directory would delete
every offsite backup — exactly the disaster the offsite copy exists to
survive. Remote retention is pruned separately with rclone delete --min-age.
Restore runbook¶
All commands run from /opt/maestro. Define the compose prefix once:
Postgres — the dump carries --clean --if-exists, so piping it in
drops and recreates every object:
$DC stop backend
gunzip -c maestro-pg-<ts>.sql.gz | $DC exec -T postgres psql -U maestro -d maestro
$DC up -d backend
Umami restores the same way into -d umami. For a cheap drill without
touching live data, restore into a scratch DB first:
$DC exec -T postgres createdb -U maestro scratch then psql ... -d scratch.
MongoDB — --drop replaces each collection in place:
$DC exec -T mongo mongorestore --archive --gzip --drop --nsInclude 'maestro.*' \
-u "$MONGO_USER" -p "$MONGO_PASSWORD" --authenticationDatabase admin \
< maestro-mongo-<ts>.archive.gz
Qdrant — upload the snapshot per collection; the upload recreates the
collection even if it no longer exists (read QDRANT_API_KEY from .env.prod
rather than exporting it into shell history):
cat maestro-qdrant-<collection>-<ts>.snapshot | docker run --rm -i \
--network maestro_default curlimages/curl:8.11.1 -fsS \
-H "api-key: $QDRANT_API_KEY" \
-F "snapshot=@-;filename=restore.snapshot" \
"http://qdrant:6333/collections/<collection>/snapshots/upload?priority=snapshot"
Full disaster recovery, in order: provision a fresh host → place the four
files (compose, Caddyfile, .env.prod from your encrypted copy, backup.sh)
→ up -d (this runs migrate against the empty database) → restore Postgres,
Mongo, then Qdrant from the offsite bucket (rclone copy oci:maestro-backups/daily .)
→ $DC restart backend. If the dump predates the current migration head,
restoring it rewinds the schema to the dump's state; run
$DC up -d migrate afterwards to bring it forward again.
Local smoke test¶
Verify the production stack on your own machine before touching a server. Set
DOMAIN=http://localhost in .env.prod — the http:// prefix tells Caddy to
serve plain HTTP and skip certificate provisioning.
cp .env.prod.example .env.prod # DOMAIN=http://localhost, fill the secrets
docker compose -f docker-compose.prod.yml --env-file .env.prod up -d --build
Keep real secrets out of the repo tree. A
.env.prodfilled with production values (notablyAPI_KEY_MASTER_KEY, which decrypts every stored BYOK key) must not sit inside the working tree — it is.gitignored, but onegit add -faway from being committed. Keep the real file outside the repo (e.g.~/.maestro/.env.prod, locked to your user) and pass its absolute path to--env-file. On the production host the file already lives beside the compose file at/opt/maestro/.env.prod, so the relative form is correct there; only local runs off an out-of-tree copy need the absolute path:
Then check, in order:
docker compose -f docker-compose.prod.yml --env-file .env.prod ps—migrateandollama-pullexited (0), the resthealthy.curl http://localhost/health→{"status":"ok"}.curl http://localhost/api/v1/billing/plans/public→ JSON.- Open
http://localhost/pricing. Prices must render — the "unreachable" banner means the server-side fetch failed. - Open
http://localhost/docs. You should get the marketing page, not Swagger. - Register, log in, start a task. In DevTools → Network → WS, the connection to
/api/v1/tasks/<id>/streammust report 101 Switching Protocols and stream events. - Upload a
.txtunder Documents. A200proves the embedding service is wired up. - If testing analytics (
COMPOSE_PROFILES=analytics+ theUMAMI_*vars): in a fresh browser profile the notice must show Accept/Reject with no/a/script.jsrequest before you answer. After Accept, each marketing page navigation sends onePOST /a/api/send; navigating into/loginor/dashboardmust send zero.curl http://localhost/api/sendmust still reach the backend (a FastAPI 404/405, not Umami) — that proves the/amatcher split.
Running the frontend elsewhere¶
The frontend image is domain-agnostic, and the code supports a split deploy without a rebuild. On Vercel (or any Node host), set:
| Variable | Value |
|---|---|
BACKEND_ORIGIN |
https://api.your-domain — proxies /api/*, keeping the browser same-origin |
INTERNAL_API_ORIGIN |
https://api.your-domain — used by the server-rendered /pricing and /templates |
NEXT_PUBLIC_WS_BASE_URL |
wss://api.your-domain — WebSockets cannot be proxied through Next rewrites |
SITE_URL |
https://your-domain — origin for canonical URLs, sitemap.xml, OG tags |
The backend then needs CORS_ORIGINS set to the app's origin, since the
WebSocket handshake no longer shares it. The API itself still has to run as a
long-lived container somewhere.
BACKEND_ORIGIN is not optional in this topology, and the refresh cookie is why:
it is SameSite, so it only travels on requests the browser considers same-site.
Proxying /api/* through Next keeps them that way. Calling the API host directly
from the browser would leave every sign-in unable to survive a reload. If app and
API share one registrable domain (app.example.com / api.example.com), the
alternative is REFRESH_COOKIE_DOMAIN=.example.com — at the cost of handing the
session cookie to every subdomain. There is deliberately no SameSite=None
option; see CLAUDE.md §8.
SITE_URL is what keeps the image domain-agnostic despite SEO needing an
absolute origin: it is read at request time, not baked in, so the same image
serves any domain and changing it needs only a container restart, not a
rebuild. It is deliberately not NEXT_PUBLIC_ (that would inline it at build
time) and not derived from DOMAIN (which may carry a scheme). Unset, the
pages fall back to a placeholder domain — visibly wrong rather than silently
plausible.