Skip to main content

Installation

This page deploys a fresh instance with Docker Compose. When you are done you will have four containers running, a database with its schema applied, and an administrator account you created through the browser.

Prerequisites​

  • A Linux host with Docker Engine and the Compose v2 plugin (docker compose version must work — docker-compose v1 is not enough).
  • Disk. Metric history is the only thing that grows without your help. With the shipped retention the raw table holds about a week and the rollups hold the rest; budget tens of gigabytes for a few dozen hosts and read Backup and Restore for the retention numbers.
  • A DNS name and a TLS-terminating reverse proxy in front of it. The compose file publishes every port on 127.0.0.1 only, so nothing is reachable from outside the host until you put a proxy there.
  • openssl for generating secrets.

The compose file​

Copy docker-compose.yml from the repository next to an .env file on the host. It defines four services:

ServiceContainerWhat it is for
appselvara-appThe dashboard and the API. Runs migrations at start, then serves on port 3000.
postgresselvara-dbtimescale/timescaledb:latest-pg16. Holds everything: users, customers, systems, settings, metrics.
redisselvara-redisPub/sub for the live updates and the rate-limit counters. Nothing durable.
cronselvara-cronAn Alpine container with a one-line crontab that calls POST /api/cron/evaluate-alerts every minute. Without it no alert is ever raised or resolved.

Two things in that file are deliberate and worth not changing:

  • name: selvara at the top pins the Compose project name. Volume names are derived from it, so a checkout into a differently named directory keeps finding the same database instead of quietly starting an empty one.
  • container_name: selvara-app pins the app to a single container. See Upgrading for why two must never run at once.

The file names a prebuilt image. To build from source instead, uncomment the build: . line under the app service. The Docker build compiles the Go agent in its own stage and writes the linux/amd64 and linux/arm64 binaries into the image, so the agent installer works straight after the build.

Generate the secrets​

Create .env next to the compose file. Compose refuses to start when AUTH_SECRET, WEBHOOK_SECRET or CRON_SECRET is missing, so generate all of them now:

cat > .env <<EOF
POSTGRES_PASSWORD=$(openssl rand -hex 24)
AUTH_SECRET=$(openssl rand -base64 48)
WEBHOOK_SECRET=$(openssl rand -base64 48)
CRON_SECRET=$(openssl rand -base64 32)
SETTINGS_ENCRYPTION_KEY=$(openssl rand -base64 48)
AUTH_URL=https://<your-domain>
EOF
chmod 600 .env

SETTINGS_ENCRYPTION_KEY is not in the shipped compose file and it is not required — but set it anyway, on day one. Without it the code falls back to AUTH_SECRET for encrypting stored secrets, which ties two unrelated lifetimes together: rotating your session secret then makes every stored SMTP password, agent secret and TOTP secret undecryptable. Configuration explains the fallback in full.

To pass it through, add one line to the app service's environment: block:

SETTINGS_ENCRYPTION_KEY: ${SETTINGS_ENCRYPTION_KEY:?SETTINGS_ENCRYPTION_KEY is required}

POSTGRES_PASSWORD is only ever read when PostgreSQL initialises an empty data directory. Changing it in .env later does not change the password inside an existing volume — it only changes the one the app tries to connect with, and the app then fails to reach the database. Decide it once, here.

Bring it up​

docker compose pull
docker compose up -d
docker compose logs -f app

What the first start does​

There is no separate migrate command and you should not look for one. The image's entrypoint runs, in this order, before anything listens on port 3000:

  1. node scripts/migrate.mjs — applies every pending schema migration. The migrations are inline in that script, not files in a directory, and the Kysely migrator tracks which ones have run in the database.
  2. The same script then seeds the default alert rules and event rules, once, into an instance that has none. A marker in app_settings stops it happening again, so rules you delete stay deleted.
  3. It configures TimescaleDB: makes metrics a hypertable, enables compression, creates the metrics_5m and metrics_1h rollups, and registers the retention and compression policies. This step is idempotent and runs on every boot. If it fails it is logged and the app still starts — it is not allowed to block startup.
  4. node scripts/seed.mjs — without --demo this only ensures the default alert rules exist. It creates no users and no sample servers.
  5. node server.js — the dashboard starts.

A healthy first boot looks like this:

Running database migrations...
Applied: 0001_baseline
...
Applied: 0007_customers_and_server_region
Applied 6 migration(s)
Migrations completed successfully
Created 36 default alert rules
Created 1 default event rules
Configuring TimescaleDB retention...
- metrics is now a hypertable (1 day chunks)
- created rollup metrics_5m (5 minutes buckets)
- created rollup metrics_1h (1 hour buckets)
Retention active: raw 7 days (compressed after 2 days), 5m rollup 30 days, 1h rollup 400 days
Running database seed...
Seeding database...
Production mode: Skipping sample data (use --demo for sample data)
Seeding complete!
Starting application...

On an empty database this takes seconds. The container healthcheck allows a 900-second start period anyway, because an instance that is being converted from a large pre-existing metrics table spends minutes backfilling the rollups. Silence during backfilling metrics_1h ... this can take a while is work, not a hang.

Create the administrator​

There is no seeded admin account and no environment variable that creates one. Shipping credentials in the environment meant every install started with a password that had been typed into a shell, so that path was removed.

Open https://<your-domain>/ in a browser. With no account in the database, /login redirects to /setup, which asks for a name, an e-mail address, an optional username and a password — see Users and Permissions for what one must be. The account it creates has the ADMIN role.

That page closes itself. It is guarded by whether the users table is empty, and the API behind it takes a Postgres advisory lock before it checks, so two requests arriving together cannot both create an owner. Once an account exists, /setup redirects to /login and POST /api/setup answers 409 Already set up. The endpoint is also rate-limited to ten attempts per ten minutes per client address.

If you never reach the setup screen, check Troubleshooting — a wrong AUTH_URL is the usual reason.

Because the page closes itself, a lost password is not recovered here. With SMTP configured the login screen offers a reset link; without it, scripts/reset-password.mjs sets a new password from a shell that reaches the database — Troubleshooting has the commands.

Health endpoint​

GET /api/health needs no authentication and is what the container healthcheck polls:

curl -s http://127.0.0.1:3000/api/health
{"status":"ok","timestamp":"...","services":{"database":"ok","redis":"ok"}}

It returns HTTP 200 when both checks pass and HTTP 503 with "status":"degraded" when either the database or Redis does not answer. Point your external uptime check at it. The cron container waits for this healthcheck before it starts, so an app that never becomes healthy is also an app whose alerts are never evaluated.

Behind a reverse proxy​

The app listens on 127.0.0.1:3000. Terminate TLS in front of it and forward to that address. The proxy must pass X-Forwarded-For; the rate limiter reads the last entry of that header, on the grounds that a client can forge the leading entries but not the one your own proxy appends.

AUTH_URL must be the exact public URL, scheme and host included, with no trailing slash. It is what NextAuth builds its callback and CSRF endpoints from, and a mismatch does not degrade gracefully — the login form simply stops working. https://<your-domain> is right; http://<your-domain>, https://<your-domain>/ and the container's internal address are all wrong.

Set NEXTAUTH_URL to the same value. The shipped compose file already does this from one variable, and it matters: a couple of routes read NEXTAUTH_URL without falling back to AUTH_URL, and with it unset they emit agent install commands and download URLs with an empty host.

WebSockets​

Every open dashboard keeps one WebSocket to /api/live, and every page and list follows the server over it: a new alert, a system going offline, a rule someone else just edited, a chart's newest bucket — none of it needs a reload, and there is no refresh button. The proxy has to forward the WebSocket upgrade, or the pages load but never change until reloaded.

Caddy and Traefik forward upgrades without extra configuration. nginx does not; the location that proxies to the app needs:

location / {
proxy_pass http://127.0.0.1:3000;
proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection $connection_upgrade;
proxy_set_header Host $host;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
}

with, once in the http block:

map $http_upgrade $connection_upgrade {
default upgrade;
'' close;
}

The server pings every open socket every 25 seconds, which keeps it inside nginx's default proxy_read_timeout of 60 seconds.

The socket is authenticated by the same session cookie as the pages, and accepted only from a page whose origin is this dashboard: the host the browser connected to, or the host of AUTH_URL. A proxy that rewrites Host has to pass X-Forwarded-Host or AUTH_URL has to be right — otherwise the socket is refused with 401 while the pages themselves work.

A maintenance page during upgrades​

While the app container restarts, nothing listens on port 3000 and the proxy answers with a bare 502. The repository ships a self-contained page for that moment, deploy/caddy/maintenance.html: logo, a short note in German or English depending on the browser, and a check of /api/health every five seconds that reloads the page once the new container answers. Copy it to the proxy host, for example to /etc/caddy/selvara/maintenance.html, and add this to the site block:

handle_errors {
@down `{err.status_code} in [502, 503, 504]`
handle @down {
header Cache-Control no-store
header Retry-After 30
root * /etc/caddy/selvara
rewrite * /maintenance.html
file_server {
status 503
}
}
}

handle_errors only fires when Caddy itself cannot reach the app. An error the app returns — a 500 or the 503 of a degraded health check — reaches the client unchanged. Agents get the 503 in place of the 502 and retry as before. file_server's status needs Caddy 2.7 or later.

Next steps​