Installation
This page deploys a fresh instance with Docker Compose. When you are done you will have four containers running, a database with its schema applied, and an administrator account you created through the browser.
Prerequisites
- A Linux host with Docker Engine and the Compose v2 plugin
(
docker compose versionmust work —docker-composev1 is not enough). - Disk. Metric history is the only thing that grows without your help. With the shipped retention the raw table holds about a week and the rollups hold the rest; budget tens of gigabytes for a few dozen hosts and read Backup and Restore for the retention numbers.
- A DNS name and a TLS-terminating reverse proxy in front of it. The compose
file publishes every port on
127.0.0.1only, so nothing is reachable from outside the host until you put a proxy there. opensslfor generating secrets.
The compose file
Copy docker-compose.yml from the repository next to an .env file on the
host. It defines four services:
| Service | Container | What it is for |
|---|---|---|
app | selvara-app | The dashboard and the API. Runs migrations at start, then serves on port 3000. |
postgres | selvara-db | timescale/timescaledb:latest-pg16. Holds everything: users, customers, systems, settings, metrics. |
redis | selvara-redis | Pub/sub for the live updates and the rate-limit counters. Nothing durable. |
cron | selvara-cron | An Alpine container with a one-line crontab that calls POST /api/cron/evaluate-alerts every minute. Without it no alert is ever raised or resolved. |
Two things in that file are deliberate and worth not changing:
name: selvaraat the top pins the Compose project name. Volume names are derived from it, so a checkout into a differently named directory keeps finding the same database instead of quietly starting an empty one.container_name: selvara-apppins the app to a single container. See Upgrading for why two must never run at once.
The file names a prebuilt image. To build from source instead, uncomment the
build: . line under the app service. The Docker build compiles the Go agent
in its own stage and writes the linux/amd64 and linux/arm64 binaries into
the image, so the agent installer works straight after the build.
Generate the secrets
Create .env next to the compose file. Compose refuses to start when
AUTH_SECRET, WEBHOOK_SECRET or CRON_SECRET is missing, so generate all of
them now:
cat > .env <<EOF
POSTGRES_PASSWORD=$(openssl rand -hex 24)
AUTH_SECRET=$(openssl rand -base64 48)
WEBHOOK_SECRET=$(openssl rand -base64 48)
CRON_SECRET=$(openssl rand -base64 32)
SETTINGS_ENCRYPTION_KEY=$(openssl rand -base64 48)
AUTH_URL=https://<your-domain>
EOF
chmod 600 .env
SETTINGS_ENCRYPTION_KEY is not in the shipped compose file and it is not
required — but set it anyway, on day one. Without it the code falls back to
AUTH_SECRET for encrypting stored secrets, which ties two unrelated lifetimes
together: rotating your session secret then makes every stored SMTP password,
agent secret and TOTP secret undecryptable. Configuration
explains the fallback in full.
To pass it through, add one line to the app service's environment: block:
SETTINGS_ENCRYPTION_KEY: ${SETTINGS_ENCRYPTION_KEY:?SETTINGS_ENCRYPTION_KEY is required}
POSTGRES_PASSWORD is only ever read when PostgreSQL initialises an empty data
directory. Changing it in .env later does not change the password inside an
existing volume — it only changes the one the app tries to connect with, and
the app then fails to reach the database. Decide it once, here.
Bring it up
docker compose pull
docker compose up -d
docker compose logs -f app
What the first start does
There is no separate migrate command and you should not look for one. The image's entrypoint runs, in this order, before anything listens on port 3000:
node scripts/migrate.mjs— applies every pending schema migration. The migrations are inline in that script, not files in a directory, and the Kysely migrator tracks which ones have run in the database.- The same script then seeds the default alert rules and event rules, once,
into an instance that has none. A marker in
app_settingsstops it happening again, so rules you delete stay deleted. - It configures TimescaleDB: makes
metricsa hypertable, enables compression, creates themetrics_5mandmetrics_1hrollups, and registers the retention and compression policies. This step is idempotent and runs on every boot. If it fails it is logged and the app still starts — it is not allowed to block startup. node scripts/seed.mjs— without--demothis only ensures the default alert rules exist. It creates no users and no sample servers.node server.js— the dashboard starts.
A healthy first boot looks like this:
Running database migrations...
Applied: 0001_baseline
...
Applied: 0007_customers_and_server_region
Applied 6 migration(s)
Migrations completed successfully
Created 36 default alert rules
Created 1 default event rules
Configuring TimescaleDB retention...
- metrics is now a hypertable (1 day chunks)
- created rollup metrics_5m (5 minutes buckets)
- created rollup metrics_1h (1 hour buckets)
Retention active: raw 7 days (compressed after 2 days), 5m rollup 30 days, 1h rollup 400 days
Running database seed...
Seeding database...
Production mode: Skipping sample data (use --demo for sample data)
Seeding complete!
Starting application...
On an empty database this takes seconds. The container healthcheck allows a
900-second start period anyway, because an instance that is being converted
from a large pre-existing metrics table spends minutes backfilling the
rollups. Silence during backfilling metrics_1h ... this can take a while is
work, not a hang.
Create the administrator
There is no seeded admin account and no environment variable that creates one. Shipping credentials in the environment meant every install started with a password that had been typed into a shell, so that path was removed.
Open https://<your-domain>/ in a browser. With no account in the database,
/login redirects to /setup, which asks for a name, an e-mail address, an
optional username and a password — see
Users and Permissions for what one must be.
The account it creates has the ADMIN role.
That page closes itself. It is guarded by whether the users table is empty,
and the API behind it takes a Postgres advisory lock before it checks, so two
requests arriving together cannot both create an owner. Once an account exists,
/setup redirects to /login and POST /api/setup answers 409 Already set up. The endpoint is also rate-limited to ten attempts per ten minutes per
client address.
If you never reach the setup screen, check Troubleshooting —
a wrong AUTH_URL is the usual reason.
Because the page closes itself, a lost password is not recovered here. With
SMTP configured the login screen offers a reset link; without it,
scripts/reset-password.mjs sets a new password from a shell that reaches the
database — Troubleshooting has the commands.
Health endpoint
GET /api/health needs no authentication and is what the container healthcheck
polls:
curl -s http://127.0.0.1:3000/api/health
{"status":"ok","timestamp":"...","services":{"database":"ok","redis":"ok"}}
It returns HTTP 200 when both checks pass and HTTP 503 with
"status":"degraded" when either the database or Redis does not answer. Point
your external uptime check at it. The cron container waits for this
healthcheck before it starts, so an app that never becomes healthy is also an
app whose alerts are never evaluated.
Behind a reverse proxy
The app listens on 127.0.0.1:3000. Terminate TLS in front of it and forward
to that address. The proxy must pass X-Forwarded-For; the rate limiter reads
the last entry of that header, on the grounds that a client can forge the
leading entries but not the one your own proxy appends.
AUTH_URL must be the exact public URL, scheme and host included, with no
trailing slash. It is what NextAuth builds its callback and CSRF endpoints
from, and a mismatch does not degrade gracefully — the login form simply stops
working. https://<your-domain> is right; http://<your-domain>,
https://<your-domain>/ and the container's internal address are all wrong.
Set NEXTAUTH_URL to the same value. The shipped compose file already does
this from one variable, and it matters: a couple of routes read NEXTAUTH_URL
without falling back to AUTH_URL, and with it unset they emit agent install
commands and download URLs with an empty host.
WebSockets
Every open dashboard keeps one WebSocket to /api/live, and every page and
list follows the server over it: a new alert, a system going offline, a rule
someone else just edited, a chart's newest bucket — none of it needs a reload,
and there is no refresh button. The proxy has to forward the WebSocket
upgrade, or the pages load but never change until reloaded.
Caddy and Traefik forward upgrades without extra configuration. nginx does not; the location that proxies to the app needs:
location / {
proxy_pass http://127.0.0.1:3000;
proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection $connection_upgrade;
proxy_set_header Host $host;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
}
with, once in the http block:
map $http_upgrade $connection_upgrade {
default upgrade;
'' close;
}
The server pings every open socket every 25 seconds, which keeps it inside
nginx's default proxy_read_timeout of 60 seconds.
The socket is authenticated by the same session cookie as the pages, and
accepted only from a page whose origin is this dashboard: the host the browser
connected to, or the host of AUTH_URL. A proxy that rewrites Host has to
pass X-Forwarded-Host or AUTH_URL has to be right — otherwise the socket is
refused with 401 while the pages themselves work.
A maintenance page during upgrades
While the app container restarts, nothing listens on port 3000 and the proxy
answers with a bare 502. The repository ships a self-contained page for that
moment, deploy/caddy/maintenance.html: logo, a short note in German or
English depending on the browser, and a check of /api/health every five
seconds that reloads the page once the new container answers. Copy it to the
proxy host, for example to /etc/caddy/selvara/maintenance.html, and add this
to the site block:
handle_errors {
@down `{err.status_code} in [502, 503, 504]`
handle @down {
header Cache-Control no-store
header Retry-After 30
root * /etc/caddy/selvara
rewrite * /maintenance.html
file_server {
status 503
}
}
}
handle_errors only fires when Caddy itself cannot reach the app. An error
the app returns — a 500 or the 503 of a degraded health check — reaches the
client unchanged. Agents get the 503 in place of the 502 and retry as
before. file_server's status needs Caddy 2.7 or later.
Next steps
- Put the agent on a host: Agent Installation.
- Configure SMTP so password resets and e-mail notifications work: Configuration.
- Take a backup before you need one: Backup and Restore.