Skip to main content

Troubleshooting

Symptom, check, fix. Commands assume the shipped compose file, so the containers are selvara-app, selvara-db, selvara-redis and selvara-cron.

Start here, always:

docker compose ps
docker compose logs --tail=100 app
curl -s http://127.0.0.1:3000/api/health

/api/health returns 200 when both the database and Redis answer, and 503 with "status":"degraded" when either does not. It tells you which one.

The app will not start​

The entrypoint runs migrations before the server, so a container that never listens on port 3000 has usually failed in that first step. Read the log from the top, not the tail.

docker compose logs app | head -60

DATABASE_URL is not set — the variable is missing from the app's environment. The compose file builds it from POSTGRES_PASSWORD; check that .env exists next to docker-compose.yml and that you ran docker compose up from that directory.

Connection refused / password authentication failed — the app cannot reach PostgreSQL. Check the database container first:

docker compose logs --tail=50 postgres
docker exec -i selvara-db pg_isready -U monitoring -d selvara

If pg_isready is fine but the app is not, the password is the likely cause. PostgreSQL only applies POSTGRES_PASSWORD when it initialises an empty data directory, so changing it in .env after the first start changes what the app sends and not what the server expects. Fix it inside PostgreSQL:

docker exec -i selvara-db psql -U monitoring -d postgres \
-c "ALTER ROLE monitoring WITH PASSWORD 'the-value-in-your-env';"

Migration failed: — the script exited 1 and the server was never started, so the container is restarting in a loop. On PostgreSQL the whole migration run is one transaction, so the schema was rolled back and the previous image still works against it. Read the error, and if you cannot resolve it, pin the previous image tag and restore from your dump — Upgrading.

TimescaleDB setup failed (app will still start): — this one does not stop the app, on purpose. But retention and compression may not be configured, and an unconfigured metrics table grows without bound. See The disk fills up.

The container is healthy but slow to become so — on the very first start against an existing large metrics table, converting it to a hypertable and backfilling the rollups takes minutes. The healthcheck allows 900 seconds for this. The log says backfilling metrics_1h over the last 400 days, this can take a while... while it happens. Let it finish.

Login returns a 500, or the form does nothing​

Almost always AUTH_URL. NextAuth builds its own endpoints from it, and if what it derives does not match the address in the browser, sign-in fails.

Curl the CSRF endpoint through your proxy, using the exact public URL:

curl -si https://<your-domain>/api/auth/csrf

A working instance returns 200 and a JSON body with a csrfToken. A 500, an HTML error page, or a redirect somewhere unexpected means the URL the app thinks it is serving and the one you asked for disagree.

Compare them:

docker exec selvara-app env | grep -E '^(AUTH_URL|NEXTAUTH_URL)='

Both must be the exact public URL — scheme included, no trailing slash, https if the browser uses https. Set them to the same value, then docker compose up -d app.

If /api/auth/csrf is fine and sign-in still fails, check that your proxy forwards cookies and does not strip Set-Cookie, and that the browser is on https if the cookie is marked Secure.

A blank page at / on a brand-new instance is not a fault: with no account in the database, /login redirects to /setup. See Installation.

An agent reports nothing​

The system shows as offline, or as never having reported. Systems are marked OFFLINE when they have not been seen for two minutes, by the same cron run that evaluates alerts — so if every system went offline at once, read Alerts never fire instead. A system whose offline alerts are off shows as Powered off (STANDBY) instead; see Customers and Systems.

Work from the host outwards.

1. Is the agent running?

systemctl status selvara-agent
journalctl -u selvara-agent -n 100 --no-pager

2. Webhook secret. The agent signs a JWT with the webhook_secret in its config.yaml; the dashboard verifies it with its own WEBHOOK_SECRET. A mismatch makes every request 401 and there is nothing on the dashboard side to see, because the request never identifies a system.

docker exec selvara-app env | grep '^WEBHOOK_SECRET='
sudo grep webhook_secret /opt/selvara-agent/config.yaml

They must be identical. If you rotated WEBHOOK_SECRET, every agent's config is now stale — re-run the installer on each host, which fetches a fresh config.yaml. See Agent Installation.

3. Clock skew. The agent's token carries exp five minutes after it is issued. A host whose clock is more than that ahead of or behind the dashboard produces tokens that are already expired or not yet valid, and every request is rejected as an invalid signature. This is the failure that looks like nothing at all: the agent logs a send failure, the dashboard logs nothing.

# on the monitored host
timedatectl
# compare against the dashboard host
date -u

Fix it with NTP (timedatectl set-ntp true) rather than by setting the clock once.

4. Reachability. The agent pushes to the dashboard; nothing connects inward. From the monitored host:

curl -sS -o /dev/null -w '%{http_code}\n' https://<your-domain>/api/health

Anything but 200 is a network, DNS, firewall or proxy problem between that host and the dashboard. An egress firewall that blocks outbound 443 is the common one.

5. Wrong endpoint. Check that webhook_url in config.yaml names your public URL and ends in /api/webhook/metrics. An empty or wrong host there usually means NEXTAUTH_URL was unset when the config was generated — see Configuration.

The same five checks on a Windows host​

A Windows agent is a service, not a unit, and it writes a log file because there is no journal to write to.

1. Is the agent running?

sc query SelvaraAgent
Get-Content C:\ProgramData\Selvara\logs\agent.log -Tail 100
Get-EventLog -LogName Application -Source SelvaraAgent -Newest 20

STOPPED for a service set to start automatically is the whole answer, and the last lines of the log say why. The event log carries four milestones only — started, stopping, an update installed, a configuration it could not read — so a stopped service with nothing in the event log points at something that killed the process rather than at the agent deciding to stop. An endpoint product quarantining the unsigned binary looks exactly like that: the service does not start and the log stays empty. See Agent Installation.

A service that is stopped right after an update, with Update installed, exiting so the service is started again as its last event, is missing its recovery action or its failure flag — see Agent Rollout.

2. Webhook secret and 5. wrong endpoint are the same file at a different path, and reading it needs an elevated shell: the installer restricts it to SYSTEM and the local administrators because it holds the secret.

Select-String webhook C:\ProgramData\Selvara\config.yaml

3. Clock skew is w32tm /query /status, and the fix is still to have the host synchronise rather than to set the clock once.

4. Reachability from the monitored host:

(Invoke-WebRequest https://<your-domain>/api/health -UseBasicParsing).StatusCode

To watch a collection cycle as it happens, stop the service and run the binary with -foreground: it then logs to the console. Start the service again afterwards — a console agent does not restart itself after an update.

An agent stopped after an update and does not come back​

systemctl status selvara-agent says failed (Result: start-limit-hit). systemd restarted the agent several times within seconds and then gave up — the pattern of an update the agent kept installing, as with 2.27.0, whose agent binary reported an older version than it was offered. Nothing on the dashboard can reach an agent that is not running; start it once on the host:

systemctl reset-failed selvara-agent && systemctl start selvara-agent

Units written by an installer from 2.28.1 on set StartLimitIntervalSec=0, so systemd keeps restarting the agent instead. From agent 2.9.1 on, the agent adds the same setting to an older unit itself on its next start, as the drop-in /etc/systemd/system/selvara-agent.service.d/start-limit.conf.

A host goes offline for a few minutes, then catches up​

The system flips to offline and back on its own. The agent logs a pair like this, and its version check and settings fetch time out in between:

Webhook unreachable, buffering reports: … context deadline exceeded
Webhook reachable again, 4 reports waiting

Nothing is lost: the reports buffered in between are sent once the dashboard answers again. context deadline exceeded means the dashboard never answered at all. A rate limit or a rejected token looks different, because it comes back at once with a status code in the log line.

The agent sends everything over one HTTP/2 connection. When a burst of packet loss stalls that connection, the agent notices within about 20 seconds: it pings a connection that has gone quiet and dials a new one when the ping gets no answer. Before 2.4.2 it kept waiting on the stalled connection until the kernel gave up on it, which took two to four minutes. That was long enough to cross the two-minute offline mark, so a few seconds of loss showed up as an outage. Update the agent first if it is older.

If the host still drops out, the loss itself is worth finding. The reverse proxy's access log shows the request that hung. With Caddy's JSON log it has status 0 and a duration in minutes:

zcat -f /var/log/caddy/*.log* | jq -c 'select(.status==0)
| [(.ts|todate), .request.client_ip, .request.uri, .duration]'

If only one host or one network shows up there, the problem is on that side. Look at that host's uplink, at anything that saturates it (a storage rebalance, a backup), and run mtr towards the dashboard while it happens.

A Windows agent reports, but a collector is failing​

Two answers cover most of it.

eventlog failing on the Security channel. Reading the Security event log needs LocalSystem or membership of Event Log Readers. The installed service runs as LocalSystem and has no trouble with it; an agent started by hand in an ordinary console does, and shows as

Security: exit status 5: Zugriff verweigert

with the wording in the host's display language. The other channels are still read and still reported. Judge that collector from the installed service rather than from a hand-run binary; if you do want to run it by hand, put the account you run it as into Event Log Readers first.

A collector reporting nothing just after a start. Defender, the firewall, the scheduled tasks, the shares, BitLocker and the Windows Update count each run on their own schedule beside the collection, so the first report after a start carries nothing for them. Give it a cycle.

A host that never reports one of them is a different matter, and usually not a fault either: defender, firewall and bitlocker report nothing and raise no error on an edition, or a Server Core install, that does not have the feature behind them. Which services a Windows host does and does not detect, and the three things that do not carry over from a Linux host, are in Agent Collectors.

Alerts never fire​

Alert evaluation is not driven by a timer inside the application. Nothing in the app wakes up on its own. Every alert that is raised, every alert that is resolved, every system marked offline and every notification sent is the result of an HTTP call to:

POST /api/cron/evaluate-alerts

In the shipped compose file the selvara-cron container makes that call once a minute. If that container is not running, alerts never fire and no system ever goes offline — while the dashboard keeps showing fresh metrics, which is why this can go unnoticed for a long time.

docker compose ps cron
docker logs --tail=50 selvara-cron

The container writes a small script, installs a one-line crontab and execs crond in the foreground. If its log shows nothing at all about running the job, it is up but the schedule is not firing.

Trigger one by hand and read what comes back:

docker exec selvara-app sh -c \
'wget -qO- --header="Authorization: Bearer $CRON_SECRET" \
http://localhost:3000/api/cron/evaluate-alerts'

The quoting matters: $CRON_SECRET has to expand inside the container, where it is set, not in your own shell.

{"success":true,"timestamp":"...","alerts":{"evaluated":36,"triggered":0,"resolved":0},"health":{"marked_offline":0},"clusters":{...}}
  • {"error":"Unauthorized"} with 401 — the token does not match the app's CRON_SECRET. The cron container and the app container must be given the same value; the compose file does this from one variable.
  • {"success":false,"error":"Evaluation failed"} with 500 — the run threw. docker compose logs app | grep '\[Cron\]' has the reason.
  • A response, but "evaluated":0 — there are no enabled alert rules. See Alert Rules.
  • A response, alerts evaluated, but nothing triggers when it should — check whether a maintenance window is open, or whether the customer has been switched inactive. Both suppress alerting on purpose and both are visible in the UI. See Maintenance Windows and Customers and Systems.

The cron container starts only once the app reports healthy (depends_on: service_healthy). An app that never becomes healthy therefore also has no alerting at all, even though it is serving pages.

If alerts fire but nobody hears about them, the problem is delivery, not evaluation — see Notifications.

The disk fills up​

Metric history is the only thing here that grows on its own, and it grows in the indexes as much as in the rows.

Find out where it went. scripts/db-usage.sql is read-only and safe to run on a full disk:

docker exec -i selvara-db psql -U monitoring -d selvara < scripts/db-usage.sql

Read three things in its output:

  • Is metrics a hypertable? — an empty result means it is a plain table. Nothing is ever dropped from a plain table, and this is the root cause of every runaway case.
  • Retention / compression policies — an empty result means nothing is ever deleted.
  • Largest relations — if the index column dwarfs the heap, indexes are the problem, not the samples.

If retention is simply not aggressive enough, shorten it. The METRICS_RAW_RETENTION, METRICS_5M_RETENTION and METRICS_1H_RETENTION variables are in Configuration, together with the caveat that an existing policy is not rewritten just because the variable changed.

If the table was never converted, or the disk is already full enough that ordinary measures will not run, use scripts/emergency-reclaim.sh. It only uses operations that hand files back to the operating system immediately — DROP INDEX, DROP TABLE, CHECKPOINT — because DELETE on a full disk is the worst available move: it writes WAL, leaves dead tuples, and returns nothing until a VACUUM FULL that needs a second copy of the table.

# prints the plan, changes nothing
./scripts/emergency-reclaim.sh

# actually do it
./scripts/emergency-reclaim.sh --apply

What it does: stops selvara-app and selvara-cron but leaves PostgreSQL running; if metrics is already a hypertable it just drops old chunks and stops there; otherwise it drops the redundant indexes, rebuilds metrics with the last KEEP_HOURS at full resolution plus everything older downsampled to one sample per hour, and restarts the app. History is not thrown away by default. Two knobs:

KEEP_HOURS=72 ./scripts/emergency-reclaim.sh --apply # keep 3 days full-res
DOWNSAMPLE=0 ./scripts/emergency-reclaim.sh --apply # discard old data

Run db-usage.sql from the same directory or copy it next to the script — it looks for it beside itself and skips the report silently if it is not there.

Afterwards, start the current image. Its migration converts metrics into a hypertable, builds the rollups from the history you kept, and installs the retention policies. Do not start an older image against the reclaimed database: it will start refilling the table with no retention at all.

There is one refusal worth knowing about. If metrics is still a plain table and larger than METRICS_MAX_MIGRATE_BYTES (20 GiB by default), the migration prints

!! metrics is 41.2 GB and is NOT a hypertable.
!! Converting it in place would rewrite the whole table and needs
!! that much free disk again. Refusing.
!! Run scripts/emergency-reclaim.sh first, then restart the app.

and skips the conversion. That is the safety rail doing its job, not a bug. Reclaim first.

I have lost the admin password​

If SMTP is configured, use the reset link on the login screen. The "forgot password" link only appears when SMTP is enabled, and POST /api/auth/password/forgot answers 503 Mail is not configured when it is not.

Otherwise, set a new one with scripts/reset-password.mjs. It writes the same bcrypt hash the application writes into users.password_hash, and it needs DATABASE_URL and nothing else. That is also why it is no way around the login: anyone who can run it can already read and write the database by hand.

The script ships in the image beside the migration and seed scripts, so it is already in the container.

NODE_PATH points at the pg and bcryptjs the scripts use. The entrypoint sets it for its own run; docker exec starts a fresh environment, so pass it in. Begin by listing the accounts — the address you remember and the address in the database are not always the same:

docker exec -e NODE_PATH=/app/scripts/node_modules selvara-app \
node scripts/reset-password.mjs --list
2 account(s):
ADMIN you@example.com - authenticator enrolled
VIEWER ops@example.com ops no authenticator

Then set the password. -it matters: the script asks for it on the terminal and does not accept it as an argument, so it stays out of your shell history and out of the process list.

docker exec -it -e NODE_PATH=/app/scripts/node_modules selvara-app \
node scripts/reset-password.mjs you@example.com
New password:
Repeat password:
Password updated for you@example.com, role ADMIN

The account is matched the way the login form matches it: a value containing @ against the e-mail address, anything else against the username, both without regard to case. The password must pass the same rule as one set in the interface (Passwords); the script says why when it refuses one. Where there is no terminal, the first line of standard input is used instead:

printf '%s\n' "$NEW_PASSWORD" | docker exec -i \
-e NODE_PATH=/app/scripts/node_modules selvara-app \
node scripts/reset-password.mjs you@example.com

On a host that has the repository checked out and its dependencies installed, running it against the published database port does the same job — useful when the container itself will not start:

DATABASE_URL=postgresql://monitoring:<POSTGRES_PASSWORD>@127.0.0.1:5432/selvara \
node scripts/reset-password.mjs you@example.com

Two follow-ups.

If the account has an authenticator app enrolled, the new password alone will not get you in — sign-in still asks for a code, and the script says so when it finds one. --clear-2fa unenrols it as part of the reset. It is never the default, because dropping a second factor is not something a password reset should decide on its own:

docker exec -it -e NODE_PATH=/app/scripts/node_modules selvara-app \
node scripts/reset-password.mjs you@example.com --clear-2fa

With SMTP configured the account then falls back to e-mailed codes; without it the password alone signs in. Enrol an authenticator again under Profile → Security afterwards.

And if no account is left with the ADMIN role, promote one. The script sets passwords, not roles:

docker exec -i selvara-db psql -U monitoring -d selvara -v ON_ERROR_STOP=1 -c \
"UPDATE users SET role = 'ADMIN' WHERE email = 'you@<your-domain>';"

If the users table is empty — every account deleted — the setup screen opens again by itself, because it is gated on that table being empty. Visiting / will take you to /setup. See Installation.