Troubleshooting
Symptom, check, fix. Commands assume the shipped compose file, so the
containers are selvara-app, selvara-db, selvara-redis and
selvara-cron.
Start here, always:
docker compose ps
docker compose logs --tail=100 app
curl -s http://127.0.0.1:3000/api/health
/api/health returns 200 when both the database and Redis answer, and 503 with
"status":"degraded" when either does not. It tells you which one.
The app will not start
The entrypoint runs migrations before the server, so a container that never listens on port 3000 has usually failed in that first step. Read the log from the top, not the tail.
docker compose logs app | head -60
DATABASE_URL is not set — the variable is missing from the app's
environment. The compose file builds it from POSTGRES_PASSWORD; check that
.env exists next to docker-compose.yml and that you ran docker compose up
from that directory.
Connection refused / password authentication failed — the app cannot reach PostgreSQL. Check the database container first:
docker compose logs --tail=50 postgres
docker exec -i selvara-db pg_isready -U monitoring -d selvara
If pg_isready is fine but the app is not, the password is the likely cause.
PostgreSQL only applies POSTGRES_PASSWORD when it initialises an empty data
directory, so changing it in .env after the first start changes what the app
sends and not what the server expects. Fix it inside PostgreSQL:
docker exec -i selvara-db psql -U monitoring -d postgres \
-c "ALTER ROLE monitoring WITH PASSWORD 'the-value-in-your-env';"
Migration failed: — the script exited 1 and the server was never
started, so the container is restarting in a loop. On PostgreSQL the whole
migration run is one transaction, so the schema was rolled back and the
previous image still works against it. Read the error, and if you cannot
resolve it, pin the previous image tag and restore from your dump —
Upgrading.
TimescaleDB setup failed (app will still start): — this one does not
stop the app, on purpose. But retention and compression may not be configured,
and an unconfigured metrics table grows without bound. See
The disk fills up.
The container is healthy but slow to become so — on the very first start
against an existing large metrics table, converting it to a hypertable and
backfilling the rollups takes minutes. The healthcheck allows 900 seconds for
this. The log says backfilling metrics_1h over the last 400 days, this can take a while... while it happens. Let it finish.
Login returns a 500, or the form does nothing
Almost always AUTH_URL. NextAuth builds its own endpoints from it, and if
what it derives does not match the address in the browser, sign-in fails.
Curl the CSRF endpoint through your proxy, using the exact public URL:
curl -si https://<your-domain>/api/auth/csrf
A working instance returns 200 and a JSON body with a csrfToken. A 500, an
HTML error page, or a redirect somewhere unexpected means the URL the app
thinks it is serving and the one you asked for disagree.
Compare them:
docker exec selvara-app env | grep -E '^(AUTH_URL|NEXTAUTH_URL)='
Both must be the exact public URL — scheme included, no trailing slash,
https if the browser uses https. Set them to the same value, then
docker compose up -d app.
If /api/auth/csrf is fine and sign-in still fails, check that your proxy
forwards cookies and does not strip Set-Cookie, and that the browser is on
https if the cookie is marked Secure.
A blank page at / on a brand-new instance is not a fault: with no account in
the database, /login redirects to /setup. See
Installation.
An agent reports nothing
The system shows as offline, or as never having reported. Systems are marked
OFFLINE when they have not been seen for two minutes, by the same cron run
that evaluates alerts — so if every system went offline at once, read
Alerts never fire instead. A system whose offline
alerts are off shows as Powered off (STANDBY) instead; see
Customers and Systems.
Work from the host outwards.
1. Is the agent running?
systemctl status selvara-agent
journalctl -u selvara-agent -n 100 --no-pager
2. Webhook secret. The agent signs a JWT with the webhook_secret in its
config.yaml; the dashboard verifies it with its own WEBHOOK_SECRET. A
mismatch makes every request 401 and there is nothing on the dashboard side to
see, because the request never identifies a system.
docker exec selvara-app env | grep '^WEBHOOK_SECRET='
sudo grep webhook_secret /opt/selvara-agent/config.yaml
They must be identical. If you rotated WEBHOOK_SECRET, every agent's config
is now stale — re-run the installer on each host, which fetches a fresh
config.yaml. See Agent Installation.
3. Clock skew. The agent's token carries exp five minutes after it is
issued. A host whose clock is more than that ahead of or behind the dashboard
produces tokens that are already expired or not yet valid, and every request
is rejected as an invalid signature. This is the failure that looks like
nothing at all: the agent logs a send failure, the dashboard logs nothing.
# on the monitored host
timedatectl
# compare against the dashboard host
date -u
Fix it with NTP (timedatectl set-ntp true) rather than by setting the clock
once.
4. Reachability. The agent pushes to the dashboard; nothing connects inward. From the monitored host:
curl -sS -o /dev/null -w '%{http_code}\n' https://<your-domain>/api/health
Anything but 200 is a network, DNS, firewall or proxy problem between that
host and the dashboard. An egress firewall that blocks outbound 443 is the
common one.
5. Wrong endpoint. Check that webhook_url in config.yaml names your
public URL and ends in /api/webhook/metrics. An empty or wrong host there
usually means NEXTAUTH_URL was unset when the config was generated — see
Configuration.
The same five checks on a Windows host
A Windows agent is a service, not a unit, and it writes a log file because there is no journal to write to.
1. Is the agent running?
sc query SelvaraAgent
Get-Content C:\ProgramData\Selvara\logs\agent.log -Tail 100
Get-EventLog -LogName Application -Source SelvaraAgent -Newest 20
STOPPED for a service set to start automatically is the whole answer, and
the last lines of the log say why. The event log carries four milestones only
— started, stopping, an update installed, a configuration it could not read —
so a stopped service with nothing in the event log points at something
that killed the process rather than at the agent deciding to stop. An endpoint
product quarantining the unsigned binary looks exactly like that: the service
does not start and the log stays empty. See
Agent Installation.
A service that is stopped right after an update, with Update installed, exiting so the service is started again as its last event, is missing its recovery action or its failure flag — see Agent Rollout.
2. Webhook secret and 5. wrong endpoint are the same file at a different path, and reading it needs an elevated shell: the installer restricts it to SYSTEM and the local administrators because it holds the secret.
Select-String webhook C:\ProgramData\Selvara\config.yaml
3. Clock skew is w32tm /query /status, and the fix is still to have the
host synchronise rather than to set the clock once.
4. Reachability from the monitored host:
(Invoke-WebRequest https://<your-domain>/api/health -UseBasicParsing).StatusCode
To watch a collection cycle as it happens, stop the service and run the binary
with -foreground: it then logs to the console. Start the service again
afterwards — a console agent does not restart itself after an update.
An agent stopped after an update and does not come back
systemctl status selvara-agent says failed (Result: start-limit-hit).
systemd restarted the agent several times within seconds and then gave up —
the pattern of an update the agent kept installing, as with 2.27.0, whose
agent binary reported an older version than it was offered. Nothing on the
dashboard can reach an agent that is not running; start it once on the host:
systemctl reset-failed selvara-agent && systemctl start selvara-agent
Units written by an installer from 2.28.1 on set StartLimitIntervalSec=0,
so systemd keeps restarting the agent instead. From agent 2.9.1 on, the agent
adds the same setting to an older unit itself on its next start, as the
drop-in /etc/systemd/system/selvara-agent.service.d/start-limit.conf.
A host goes offline for a few minutes, then catches up
The system flips to offline and back on its own. The agent logs a pair like this, and its version check and settings fetch time out in between:
Webhook unreachable, buffering reports: … context deadline exceeded
Webhook reachable again, 4 reports waiting
Nothing is lost: the reports buffered in between are sent once the dashboard
answers again. context deadline exceeded means the dashboard never answered
at all. A rate limit or a rejected token looks different, because it comes
back at once with a status code in the log line.
The agent sends everything over one HTTP/2 connection. When a burst of packet loss stalls that connection, the agent notices within about 20 seconds: it pings a connection that has gone quiet and dials a new one when the ping gets no answer. Before 2.4.2 it kept waiting on the stalled connection until the kernel gave up on it, which took two to four minutes. That was long enough to cross the two-minute offline mark, so a few seconds of loss showed up as an outage. Update the agent first if it is older.
If the host still drops out, the loss itself is worth finding. The reverse
proxy's access log shows the request that hung. With Caddy's JSON log it has
status 0 and a duration in minutes:
zcat -f /var/log/caddy/*.log* | jq -c 'select(.status==0)
| [(.ts|todate), .request.client_ip, .request.uri, .duration]'
If only one host or one network shows up there, the problem is on that side.
Look at that host's uplink, at anything that saturates it (a storage
rebalance, a backup), and run mtr towards the dashboard while it happens.
A Windows agent reports, but a collector is failing
Two answers cover most of it.
eventlog failing on the Security channel. Reading the Security event
log needs LocalSystem or membership of Event Log Readers. The installed
service runs as LocalSystem and has no trouble with it; an agent started by
hand in an ordinary console does, and shows as
Security: exit status 5: Zugriff verweigert
with the wording in the host's display language. The other channels are still read and still reported. Judge that collector from the installed service rather than from a hand-run binary; if you do want to run it by hand, put the account you run it as into Event Log Readers first.
A collector reporting nothing just after a start. Defender, the firewall, the scheduled tasks, the shares, BitLocker and the Windows Update count each run on their own schedule beside the collection, so the first report after a start carries nothing for them. Give it a cycle.
A host that never reports one of them is a different matter, and usually not
a fault either: defender, firewall and bitlocker report nothing and
raise no error on an edition, or a Server Core install, that does not have the
feature behind them. Which services a Windows host does and does not detect,
and the three things that do not carry over from a Linux host, are in
Agent Collectors.
Alerts never fire
Alert evaluation is not driven by a timer inside the application. Nothing in the app wakes up on its own. Every alert that is raised, every alert that is resolved, every system marked offline and every notification sent is the result of an HTTP call to:
POST /api/cron/evaluate-alerts
In the shipped compose file the selvara-cron container makes that call once a
minute. If that container is not running, alerts never fire and no system ever
goes offline — while the dashboard keeps showing fresh metrics, which is why
this can go unnoticed for a long time.
docker compose ps cron
docker logs --tail=50 selvara-cron
The container writes a small script, installs a one-line crontab and execs
crond in the foreground. If its log shows nothing at all about running the
job, it is up but the schedule is not firing.
Trigger one by hand and read what comes back:
docker exec selvara-app sh -c \
'wget -qO- --header="Authorization: Bearer $CRON_SECRET" \
http://localhost:3000/api/cron/evaluate-alerts'
The quoting matters: $CRON_SECRET has to expand inside the container, where
it is set, not in your own shell.
{"success":true,"timestamp":"...","alerts":{"evaluated":36,"triggered":0,"resolved":0},"health":{"marked_offline":0},"clusters":{...}}
{"error":"Unauthorized"}with 401 — the token does not match the app'sCRON_SECRET. Thecroncontainer and theappcontainer must be given the same value; the compose file does this from one variable.{"success":false,"error":"Evaluation failed"}with 500 — the run threw.docker compose logs app | grep '\[Cron\]'has the reason.- A response, but
"evaluated":0— there are no enabled alert rules. See Alert Rules. - A response, alerts evaluated, but nothing triggers when it should — check whether a maintenance window is open, or whether the customer has been switched inactive. Both suppress alerting on purpose and both are visible in the UI. See Maintenance Windows and Customers and Systems.
The cron container starts only once the app reports healthy
(depends_on: service_healthy). An app that never becomes healthy therefore
also has no alerting at all, even though it is serving pages.
If alerts fire but nobody hears about them, the problem is delivery, not evaluation — see Notifications.
The disk fills up
Metric history is the only thing here that grows on its own, and it grows in the indexes as much as in the rows.
Find out where it went. scripts/db-usage.sql is read-only and safe to run
on a full disk:
docker exec -i selvara-db psql -U monitoring -d selvara < scripts/db-usage.sql
Read three things in its output:
- Is metrics a hypertable? — an empty result means it is a plain table. Nothing is ever dropped from a plain table, and this is the root cause of every runaway case.
- Retention / compression policies — an empty result means nothing is ever deleted.
- Largest relations — if the index column dwarfs the heap, indexes are the problem, not the samples.
If retention is simply not aggressive enough, shorten it. The
METRICS_RAW_RETENTION, METRICS_5M_RETENTION and METRICS_1H_RETENTION
variables are in Configuration, together with the caveat that
an existing policy is not rewritten just because the variable changed.
If the table was never converted, or the disk is already full enough that
ordinary measures will not run, use scripts/emergency-reclaim.sh. It only
uses operations that hand files back to the operating system immediately —
DROP INDEX, DROP TABLE, CHECKPOINT — because DELETE on a full disk is
the worst available move: it writes WAL, leaves dead tuples, and returns
nothing until a VACUUM FULL that needs a second copy of the table.
# prints the plan, changes nothing
./scripts/emergency-reclaim.sh
# actually do it
./scripts/emergency-reclaim.sh --apply
What it does: stops selvara-app and selvara-cron but leaves PostgreSQL
running; if metrics is already a hypertable it just drops old chunks and
stops there; otherwise it drops the redundant indexes, rebuilds metrics with
the last KEEP_HOURS at full resolution plus everything older downsampled to
one sample per hour, and restarts the app. History is not thrown away by
default. Two knobs:
KEEP_HOURS=72 ./scripts/emergency-reclaim.sh --apply # keep 3 days full-res
DOWNSAMPLE=0 ./scripts/emergency-reclaim.sh --apply # discard old data
Run db-usage.sql from the same directory or copy it next to the script — it
looks for it beside itself and skips the report silently if it is not there.
Afterwards, start the current image. Its migration converts metrics into a
hypertable, builds the rollups from the history you kept, and installs the
retention policies. Do not start an older image against the reclaimed
database: it will start refilling the table with no retention at all.
There is one refusal worth knowing about. If metrics is still a plain table
and larger than METRICS_MAX_MIGRATE_BYTES (20 GiB by default), the migration
prints
!! metrics is 41.2 GB and is NOT a hypertable.
!! Converting it in place would rewrite the whole table and needs
!! that much free disk again. Refusing.
!! Run scripts/emergency-reclaim.sh first, then restart the app.
and skips the conversion. That is the safety rail doing its job, not a bug. Reclaim first.
I have lost the admin password
If SMTP is configured, use the reset link on the login screen. The
"forgot password" link only appears when SMTP is enabled, and
POST /api/auth/password/forgot answers 503 Mail is not configured when it
is not.
Otherwise, set a new one with scripts/reset-password.mjs. It writes the
same bcrypt hash the application writes into users.password_hash, and it
needs DATABASE_URL and nothing else. That is also why it is no way around the
login: anyone who can run it can already read and write the database by hand.
The script ships in the image beside the migration and seed scripts, so it is already in the container.
NODE_PATH points at the pg and bcryptjs the scripts use. The entrypoint
sets it for its own run; docker exec starts a fresh environment, so pass it
in. Begin by listing the accounts — the address you remember and the address in
the database are not always the same:
docker exec -e NODE_PATH=/app/scripts/node_modules selvara-app \
node scripts/reset-password.mjs --list
2 account(s):
ADMIN you@example.com - authenticator enrolled
VIEWER ops@example.com ops no authenticator
Then set the password. -it matters: the script asks for it on the terminal
and does not accept it as an argument, so it stays out of your shell history
and out of the process list.
docker exec -it -e NODE_PATH=/app/scripts/node_modules selvara-app \
node scripts/reset-password.mjs you@example.com
New password:
Repeat password:
Password updated for you@example.com, role ADMIN
The account is matched the way the login form matches it: a value containing
@ against the e-mail address, anything else against the username, both
without regard to case. The password must pass the same rule as one set in
the interface (Passwords); the script says
why when it refuses one. Where there is no terminal, the first line of standard input
is used instead:
printf '%s\n' "$NEW_PASSWORD" | docker exec -i \
-e NODE_PATH=/app/scripts/node_modules selvara-app \
node scripts/reset-password.mjs you@example.com
On a host that has the repository checked out and its dependencies installed, running it against the published database port does the same job — useful when the container itself will not start:
DATABASE_URL=postgresql://monitoring:<POSTGRES_PASSWORD>@127.0.0.1:5432/selvara \
node scripts/reset-password.mjs you@example.com
Two follow-ups.
If the account has an authenticator app enrolled, the new password alone will
not get you in — sign-in still asks for a code, and the script says so when it
finds one. --clear-2fa unenrols it as part of the reset. It is never the
default, because dropping a second factor is not something a password reset
should decide on its own:
docker exec -it -e NODE_PATH=/app/scripts/node_modules selvara-app \
node scripts/reset-password.mjs you@example.com --clear-2fa
With SMTP configured the account then falls back to e-mailed codes; without it the password alone signs in. Enrol an authenticator again under Profile → Security afterwards.
And if no account is left with the ADMIN role, promote one. The script sets
passwords, not roles:
docker exec -i selvara-db psql -U monitoring -d selvara -v ON_ERROR_STOP=1 -c \
"UPDATE users SET role = 'ADMIN' WHERE email = 'you@<your-domain>';"
If the users table is empty — every account deleted — the setup screen opens
again by itself, because it is gated on that table being empty. Visiting /
will take you to /setup. See Installation.