Skip to main content

Backup and Restore

There is one thing to back up: the PostgreSQL database. Everything Selvara knows is in there. There is also one thing that is not in it and that a dump is useless without, so read the first two sections before you write a cron job.

What has to be backed up​

The database holds all of it:

  • Accounts — users, password hashes, usernames, roles, TOTP secrets (encrypted), trusted devices, pending password-reset tokens.
  • Your inventory — customers, systems, clusters, and each system's region field and tags.
  • Agent credentials — every system's agent key and its hash. Lose these and every agent has to be re-keyed on its host.
  • Everything configured in the dashboard — SMTP settings including the encrypted password, the agent settings catalogue with its encrypted secrets, the agent rollout mode.
  • Rules and state — alert rules, event rules, maintenance windows, open and historical alerts, the event timeline and notification preferences.
  • Metrics — the metrics hypertable and the metrics_5m / metrics_1h rollups.

What you do not need to back up:

  • Redis. It holds the pub/sub stream and rate-limit counters. Losing it costs you a few seconds of live updates and resets some counters.
  • The application image and the compiled agent binaries. The image is rebuilt from the repository, and the agent binaries are compiled into it.
  • Anything on the monitored hosts. An agent's config.yaml is regenerated from the dashboard by re-running the installer.

The thing that is not in the dump​

The encryption key. Secrets in the database are encrypted with a key derived from SETTINGS_ENCRYPTION_KEY, or from AUTH_SECRET when that is unset, and neither is stored in the database. A dump restored with a different key gives you back rows you cannot read: SMTP password, agent secrets, every user's TOTP secret.

Back up the .env file alongside every dump, and keep it somewhere the dump is not. Configuration explains the fallback and what breaks when the key changes.

Taking a dump​

Use the custom format (-Fc). It restores selectively, compresses, and is what pg_restore wants.

docker exec -i selvara-db pg_dump -U monitoring -Fc selvara \
> selvara-$(date +%F-%H%M).dump

Do not add -t to docker exec. It allocates a pseudo-TTY, which translates line endings in the stream and hands you a corrupt archive that only fails when you try to restore it.

Check the result before you trust it:

ls -lh selvara-*.dump
pg_restore --list selvara-2026-01-01-0300.dump | head

pg_restore --list reading the table of contents is the cheapest proof that the file is a valid archive and not a truncated pipe.

If pg_restore is not installed on the host, run it through the database image:

docker run --rm -i -v "$PWD:/b" timescale/timescaledb:latest-pg16 \
pg_restore --list /b/selvara-2026-01-01-0300.dump | head

The TimescaleDB trap​

This is not an ordinary PostgreSQL database and an ordinary restore will not give you back what you dumped.

metrics is a hypertable. The rows do not live in a table called metrics — they live in chunk tables under the _timescaledb_internal schema, and metrics is a view-like parent over them. metrics_5m and metrics_1h are continuous aggregates, which are materialised views plus catalog entries plus their own hypertables plus background jobs. All of that is in the dump, and restoring it plainly makes PostgreSQL replay catalog rows and background-job definitions in an order TimescaleDB is not expecting.

TimescaleDB ships two functions for exactly this. They put the extension into a restoring mode that suspends its background workers and relaxes the checks that would otherwise reject the incoming catalog rows:

SELECT timescaledb_pre_restore();
-- run pg_restore here
SELECT timescaledb_post_restore();

timescaledb_pre_restore() changes a database-level setting, so it stays in effect for sessions that connect afterwards — pg_restore is a separate connection and is still covered. Do not skip timescaledb_post_restore(). Until you run it the background workers stay off, which means no retention, no compression and no rollup refresh, on a database that otherwise looks fine.

Two more conditions on the target:

  • The same PostgreSQL major version. The shipped image is pg16.
  • A TimescaleDB version at least as new as the one the dump was taken with, and the extension created before the restore, not by it.

Dumping without the metric history​

Metric history is almost all of the volume and the least valuable part of it — a restored instance without it has every account, customer, system, rule and agent key, and simply starts its charts from empty.

The obvious exclusion does not do what it looks like:

# Does NOT exclude the metric rows.
pg_dump -U monitoring -Fc --exclude-table-data='metrics*' selvara

That pattern matches the names in the public schema. The rows are in the chunks, so they come along anyway. Exclude the chunks too:

docker exec -i selvara-db pg_dump -U monitoring -Fc \
--exclude-table-data='public.metrics*' \
--exclude-table-data='_timescaledb_internal.*' \
selvara > selvara-config-$(date +%F).dump

Compare the sizes of the two dumps on your own instance before you rely on this. If the "config" dump is not dramatically smaller than the full one, the exclusion did not match and you should find out why rather than assume.

What it costs: every chart is empty until agents have reported for long enough to fill one. The current values that stat cards and health badges read come from metric_series, a separate table with one row per series that the metrics* pattern does not match, so those survive — the restored instance looks current and has no past. Alerts that average over a duration window need that window to pass before they can fire again, which with the shipped rules is between 30 seconds and five minutes.

Keep both kinds. A small config dump often, a full dump less often.

Restoring​

Stop the writers first. PostgreSQL stays up; nothing else should be touching the database.

docker compose stop app cron

Restore into a freshly created, empty database. Restoring over a database that already has Selvara's schema in it is how you get half-applied objects and a confusing set of errors.

# 1. Create an empty database with the extension already present.
docker exec -i selvara-db psql -U monitoring -d postgres -v ON_ERROR_STOP=1 \
-c "CREATE DATABASE selvara_restore OWNER monitoring;"
docker exec -i selvara-db psql -U monitoring -d selvara_restore -v ON_ERROR_STOP=1 \
-c "CREATE EXTENSION IF NOT EXISTS timescaledb;"

# 2. Enter restore mode.
docker exec -i selvara-db psql -U monitoring -d selvara_restore -v ON_ERROR_STOP=1 \
-c "SELECT timescaledb_pre_restore();"

# 3. Restore.
docker exec -i selvara-db pg_restore -U monitoring -d selvara_restore \
--no-owner --no-privileges < selvara-2026-01-01-0300.dump

# 4. Leave restore mode. Not optional.
docker exec -i selvara-db psql -U monitoring -d selvara_restore -v ON_ERROR_STOP=1 \
-c "SELECT timescaledb_post_restore();"

Then point the instance at it. Either rename the databases, or change the database name in DATABASE_URL and restart:

docker compose up -d
docker compose logs -f app
curl -s http://127.0.0.1:3000/api/health

The app runs its migrations against the restored database on start, so a dump taken from an older version is brought forward automatically. A dump taken from a newer version is not brought backward — see Upgrading.

Verifying that a backup is restorable​

A dump you have never restored is a hypothesis. Test it somewhere that is not production, on a schedule you will actually keep.

docker run -d --name selvara-restore-test \
-e POSTGRES_USER=monitoring \
-e POSTGRES_PASSWORD=throwaway \
-e POSTGRES_DB=selvara \
timescale/timescaledb:latest-pg16

# wait for it to accept connections
docker exec -i selvara-restore-test pg_isready -U monitoring -d selvara

docker exec -i selvara-restore-test psql -U monitoring -d selvara -v ON_ERROR_STOP=1 \
-c "CREATE EXTENSION IF NOT EXISTS timescaledb;" \
-c "SELECT timescaledb_pre_restore();"
docker exec -i selvara-restore-test pg_restore -U monitoring -d selvara \
--no-owner --no-privileges < selvara-2026-01-01-0300.dump
docker exec -i selvara-restore-test psql -U monitoring -d selvara -v ON_ERROR_STOP=1 \
-c "SELECT timescaledb_post_restore();"

Then check that it is the database you think it is, not merely a database:

docker exec -i selvara-restore-test psql -U monitoring -d selvara -c "
SELECT (SELECT count(*) FROM users) AS users,
(SELECT count(*) FROM customers) AS customers,
(SELECT count(*) FROM systems) AS systems,
(SELECT count(*) FROM alert_rules) AS alert_rules,
(SELECT count(*) FROM app_settings) AS settings;
SELECT hypertable_name FROM timescaledb_information.hypertables;
SELECT view_name FROM timescaledb_information.continuous_aggregates;"

metrics must appear as a hypertable and both metrics_5m and metrics_1h as continuous aggregates. If they came back as plain tables and plain views, the restore silently degraded and the retention policies are not there either — which on a live instance means the metrics table grows until the disk is full.

Clean up with docker rm -f selvara-restore-test.

Retention makes old backups different from the live database​

Metric history is pruned on a schedule, set by the METRICS_* variables and described in Configuration. With the shipped defaults:

DataKept for
Raw samples in metrics7 days (compressed after 2)
metrics_5m rollup30 days
metrics_1h rollup400 days

Two consequences.

A backup from three months ago contains raw samples the live database dropped eleven weeks ago. That is occasionally exactly what you want — the only copy of full-resolution data from an incident. Keep an offline dump of an interesting week if the detail matters to you; nothing else will.

And restoring that backup does not keep that history. The retention policies come back with it and run again on their normal schedule, so raw samples older than the raw retention are dropped shortly after the restore. If you restored specifically to look at old raw data, copy the rows you need out first, or remove the retention policy on the restored copy before you start it up:

SELECT remove_retention_policy('metrics', if_exists => true);

Moving to a new server​

A restore onto a new machine gives you the dashboard back; the agents still report to the old address in their config.yaml. Rather than re-running the installer on every host, the old server can tell its agents where to go, and each agent moves only once the new server has proved it knows that host.

How the move works​

An administrator enters the new server's address on the old server, under Agents → Server move, in the card Agents' server address. The address is scheme, host and optional port — https://monitoring.example.com — with no path. It is stored as the instance setting agent_server_url and can also be read and written through GET and PUT /api/agents/server-url, administrators only.

Saving it bumps every system's configuration version, so every agent re-reads its settings with its next report, and the settings answer then carries the new address. An agent that sees an address other than the server it reports to:

  1. Checks the new server by fetching its own settings there, signed with its own credentials. Only an HTTP 200 counts. That is why the new server must run a restored copy of this database with the same secrets: a server that does not know the system id, or verifies with a different WEBHOOK_SECRET, turns the agent away.
  2. On success writes the new address to a file named server-url next to its binary — /opt/selvara-agent/server-url on Linux, C:\Program Files\Selvara\server-url on Windows — records an agent.server_switched event on the new server, and restarts. From then on that file outranks the host in webhook_url, and from agent 2.8.0 the restart also writes the new address into config.yaml. See Agent Installation.
  3. On failure stays where it is and records an agent.server_switch_failed warning with the reason on the old server's timeline, then tries again, at most every five minutes.

While the address is set, the card lists the agents that have not arrived: each with the address it reports through, its agent version, when it was last seen and why it has not moved — too old, held back from updates, the reason its last attempt failed, or simply waiting for its next report. Cancel the move clears the address; agents that already moved stay moved.

Agents older than 2.4.0 ignore the setting and simply keep reporting where they are. The list marks them; update them first.

Above the list, the card always shows which addresses agents reported through in the last 24 hours, with a count each — the dashboard records the address of every report, as the reverse proxy passed it on.

The procedure​

  1. Update every agent first. The move needs the current agent version everywhere. The Agents page must show no outdated agent — see Agent Rollout. A host still on an older version will not move and has to be re-installed against the new server by hand.

  2. Set up the new server as in Installation, restore the latest dump into it as described in Restoring, and start it with the old server's .env, changing only the address:

    • AUTH_URL and NEXTAUTH_URL — the new server's public address.
    • WEBHOOK_SECRET — the same as on the old server. Agents sign every request with it; a different value fails the agent's check of the new server and no agent moves.
    • SETTINGS_ENCRYPTION_KEY — the same, or, where it is unset, AUTH_SECRET — the same, because that is then the key. Otherwise the SMTP password, the agent settings' secrets and every TOTP secret cannot be read — see The thing that is not in the dump.
    • AUTH_SECRET — the same, even with a separate encryption key, so that sessions and trusted devices stay valid.
    • CRON_SECRET — whatever the new server's cron container is given; copying the .env keeps the two in step.

    Check it answers: curl -s https://<new-server>/api/health.

  3. On the old server, open Agents → Server move, enter the new server's address and confirm.

  4. Watch the agents move. Within one or two report intervals each agent goes quiet on the old server and appears on the new one, with agent.server_switched on its timeline there. An agent that cannot move leaves agent.server_switch_failed on the old server's timeline, with the reason.

  5. Keep the restore and the switch close together. Metrics, alerts and events that reach the old server between the dump and the moment an agent moves are not on the new server. Take the dump, restore it and set the address in one sitting.

  6. Switch the old server off once the card says no agent reports there any more.

Renaming this server​

The same mechanism moves agents to a new name of the same server — a new domain in front of an unchanged container and database:

  1. Point the new name at the server and add it to the reverse proxy next to the old one, so both answer.
  2. Set AUTH_URL and NEXTAUTH_URL in .env to the new address and recreate the app container. Install commands, agent downloads and updates use that address from then on.
  3. Under Agents → Server move, enter the new address — this server's own. The card calls it a rename and lists every agent that still reports through another address, online or not: an agent that is switched off today moves when it next reports.
  4. When the card says every agent uses the new address, and the addresses of the last 24 hours show only the new one, remove the old name from the reverse proxy. Leave the address set; it costs nothing, and an agent restored from an old backup finds its way over as long as it can still reach the old name.

An agent that reports through a name the card does not expect — an internal address such as http://10.0.0.5:3000 that bypasses the proxy — is listed as well and moves to the public name too, provided it can reach it.

Undoing it​

  • For every agent: on the new server, enter the old server's address under Agents → Server move. The agents move back the same way, with the same check against the old server.
  • For one host: re-run the installer with the agent key against the server it should report to. It writes a fresh config.yaml and deletes server-url. Deleting server-url alone is no longer enough from agent 2.8.0 on, because config.yaml names the new server as well.