Skip to main content

Upgrading

An upgrade is: take a backup, pull the new image, recreate the app container, watch the migrations apply. There is no separate migration step to run — the entrypoint does it — and there is no way to undo one, which is why the backup is first and not optional.

Before you start​

Take a dump. Backup and Restore has the command and the TimescaleDB caveats. The rollback path for a bad upgrade is that dump; see Rolling back below.

Read CHANGELOG.md for the versions between the one you are running and the one you are moving to. Entries are written for the person running the instance and say what they have to do — the 2.6.0 entry, for example, tells you that what used to be called a region is now a customer, that /api/regions is gone and answers 404, and that you should rename your customers afterwards because they still carry the names they were given as regions.

The sequence​

docker compose pull
docker compose up -d
docker compose logs -f app

docker compose up -d stops the old app container and starts the new one. It does not run them side by side, and you must not make it.

Never run two app containers against the same database during an upgrade. Migrations run in the entrypoint, before the new container listens on port 3000, so the moment the new container starts it has already changed the schema underneath anything else that is connected. An older container that is still running does not degrade — it starts throwing errors on queries that name columns the migration just renamed. The 2.6.0 migration renames regions to customers and systems.region_id to systems.customer_id; a 2.5.x container surviving that rename is broken instantly and loudly.

The compose file protects you from this by pinning container_name: selvara-app, which makes a second replica impossible to start. Do not remove that pin to get a zero-downtime deploy: this design has no zero-downtime deploy. The gap is the length of one migration run, which on a normal upgrade is seconds.

Browsers that hit the gap see the proxy's error page. To show a Selvara page that reloads itself instead, see Installation → A maintenance page during upgrades.

The cron container waits on the app's healthcheck (depends_on: service_healthy), so it will not fire alert evaluations at a half-migrated database. Leave that dependency in place.

What the logs look like when it worked​

Running database migrations...
Skipped: 0001_baseline
Skipped: 0002_metric_series_and_index_cleanup
Applied: 0007_customers_and_server_region
Applied 1 migration(s)
Migrations completed successfully
Configuring TimescaleDB retention...
Retention active: raw 7 days (compressed after 2 days), 5m rollup 30 days, 1h rollup 400 days
Running database seed...
Seeding complete!
Starting application...

Skipped means already applied — the migrator tracks what has run in the database and only executes the rest. No pending migrations is what a restart with nothing new to do prints.

Two lines deserve attention:

  • Migration failed: followed by an error means the script exited 1 and the server was never started. The container will be restarting in a loop. On PostgreSQL the whole migration run happens inside one transaction, so the schema is rolled back to where it was rather than left half-migrated — the old image will still run against it. Stop the container, read the error, and either fix the cause or go back to the previous tag.
  • TimescaleDB setup failed (app will still start): is deliberately non-fatal. The app comes up, but retention and compression may not be configured — which means the metrics table can grow without bound. Do not leave it. Troubleshooting covers the disk case.

Then confirm the instance is actually serving:

curl -s http://127.0.0.1:3000/api/health
docker compose ps

Two version numbers​

They live in package.json at the repository root and they move independently:

FieldWhat it tracksChangelog
versionThe dashboard: UI, API, schema, deployment.CHANGELOG.md, newest first.
agentVersionThe Go agent under agent/. Bumped only when that binary changes.Noted under the dashboard release that ships it, as an ### Agent x.y.z section.

They are not meant to match and never have. The agent ran ahead for a long stretch because it shipped updates while the dashboard's number sat still.

Images are tagged latest and <YYYYMMDD>-<short-sha> by the build pipeline. Pin the dated tag in docker-compose.yml if you want upgrades to be a deliberate edit rather than whatever latest resolved to this morning.

Upgrading the agents​

You do not upgrade agents separately. The dashboard image contains the compiled agent binaries, and /api/agent/version advertises the agentVersion baked into that image. Agents check it and update themselves, restarting whichever systemd unit actually started them. So pulling a new dashboard image is what puts a new agent version in flight.

Whether it goes out to everything at once or one canary per customer first is a setting, not a property of the upgrade. See Agent Rollout if you want to stage it, and hold updates on individual systems before you pull rather than after.

Rolling back​

There is no automatic down-migration. The migrations do define down blocks, but nothing in the deployment ever calls them: scripts/migrate.mjs only calls migrateToLatest(). Starting an older image against a newer database does not undo anything — it just runs the old code against a schema it does not know.

So the rollback is:

  1. docker compose stop app cron
  2. Restore the dump you took before the upgrade, following Backup and Restore.
  3. Pin the previous image tag in docker-compose.yml.
  4. docker compose up -d

Everything written between the dump and the rollback is lost — metrics, alerts, events, and anything an operator changed in the dashboard. That window is exactly as long as you let it be, which is the argument for taking the dump immediately before the pull rather than the night before.

If the upgrade is only cosmetically wrong — a page you dislike, a label in the wrong language — do not roll back. A schema that has moved forward is much easier to live with than a restore.

Moving to another machine​

An upgrade keeps the server where it is. Moving the instance to a new machine is a restore on the new one plus telling the agents where to go, which the old server can do for them once every agent runs the current version. See Moving to a new server.