Skip to main content

Agent Rollout

Agents keep themselves up to date. This page is about controlling that: what an agent asks and when, the two rollout modes, how to hold one back, and what to do when one is stuck.

The target version​

There is one target version per dashboard — the agent version that shipped with the dashboard build you are running. The binaries live inside the dashboard image and are served from /api/agent/download; there is no external download. Upgrading the dashboard is what makes a new agent available, so the rollout you control here is the second half of an upgrade, not a separate one. See Upgrading.

The Agents page states it plainly in the card subtitle: N of M on <version>.

How an agent asks​

With auto-update enabled — the default — the agent calls GET /api/agent/version once at startup and then every check interval (auto_update.check_interval in config.yaml, five minutes when unset). It sends the version it is running this second in X-Agent-Version, and a bearer token signed with the webhook secret.

That token is what makes a per-host answer possible: the dashboard verifies it, recognises the system, and applies the rollout policy to that system alone. An agent too old to sign its check gets the plain comparison instead — target version, and update if it differs.

The answer carries the target version, the download URL, whether this host should update, the rollout decision behind that verdict, and the name of the canary it is waiting for, when it is waiting.

The comparison is equality, not ordering. An agent running a version the dashboard does not know — a test build, or one left over from a dashboard that was rolled back — counts as "not current" and will be told to install the dashboard's version over it.

How an agent updates itself​

  1. Downloads the binary for its own architecture to a temporary file next to the current one, so the move that follows is within one filesystem.
  2. Makes it executable and runs it with -version. A binary that does not run, or that reports another version than the one on offer, is deleted and the update fails without touching the installed one. Installed, such a binary would see the same update again after its restart and install it again and again.
  3. Renames the current binary to selvara-agent.backup, moves the new one into place, then removes the backup. If the move fails, the backup is renamed back.
  4. Gives the outbox up to five seconds to hand over the reports still queued, so an update does not swallow the events of the last cycles.
  5. Restarts its own systemd unit. On Windows this step is not the agent's to take — see How an update completes on Windows below.

Step 5 is worth knowing about. The agent does not restart a hard-coded unit name: it reads its own unit out of /proc/self/cgroup, which is how it survives being installed under a different name. If it cannot work the unit out, it re-executes itself in place rather than restarting a unit that might belong to something else. service_name in config.yaml overrides the detection, and is only for installations the cgroup does not describe. On Synology DSM 6 there is no systemd at all: the agent re-executes itself in place, keeping its process id, so upstart goes on watching it. Should the restart fail however it was attempted, the agent exits with an error and leaves starting the new binary to its service manager: systemd's Restart=always, upstart's respawn.

Everything that follows is logged with a [Updater] prefix:

journalctl -u selvara-agent | grep Updater

To take one host out of the mechanism entirely, set auto_update.enabled: false in its config.yaml and restart the agent (systemctl restart selvara-agent, or Restart-Service SelvaraAgent). That agent then stops asking, and the dashboard's rollout state for it becomes meaningless — it will sit on its version until you run the installer on it by hand. Prefer a hold (below) unless you specifically want the agent to stop talking to the version endpoint.

How an update completes on Windows​

Steps 1 to 4 are the same. Step 5 cannot be: the agent cannot restart itself there. To the service control manager a process that exits without having been told to stop looks like a crash, and a process the service spawns is not a service — it runs outside the manager's view and dies with the service that started it.

So once the new binary is in place the agent exits with a failure code, and the service's recovery action starts it again. Two settings make that work, and the installer writes both:

sc.exe failure SelvaraAgent reset= 86400 actions= restart/5000
sc.exe failureflag SelvaraAgent 1

The flag is not optional. Without it Windows runs a recovery action only for a process that dies outright, and never for a service that stops and reports a failure code — which is exactly what the agent does. A service registered without the flag updates itself once and then stays stopped. If you register SelvaraAgent by hand rather than with the installer, set both.

There is no journal, so the [Updater] lines are in the log file, and the moment the update completed is in the Windows Application event log as well:

Select-String Updater C:\ProgramData\Selvara\logs\agent.log
Get-EventLog -LogName Application -Source SelvaraAgent -Newest 20

A selvara-agent.exe.backup left beside the binary is normal. Windows allows a running image to be renamed but not deleted, so the file the agent is executing from cannot go until that process does; the next update removes it, or the reboot it was queued for does.

An agent run with -foreground does not come back​

-foreground runs the agent in the console instead of under the service control manager, and nothing there restarts it. It downloads the update, installs it, exits — and stays exited. That is not a fault to debug: start it again, and it is on the new version. Under the service the same exit is the whole design.

If you are testing a version by hand and want the host to keep itself current afterwards, stop the console agent and start the service.

The two modes​

The mode is instance-wide and set by an administrator in the Agents page header. It takes effect on the next check each agent makes.

All at once​

Every agent that is not on the target version and not held updates the next time it asks. On a five-minute check interval, a fleet is converted within five minutes. Use it for a fleet small enough to fix by hand, or for an update you have already proven elsewhere.

Staged​

One canary per customer goes first, and everything else waits until the canaries have proven the version.

The canary of a customer is its non-held system with the smallest id. It is not chosen by you and it does not rotate; it is stable as long as the membership and the holds do not change. The Agents page marks it with a Canary badge. Holding a system excludes it from being the canary, which is the one lever you have over the choice: hold the system you do not want going first, and the next-smallest id takes over.

A canary counts as proven when all three are true:

  • it reports the target version,
  • its status is online,
  • and it has been on that version for at least 15 minutes.

A canary whose version-change time the dashboard never recorded counts as proven — it was already on that version before tracking began, which is longer than any soak.

Every other system waits for every canary, not only its own. If one customer's canary is offline, or crashed back to its old version, or has been on the new one for only ten minutes, the whole fleet waits. That is deliberate: a bad agent version is a fleet-wide problem, and a canary that went offline on the new version is exactly the signal the soak exists to catch. The waiting system's row names the canary it is waiting for.

A customer whose systems are all held has no canary at all, and holds nobody else up.

So the timeline of a staged rollout of a fresh version is: within one check interval every canary updates; 15 minutes later, provided every canary is online and still on the new version, the rest update at their next check.

Holding an agent back​

The pause button on an agent's row stops that system updating, in either mode. Hold all and Release all in the page header do it for the whole fleet — useful just before a maintenance window, or the moment a new version misbehaves.

A held system:

  • never updates, whatever the mode, and shows the rollout state Held;
  • is never a canary, and is skipped when the canary of its customer is chosen;
  • is not waited for by anyone else.

Holding does not stop the agent reporting, collecting or applying settings. It only stops the version swap. Releasing puts the system back under the selected mode; it updates at its next check, subject to the staged rules.

Holding is the right tool for "not now". Disabling auto-update in config.yaml is the tool for "not ever, this host is managed elsewhere" — it is invisible in the dashboard, so it will look like a stuck agent to whoever inherits the instance.

Reading the Agents page​

One row per system, with a filter for Outdated, Held and With collector problems, and a search by system name.

ColumnWhat it tells you
Systemname, with a status dot, linking to the system
Customerthe customer code
Agent versiongreen on the target version, amber when behind, Not reported when the system has never reported one
RolloutCurrent, Update pending, Waiting for canary (with the canary's name underneath), Held, or Unknown for a system that has never reported a version. A Canary badge marks the canary of each customer
ConfigurationCurrent when the agent has applied the settings version the dashboard holds, amber Pending while it has not. Hover for the applied and desired numbers. See Agent Settings
CollectorsOK, or a red badge counting collectors in error and an amber badge counting services waiting for configuration. Hover either for the service names
Last seenwhen the last report arrived
Updatesthe hold/release toggle, for administrators and managers (their granted customers only)

Update pending means the dashboard would answer "update" — the agent has simply not asked yet. It clears within one check interval. A row that stays on Update pending for much longer than that is a stuck agent.

When an agent is stuck on an old version​

Work down this list; each step is cheap.

  1. Is it held? The Rollout column says Held. Release it.

  2. Is it waiting for a canary? The row names the canary. Look at that system: if it is offline, or fell back to the old version, the whole fleet is waiting for it by design. Fix it, hold it (which removes it as canary and passes the role on), or switch to All at once if you accept the risk.

  3. Has it reported at all recently? A stale Last seen is not an update problem — the agent is not talking to the dashboard. Go to Agent Installation and Troubleshooting.

  4. Is auto-update switched off on the host? grep -A2 auto_update /opt/selvara-agent/config.yaml. An enabled: false there explains everything and is invisible from the dashboard.

  5. Read the updater's log. journalctl -u selvara-agent | grep Updater shows what it tried; on Synology DSM 6, grep -e Updater -e restart /volume1/@selvara-agent/agent.log.

    • Update failed: download failed with status 404 — the dashboard image has no binary for that host's architecture. Check whether the host is arm64 and the image only carries amd64.
    • new binary verification failed — the downloaded binary does not run on that host. It was never installed; the old one is still in place.
    • the update reports version "…", not the "…" on offer — the dashboard offers one version and ships a binary of another. The agent keeps the one it runs and tries again at the next check; the dashboard image is the thing to fix.
    • Could not determine own systemd unit — the agent re-executed itself instead of restarting the unit. It should still be on the new version; if it is not, set service_name in config.yaml.
    • Failed to restart: failed to resolve executable path: …selvara-agent.backup — an agent before 2.11.2 installed the update but could not start it, which is what happened on Synology DSM 6. The old process keeps running without sending anything, so upstart sees no reason to step in. Run the installer as in step 7: it puts the current version in place, which restarts correctly from then on. sudo initctl restart selvara-agent brings the host back too, but on the version it had downloaded, which stalls the same way on its next update.
    • Nothing at all from [Updater] — the agent is not checking. Either auto-update is off, or it cannot reach the dashboard.
  6. Check the download URL the dashboard hands out. It is built from the dashboard's configured public URL. If that is unset or wrong, agents get a URL they cannot fetch while everything else keeps working, because reports go to the URL in their own config.yaml. See Configuration.

  7. Force it. Running the installer again replaces the binary and keeps the host's configuration:

    curl -sSL https://<your-dashboard>/api/agent/install.sh | sudo bash -s -- \
    --endpoint https://<your-dashboard>

    This bypasses the rollout policy entirely — it does not ask the version endpoint, it installs whatever the dashboard currently serves.

The same list on a Windows host​

Steps 1, 2, 3 and 6 are read in the dashboard and do not change. The three that touch the host do:

# is the service running, and what does its log say?
sc query SelvaraAgent
Select-String Updater C:\ProgramData\Selvara\logs\agent.log

# is auto-update switched off on this host?
Select-String -Context 0,2 auto_update C:\ProgramData\Selvara\config.yaml

# force the installed version, keeping the configuration
& ([scriptblock]::Create((irm https://<your-dashboard>/api/agent/install.ps1))) -Endpoint https://<your-dashboard>

Two failures are Windows' own. Update failed: download failed with status 404 there means the dashboard has no Windows binary for that host — it serves 64-bit x86 only. And a service that is stopped shortly after an update, with Update installed, exiting so the service is started again as the last event log entry, is a service whose recovery action or failure flag is missing: see How an update completes on Windows above, and re-run the installer, which sets both.

After an agent updates​

The dashboard notices the version in the next report and records an agent.updated event on that system, and the agent itself emits agent.started when it comes back up. Both can notify — see Event Rules. A rollout you want to watch is therefore readable as a timeline, not only as a table that slowly turns green.