Skip to main content

Agent Installation

The agent is a single static Go binary. It runs as a service on every host you want to monitor - a systemd unit on Linux, an upstart job on Synology DSM 6, a Windows service on Windows - looks at the host once every 30 seconds and posts what it found to the dashboard's webhook. It opens no ports and needs no inbound access.

This page covers getting one onto a host and confirming it works. What it then collects is on Agent Collectors; the tokens and logins some collectors need are on Agent Settings.

Two things that surprise people​

The system is created in the dashboard first. The agent does not self-register. It is installed with the identity of a system that already exists, and a host whose system was never created has nowhere to report to. Create it under Systems → Add system, or create a batch of them on the agent page's Install tab. See Customers and Systems.

The API key is shown once. It is created with the system, shown in the dialog that follows, and stored only as a hash — the dashboard cannot show it again. If you lose it before the install, regenerate it: Regenerate key on the system's row, or Generate API key on the agent page's Install tab. Either replaces the system's key and shows the new one.

Regenerating does not disturb an agent that is already running, despite the warning the dialog shows. The key is only used for the one-time configuration fetch during installation; everything afterwards is authenticated with the webhook secret. You need the new key only if you install, or re-fetch the configuration, on a host.

Requirements​

Everything up to Windows hosts at the end of this page describes a Linux host: the paths, the unit and the install.sh command are that installer's.

On the hostWhy
Linux on amd64 or arm64the architectures the dashboard serves a Linux binary for
systemdthe agent is installed as a systemd unit and refuses to install without one; Synology DSM 6, which has upstart instead, is the one exception, see Synology DSM 6
curlused to download the binary and the configuration
root (sudo)the installer writes to /opt and /etc/systemd/system; the service itself also runs as root
outbound HTTPS to the dashboardreports, settings and update checks all go outward from the host

The installer installs no packages of its own.

Where the install command comes from​

Creating a system hands it to you straight away: the System created – install the agent dialog shows the key and the finished command.

You can also get it later. Open Agents → Install, pick the target system, and press Generate API key. Either way the command comes out with that system's key and this dashboard's address filled in:

curl -sSL https://<your-dashboard>/api/agent/install.sh | sudo bash -s -- \
--api-key <key> --endpoint https://<your-dashboard>

Run it on the host as root. The --endpoint is always required, including on updates; the script does not remember it.

For several hosts at once, the Add multiple systems card creates a batch of systems and offers the result as CSV — one row per system with its name, key and ready-made install command. That CSV is the only place those keys are ever shown.

What the installer does​

  1. Checks the ground. Refuses without root, without systemd (/run/systemd/system) or DSM 6's upstart, or without curl. Maps uname -m to amd64 or arm64 and stops on anything else.
  2. Downloads the binary from /api/agent/download?arch=<amd64|arm64> to /opt/selvara-agent/selvara-agent.new and runs it with -version. A binary that does not run is deleted and the install aborts, so a bad download cannot leave the host without a working agent.
  3. Fetches the configuration from /api/agent/config?key=<key> and writes it to /opt/selvara-agent/config.yaml with mode 600. The dashboard looks the key up, and returns that system's id, the webhook URL and the webhook secret. A fresh configuration also deletes /opt/selvara-agent/server-url, so the host reports to the server the new file names. Without --api-key this step is skipped and the existing files are kept.
  4. Swaps the binary in. Stops a running agent, moves .new over /opt/selvara-agent/selvara-agent.
  5. Writes and starts the unit. /etc/systemd/system/selvara-agent.service, then daemon-reload, enable, restart. The unit joins the group of every PHP-FPM pool socket it finds at that moment — see What the agent needs to see. After two seconds it checks the unit is active and prints the version, or exits non-zero and points you at the journal.

Where everything lives​

PathContents
/opt/selvara-agent/selvara-agentthe binary
/opt/selvara-agent/config.yamlthe configuration, mode 600
/opt/selvara-agent/server-urlonly on a host that was moved to another server: the server it reports to now, which outranks webhook_url — see Backup and Restore
/etc/systemd/system/selvara-agent.servicethe unit
journal, identifier selvara-agentall output; the agent writes no log files

The configuration file​

The file holds only how to reach the dashboard. Which services run on the host is detected, and the tokens and logins they need come from the dashboard.

system_id: "<the system's id in the dashboard>"
webhook_url: "https://<your-dashboard>/api/webhook/metrics"
webhook_secret: "<the dashboard's WEBHOOK_SECRET>"
interval: 30
KeyRequiredDefaultMeaning
system_idyes–which system this host reports as
webhook_urlyes–where reports go. The settings endpoint is derived from it: same host, /api/agent/settings
webhook_secretyes–signs every request. Must equal the dashboard's WEBHOOK_SECRET. The WEBHOOK_SECRET environment variable overrides the file
intervalno30seconds between collection cycles
auto_update.enablednotrueset to false to stop this host updating itself
auto_update.check_intervalno5minutes between update checks
service_namenodetectedthe unit the updater restarts. Leave it out: the agent reads its own unit from /proc/self/cgroup
collectors––ignored. Dashboards before agent 2.0 wrote a full collector list into every host's file. Detection decides now; the agent logs once at start that it is ignoring the list
docker.socketno/var/run/docker.sockwhere the Docker daemon listens
patroni.url, garage.admin_url, garage.admin_token, nginx.status_url, caddy.admin_url, zfs.poolsnodetectedpin an address detection got wrong. GARAGE_ADMIN_TOKEN in the environment overrides the token

The pins in the last row exist for hosts where detection cannot see the service. Prefer the dashboard for these: a value set there wins over the same value in the file, applies without touching the host, and can be set once for a whole cluster. See Agent Settings.

The agent refuses to start without system_id, webhook_url and webhook_secret, and says which one is missing.

A host that was moved to another server, or to a new name of the same one, records the new address in a file named server-url next to the binary. On every start that file replaces the scheme and host of webhook_url, and the log says so (Server URL overridden by …). From agent 2.8.0 the start also writes the new address into config.yaml (Wrote … into …), so the old one is gone from the host even if server-url is lost; where config.yaml cannot be written, server-url alone carries the move. To send one host elsewhere, re-run the installer against that server. See Moving to a new server.

Check that it works​

On the host:

systemctl status selvara-agent
journalctl -u selvara-agent -f
/opt/selvara-agent/selvara-agent -version

A healthy start logs the version, the system id, the webhook URL and the interval, then Agent started successfully. After that the log is quiet by design: it carries changes only — a new detection result, an applied settings revision, a collector that started or stopped failing, an unreachable webhook.

In the dashboard, the system turns online within a cycle or two and its Info tab lists the detected services and the state of every collector. The Agents page shows the same hosts as one list, with the version each is running.

If the unit is running but the system stays offline, the problem is between the agent and the webhook: the URL in config.yaml, the secret (it must be the dashboard's WEBHOOK_SECRET exactly), or a proxy in front of the dashboard answering with something other than 2xx. The log line for a rejected report carries the status and body the dashboard answered with. More in Troubleshooting.

What the agent needs to see​

The unit runs the agent as root, and that is not incidental: fail2ban's socket, WireGuard, HAProxy's stats socket, the certificate files under /etc/letsencrypt and the journal are not readable otherwise. The unit limits what that root can do:

  • ProtectSystem=strict with ReadWritePaths=/opt/selvara-agent — the agent can write nowhere but its own directory.
  • ProtectHome=true, PrivateTmp=true, NoNewPrivileges=true, ProtectKernelTunables=true, ProtectControlGroups=true, RestrictSUIDSGID=true.
  • Capabilities bounded to CAP_DAC_READ_SEARCH (reading files it does not own), CAP_NET_ADMIN (wg show) and CAP_SYS_PTRACE.
  • SupplementaryGroups= with the group of every socket found in /run/php/*.sock, /run/php-fpm/*.sock and /var/run/php-fpm/*.sock (root excluded), written only when there is one.

On a Proxmox VE node the agent also starts one command outside that sandbox: lvs, every five minutes, to read how full the LVM thin pools and their metadata are. The kernel reports that only to CAP_SYS_ADMIN, which the unit leaves out on purpose, so the agent asks systemd, through systemd-run, for a transient unit that has it and runs nothing but that lvs report — itself sandboxed with a read-only system, no network and no new privileges. A node without an LVM-thin storage never runs it.

The last line exists because of what the bounding set leaves out. CAP_DAC_READ_SEARCH lets the agent read files it does not own, but connecting to a Unix socket needs write access, and without CAP_DAC_OVERRIDE root is refused by a www-data:www-data 0660 pool socket like any other account outside the group.

A pool added after the install, or a host installed before the installer collected the groups, is handled by the agent. With every detection — every five minutes — it compares the groups of the pool sockets with the groups it runs with. When one is missing it asks systemd, through systemd-run, to write /etc/systemd/system/selvara-agent.service.d/socket-groups.conf with every socket group, reloads systemd and restarts. It cannot write that file itself — the system is read-only inside its sandbox — so the drop-in is written by a short transient unit that systemd starts for it, the same service manager it asks to restart it after an update. The groups it asked for are recorded in socket-groups next to the binary, so a drop-in that does not take effect costs one restart, not one every five minutes. Remove that file to make it try again. Access can also be granted in the pool itself — see Agent Collectors.

Beyond the host's own files it talks to local service endpoints — the Docker socket, Patroni's REST API, etcd, PgBouncer, PostgreSQL, Garage's admin API, Traefik's metrics page, Caddy's admin API, CrowdSec's Local API — on loopback or on the host's own addresses. Whatever you firewall between the agent and a service on the same host, the collector for that service will not work.

If you run the agent under something other than this unit, give the process group membership for every socket it must read (docker, fail2ban's socket group, and so on). Collectors whose input it cannot read will report an error rather than fail the whole agent.

Updating an existing installation​

Run the same command without --api-key:

curl -sSL https://<your-dashboard>/api/agent/install.sh | sudo bash -s -- \
--endpoint https://<your-dashboard>

The binary is replaced, the configuration is kept. In normal operation you do not need this: agents update themselves. See Agent Rollout.

Uninstalling​

curl -sSL https://<your-dashboard>/api/agent/install.sh | sudo bash -s -- --uninstall

This disables and stops the unit, removes /etc/systemd/system/selvara-agent.service, reloads systemd and deletes /opt/selvara-agent including the configuration. By hand it is the same four steps:

systemctl disable --now selvara-agent
rm /etc/systemd/system/selvara-agent.service
systemctl daemon-reload
rm -rf /opt/selvara-agent

The system stays registered in the dashboard with all its history until you delete it there. It will go offline, and unless you silence or delete it, that is an alert. A host that is going away for good should be deleted in the dashboard as well; a host going down for maintenance belongs in a Maintenance Window.

Synology DSM 6​

DSM 7 runs systemd, so the agent installs there exactly as described above. DSM 6 predates systemd on Synology and uses upstart. The same install.sh command, with the same key, recognises it by /etc.defaults/VERSION and initctl. On the dashboard's agent page, choose Synology DSM 6 as the platform to see the matching service commands.

What is different on DSM 6:

DSM 6
Install path/volume1/@selvara-agent (the first data volume). DSM rewrites its small system partition on updates; the volume keeps the agent, and the @ keeps the directory out of the shares.
Servicethe upstart job /etc/init/selvara-agent.conf, with respawn
Start at boot/usr/local/etc/rc.d/selvara-agent.sh, which DSM runs once the volumes are mounted. If a DSM update removed the upstart job, it puts it back from upstart.conf next to the agent.
Log/volume1/@selvara-agent/agent.log, rotated at 8 MB with three old files kept: there is no journal
Hardeningnone: upstart has nothing like the unit's sandbox, so the agent runs as plain root

Checking on it:

initctl status selvara-agent
tail -f /volume1/@selvara-agent/agent.log
initctl restart selvara-agent

The agent reports its platform as Synology DSM with the version and update, for example 6.2.4-25556 Update 8. It updates itself like on any other host, restarting in place; if that fails, it exits and the job's respawn starts the new binary. Agents before 2.11.2 could not restart after an update on DSM 6 and have to be brought to the current version once with the install command, see Agent Rollout. Uninstalling is the same install.sh --uninstall; it removes the job, the boot script and the directory on the volume.

What DSM 6 does not have, the agent does without: the journal (so no SSH or sudo login events), timedatectl and apt. Software RAID health is read from /proc/mdstat as on any Linux host, see Agent Collectors.

After a major DSM update, check that the agent came back. If the update removed the boot script as well, run the install command again with --endpoint only; the configuration on the volume is kept.

Windows hosts​

On Windows the agent is a service named SelvaraAgent, installed by a PowerShell script that does what install.sh does on Linux. The dashboard serves one Windows binary, 64-bit x86; there is no 32-bit and no ARM build.

On the hostWhy
Windows on 64-bit x86the only Windows build the dashboard serves
PowerShell 5.1, which ships with Windows, or newerthe installer is a PowerShell script
an elevated shellit writes to Program Files and ProgramData and registers a service
outbound HTTPS to the dashboardreports, settings and update checks all go outward from the host

Installing​

The install command needs the key and the dashboard's address, and iex cannot pass arguments to what it runs, so the script is fetched and called as a script block. In an elevated PowerShell:

& ([scriptblock]::Create((irm https://<your-dashboard>/api/agent/install.ps1))) -ApiKey <key> -Endpoint https://<your-dashboard>

The key is the same key a Linux host uses and comes from the same two places: the dialog that follows a new system, or Agents -> Install.

exit inside a script block ends the shell that runs it, so a refused install closes the window and takes its own error message with it. Where that matters - an unattended install, or a first attempt with an argument likely to be wrong - download the script and run the file instead, where exit ends only the script:

irm https://<your-dashboard>/api/agent/install.ps1 -OutFile $env:TEMP\selvara-install.ps1
powershell -ExecutionPolicy Bypass -File $env:TEMP\selvara-install.ps1 -ApiKey <key> -Endpoint https://<your-dashboard>

The execution policy is what the second line works around: a script file is subject to it and is refused outright on a default Windows Server, while a script block built from a string never was.

What the installer does​

  1. Checks the ground. Refuses outside an elevated shell, without -Endpoint, on anything but 64-bit x86, and without -ApiKey unless a configuration file is already there.
  2. Downloads the binary from /api/agent/download?arch=amd64&os=windows to selvara-agent.exe.new and runs it with -version. A binary that does not run is deleted and the install aborts, so a bad download cannot leave the host without a working agent. An os the dashboard has no binary for is answered with 404 rather than the nearest match.
  3. Fetches the configuration from /api/agent/config?key=<key> into C:\ProgramData\Selvara\config.yaml and takes every account but SYSTEM and the local administrators off the file with icacls. That file holds the webhook secret; the restriction is the counterpart of mode 600 on Linux. A fresh configuration also deletes C:\Program Files\Selvara\server-url, as on Linux. Without -ApiKey the step is skipped and the existing files kept.
  4. Swaps the binary in. Stops a running service first - Windows locks the image file of a running process - and moves .new over C:\Program Files\Selvara\selvara-agent.exe.
  5. Registers and starts the service. sc.exe create SelvaraAgent binPath= "<exe> -config <config>" start= auto obj= LocalSystem, then sc.exe failure SelvaraAgent reset= 86400 actions= restart/5000 and sc.exe failureflag SelvaraAgent 1, then starts it and waits for it to report Running. It also registers SelvaraAgent as a source in the Application event log, so the agent's service events arrive there with their text rather than as bare numbers.

The recovery action is not only there for crashes. The agent cannot restart itself: to the service control manager a process that exits on its own looks like a failure, and a process it spawns is not a service. After it has installed an update of itself it exits with a failure code, and this action is what starts the new binary. Without it a host would update itself once and stay down. The failure flag belongs to it: without it Windows runs recovery actions only for a process that dies outright, never for a service that stops and reports a failure code, which is exactly what the agent does after an update.

LocalSystem is needed because the agent replaces its own binary under Program Files when it updates, which nothing below an administrator may write to.

Where everything lives​

PathContents
C:\Program Files\Selvara\selvara-agent.exethe binary
C:\ProgramData\Selvara\config.yamlthe configuration, readable only by SYSTEM and the administrators
C:\Program Files\Selvara\server-urlonly on a host that was moved to another server: the server it reports to now, which outranks webhook_url
C:\ProgramData\Selvara\logs\agent.logthe log, rotated at 8 MiB with three generations kept - this is where the output goes that a Linux host writes to the journal
service SelvaraAgent, display name Selvara Agentthe service itself, started automatically at boot

The contents of the configuration file are the same as on Linux; the table under The configuration file above applies unchanged.

Check that it works​

sc query SelvaraAgent
Get-Content C:\ProgramData\Selvara\logs\agent.log -Tail 50 -Wait
Get-EventLog -LogName Application -Source SelvaraAgent -Newest 20
& "C:\Program Files\Selvara\selvara-agent.exe" -version

In the dashboard the system turns online within a cycle or two, exactly as a Linux host does.

To watch a collection cycle directly, run the binary with -foreground: it then runs in the console instead of under the service control manager, which is the quickest way to see what it reports. Stop the service first, or the two report as the same system.

Updating and uninstalling​

Run the same command without -ApiKey to replace the binary and keep the configuration:

& ([scriptblock]::Create((irm https://<your-dashboard>/api/agent/install.ps1))) -Endpoint https://<your-dashboard>
& ([scriptblock]::Create((irm https://<your-dashboard>/api/agent/install.ps1))) -Uninstall

Uninstalling stops and deletes the service and removes both C:\Program Files\Selvara and C:\ProgramData\Selvara, the configuration and its webhook secret with them. By hand it is the same three steps:

Stop-Service SelvaraAgent
sc.exe delete SelvaraAgent
Remove-Item -Recurse -Force "C:\Program Files\Selvara", "C:\ProgramData\Selvara"

The system stays registered in the dashboard with all its history until you delete it there, and it will go offline - which, unless you silence or delete it, is an alert. The same applies as on Linux: see Uninstalling above.

The Windows binary is not signed​

Nothing in the build signs it, so Windows treats it as software from an unknown publisher. SmartScreen warns when someone downloads the .exe from the dashboard and starts it by hand, and Defender or another endpoint product may quarantine it - in which case the service will not start and the log stays empty. Allowing it is a decision for whoever runs the host. Signing would need a code-signing certificate, which this project does not have.