Agent Installation
The agent is a single static Go binary. It runs as a service on every host you want to monitor - a systemd unit on Linux, an upstart job on Synology DSM 6, a Windows service on Windows - looks at the host once every 30 seconds and posts what it found to the dashboard's webhook. It opens no ports and needs no inbound access.
This page covers getting one onto a host and confirming it works. What it then collects is on Agent Collectors; the tokens and logins some collectors need are on Agent Settings.
Two things that surprise people
The system is created in the dashboard first. The agent does not self-register. It is installed with the identity of a system that already exists, and a host whose system was never created has nowhere to report to. Create it under Systems → Add system, or create a batch of them on the agent page's Install tab. See Customers and Systems.
The API key is shown once. It is created with the system, shown in the dialog that follows, and stored only as a hash — the dashboard cannot show it again. If you lose it before the install, regenerate it: Regenerate key on the system's row, or Generate API key on the agent page's Install tab. Either replaces the system's key and shows the new one.
Regenerating does not disturb an agent that is already running, despite the warning the dialog shows. The key is only used for the one-time configuration fetch during installation; everything afterwards is authenticated with the webhook secret. You need the new key only if you install, or re-fetch the configuration, on a host.
Requirements
Everything up to Windows hosts at the end of this page describes a Linux
host: the paths, the unit and the install.sh command are that installer's.
| On the host | Why |
|---|---|
| Linux on amd64 or arm64 | the architectures the dashboard serves a Linux binary for |
| systemd | the agent is installed as a systemd unit and refuses to install without one; Synology DSM 6, which has upstart instead, is the one exception, see Synology DSM 6 |
curl | used to download the binary and the configuration |
root (sudo) | the installer writes to /opt and /etc/systemd/system; the service itself also runs as root |
| outbound HTTPS to the dashboard | reports, settings and update checks all go outward from the host |
The installer installs no packages of its own.
Where the install command comes from
Creating a system hands it to you straight away: the System created – install the agent dialog shows the key and the finished command.
You can also get it later. Open Agents → Install, pick the target system, and press Generate API key. Either way the command comes out with that system's key and this dashboard's address filled in:
curl -sSL https://<your-dashboard>/api/agent/install.sh | sudo bash -s -- \
--api-key <key> --endpoint https://<your-dashboard>
Run it on the host as root. The --endpoint is always required, including
on updates; the script does not remember it.
For several hosts at once, the Add multiple systems card creates a batch of systems and offers the result as CSV — one row per system with its name, key and ready-made install command. That CSV is the only place those keys are ever shown.
What the installer does
- Checks the ground. Refuses without root, without systemd
(
/run/systemd/system) or DSM 6's upstart, or withoutcurl. Mapsuname -mtoamd64orarm64and stops on anything else. - Downloads the binary from
/api/agent/download?arch=<amd64|arm64>to/opt/selvara-agent/selvara-agent.newand runs it with-version. A binary that does not run is deleted and the install aborts, so a bad download cannot leave the host without a working agent. - Fetches the configuration from
/api/agent/config?key=<key>and writes it to/opt/selvara-agent/config.yamlwith mode 600. The dashboard looks the key up, and returns that system's id, the webhook URL and the webhook secret. A fresh configuration also deletes/opt/selvara-agent/server-url, so the host reports to the server the new file names. Without--api-keythis step is skipped and the existing files are kept. - Swaps the binary in. Stops a running agent, moves
.newover/opt/selvara-agent/selvara-agent. - Writes and starts the unit.
/etc/systemd/system/selvara-agent.service, thendaemon-reload,enable,restart. The unit joins the group of every PHP-FPM pool socket it finds at that moment — see What the agent needs to see. After two seconds it checks the unit is active and prints the version, or exits non-zero and points you at the journal.
Where everything lives
| Path | Contents |
|---|---|
/opt/selvara-agent/selvara-agent | the binary |
/opt/selvara-agent/config.yaml | the configuration, mode 600 |
/opt/selvara-agent/server-url | only on a host that was moved to another server: the server it reports to now, which outranks webhook_url — see Backup and Restore |
/etc/systemd/system/selvara-agent.service | the unit |
journal, identifier selvara-agent | all output; the agent writes no log files |
The configuration file
The file holds only how to reach the dashboard. Which services run on the host is detected, and the tokens and logins they need come from the dashboard.
system_id: "<the system's id in the dashboard>"
webhook_url: "https://<your-dashboard>/api/webhook/metrics"
webhook_secret: "<the dashboard's WEBHOOK_SECRET>"
interval: 30
| Key | Required | Default | Meaning |
|---|---|---|---|
system_id | yes | – | which system this host reports as |
webhook_url | yes | – | where reports go. The settings endpoint is derived from it: same host, /api/agent/settings |
webhook_secret | yes | – | signs every request. Must equal the dashboard's WEBHOOK_SECRET. The WEBHOOK_SECRET environment variable overrides the file |
interval | no | 30 | seconds between collection cycles |
auto_update.enabled | no | true | set to false to stop this host updating itself |
auto_update.check_interval | no | 5 | minutes between update checks |
service_name | no | detected | the unit the updater restarts. Leave it out: the agent reads its own unit from /proc/self/cgroup |
collectors | – | – | ignored. Dashboards before agent 2.0 wrote a full collector list into every host's file. Detection decides now; the agent logs once at start that it is ignoring the list |
docker.socket | no | /var/run/docker.sock | where the Docker daemon listens |
patroni.url, garage.admin_url, garage.admin_token, nginx.status_url, caddy.admin_url, zfs.pools | no | detected | pin an address detection got wrong. GARAGE_ADMIN_TOKEN in the environment overrides the token |
The pins in the last row exist for hosts where detection cannot see the service. Prefer the dashboard for these: a value set there wins over the same value in the file, applies without touching the host, and can be set once for a whole cluster. See Agent Settings.
The agent refuses to start without system_id, webhook_url and
webhook_secret, and says which one is missing.
A host that was moved to another server, or to a new name of the same one,
records the new address in a file named server-url next to the binary. On
every start that file replaces the scheme and host of webhook_url, and the
log says so (Server URL overridden by …). From agent 2.8.0 the start also
writes the new address into config.yaml (Wrote … into …), so the old one
is gone from the host even if server-url is lost; where config.yaml
cannot be written, server-url alone carries the move. To send one host
elsewhere, re-run the installer against that server. See
Moving to a new server.
Check that it works
On the host:
systemctl status selvara-agent
journalctl -u selvara-agent -f
/opt/selvara-agent/selvara-agent -version
A healthy start logs the version, the system id, the webhook URL and the
interval, then Agent started successfully. After that the log is quiet by
design: it carries changes only — a new detection result, an applied
settings revision, a collector that started or stopped failing, an
unreachable webhook.
In the dashboard, the system turns online within a cycle or two and its Info tab lists the detected services and the state of every collector. The Agents page shows the same hosts as one list, with the version each is running.
If the unit is running but the system stays offline, the problem is between
the agent and the webhook: the URL in config.yaml, the secret (it must be
the dashboard's WEBHOOK_SECRET exactly), or a proxy in front of the
dashboard answering with something other than 2xx. The log line for a
rejected report carries the status and body the dashboard answered with.
More in Troubleshooting.
What the agent needs to see
The unit runs the agent as root, and that is not incidental: fail2ban's
socket, WireGuard, HAProxy's stats socket, the certificate files under
/etc/letsencrypt and the journal are not readable otherwise. The unit
limits what that root can do:
ProtectSystem=strictwithReadWritePaths=/opt/selvara-agent— the agent can write nowhere but its own directory.ProtectHome=true,PrivateTmp=true,NoNewPrivileges=true,ProtectKernelTunables=true,ProtectControlGroups=true,RestrictSUIDSGID=true.- Capabilities bounded to
CAP_DAC_READ_SEARCH(reading files it does not own),CAP_NET_ADMIN(wg show) andCAP_SYS_PTRACE. SupplementaryGroups=with the group of every socket found in/run/php/*.sock,/run/php-fpm/*.sockand/var/run/php-fpm/*.sock(root excluded), written only when there is one.
On a Proxmox VE node the agent also starts one command outside that
sandbox: lvs, every five minutes, to read how full the LVM thin pools
and their metadata are. The kernel reports that only to CAP_SYS_ADMIN,
which the unit leaves out on purpose, so the agent asks systemd, through
systemd-run, for a transient unit that has it and runs nothing but that
lvs report — itself sandboxed with a read-only system, no network and
no new privileges. A node without an LVM-thin storage never runs it.
The last line exists because of what the bounding set leaves out.
CAP_DAC_READ_SEARCH lets the agent read files it does not own, but
connecting to a Unix socket needs write access, and without
CAP_DAC_OVERRIDE root is refused by a www-data:www-data 0660 pool socket
like any other account outside the group.
A pool added after the install, or a host installed before the installer
collected the groups, is handled by the agent. With every detection — every
five minutes — it compares the groups of the pool sockets with the groups it
runs with. When one is missing it asks systemd, through systemd-run, to
write /etc/systemd/system/selvara-agent.service.d/socket-groups.conf with
every socket group, reloads systemd and restarts. It cannot write that file
itself — the system is read-only inside its sandbox — so the drop-in is
written by a short transient unit that systemd starts for it, the same
service manager it asks to restart it after an update. The groups it asked
for are recorded in socket-groups next to the binary, so a drop-in that
does not take effect costs one restart, not one every five minutes. Remove
that file to make it try again. Access can also be granted in the pool
itself — see Agent Collectors.
Beyond the host's own files it talks to local service endpoints — the Docker socket, Patroni's REST API, etcd, PgBouncer, PostgreSQL, Garage's admin API, Traefik's metrics page, Caddy's admin API, CrowdSec's Local API — on loopback or on the host's own addresses. Whatever you firewall between the agent and a service on the same host, the collector for that service will not work.
If you run the agent under something other than this unit, give the process
group membership for every socket it must read (docker, fail2ban's socket
group, and so on). Collectors whose input it cannot read will report an
error rather than fail the whole agent.
Updating an existing installation
Run the same command without --api-key:
curl -sSL https://<your-dashboard>/api/agent/install.sh | sudo bash -s -- \
--endpoint https://<your-dashboard>
The binary is replaced, the configuration is kept. In normal operation you do not need this: agents update themselves. See Agent Rollout.
Uninstalling
curl -sSL https://<your-dashboard>/api/agent/install.sh | sudo bash -s -- --uninstall
This disables and stops the unit, removes /etc/systemd/system/selvara-agent.service,
reloads systemd and deletes /opt/selvara-agent including the
configuration. By hand it is the same four steps:
systemctl disable --now selvara-agent
rm /etc/systemd/system/selvara-agent.service
systemctl daemon-reload
rm -rf /opt/selvara-agent
The system stays registered in the dashboard with all its history until you delete it there. It will go offline, and unless you silence or delete it, that is an alert. A host that is going away for good should be deleted in the dashboard as well; a host going down for maintenance belongs in a Maintenance Window.
Synology DSM 6
DSM 7 runs systemd, so the agent installs there exactly as described above.
DSM 6 predates systemd on Synology and uses upstart. The same install.sh
command, with the same key, recognises it by /etc.defaults/VERSION and
initctl. On the dashboard's agent page, choose Synology DSM 6 as the
platform to see the matching service commands.
What is different on DSM 6:
| DSM 6 | |
|---|---|
| Install path | /volume1/@selvara-agent (the first data volume). DSM rewrites its small system partition on updates; the volume keeps the agent, and the @ keeps the directory out of the shares. |
| Service | the upstart job /etc/init/selvara-agent.conf, with respawn |
| Start at boot | /usr/local/etc/rc.d/selvara-agent.sh, which DSM runs once the volumes are mounted. If a DSM update removed the upstart job, it puts it back from upstart.conf next to the agent. |
| Log | /volume1/@selvara-agent/agent.log, rotated at 8 MB with three old files kept: there is no journal |
| Hardening | none: upstart has nothing like the unit's sandbox, so the agent runs as plain root |
Checking on it:
initctl status selvara-agent
tail -f /volume1/@selvara-agent/agent.log
initctl restart selvara-agent
The agent reports its platform as Synology DSM with the version and
update, for example 6.2.4-25556 Update 8. It updates itself like on any
other host, restarting in place; if that fails, it exits and the job's
respawn starts the new binary. Agents before 2.11.2 could not restart
after an update on DSM 6 and have to be brought to the current version once
with the install command, see
Agent Rollout.
Uninstalling is the same
install.sh --uninstall; it removes the job, the boot script and the
directory on the volume.
What DSM 6 does not have, the agent does without: the journal (so no SSH or
sudo login events), timedatectl and apt. Software RAID health is read
from /proc/mdstat as on any Linux host, see
Agent Collectors.
After a major DSM update, check that the agent came back. If the update
removed the boot script as well, run the install command again with
--endpoint only; the configuration on the volume is kept.
Windows hosts
On Windows the agent is a service named SelvaraAgent, installed by a
PowerShell script that does what install.sh does on Linux. The dashboard
serves one Windows binary, 64-bit x86; there is no 32-bit and no ARM build.
| On the host | Why |
|---|---|
| Windows on 64-bit x86 | the only Windows build the dashboard serves |
| PowerShell 5.1, which ships with Windows, or newer | the installer is a PowerShell script |
| an elevated shell | it writes to Program Files and ProgramData and registers a service |
| outbound HTTPS to the dashboard | reports, settings and update checks all go outward from the host |
Installing
The install command needs the key and the dashboard's address, and iex
cannot pass arguments to what it runs, so the script is fetched and called
as a script block. In an elevated PowerShell:
& ([scriptblock]::Create((irm https://<your-dashboard>/api/agent/install.ps1))) -ApiKey <key> -Endpoint https://<your-dashboard>
The key is the same key a Linux host uses and comes from the same two places: the dialog that follows a new system, or Agents -> Install.
exit inside a script block ends the shell that runs it, so a refused
install closes the window and takes its own error message with it. Where
that matters - an unattended install, or a first attempt with an argument
likely to be wrong - download the script and run the file instead, where
exit ends only the script:
irm https://<your-dashboard>/api/agent/install.ps1 -OutFile $env:TEMP\selvara-install.ps1
powershell -ExecutionPolicy Bypass -File $env:TEMP\selvara-install.ps1 -ApiKey <key> -Endpoint https://<your-dashboard>
The execution policy is what the second line works around: a script file is subject to it and is refused outright on a default Windows Server, while a script block built from a string never was.
What the installer does
- Checks the ground. Refuses outside an elevated shell, without
-Endpoint, on anything but 64-bit x86, and without-ApiKeyunless a configuration file is already there. - Downloads the binary from
/api/agent/download?arch=amd64&os=windowstoselvara-agent.exe.newand runs it with-version. A binary that does not run is deleted and the install aborts, so a bad download cannot leave the host without a working agent. Anosthe dashboard has no binary for is answered with 404 rather than the nearest match. - Fetches the configuration from
/api/agent/config?key=<key>intoC:\ProgramData\Selvara\config.yamland takes every account but SYSTEM and the local administrators off the file withicacls. That file holds the webhook secret; the restriction is the counterpart of mode 600 on Linux. A fresh configuration also deletesC:\Program Files\Selvara\server-url, as on Linux. Without-ApiKeythe step is skipped and the existing files kept. - Swaps the binary in. Stops a running service first - Windows locks the
image file of a running process - and moves
.newoverC:\Program Files\Selvara\selvara-agent.exe. - Registers and starts the service.
sc.exe create SelvaraAgent binPath= "<exe> -config <config>" start= auto obj= LocalSystem, thensc.exe failure SelvaraAgent reset= 86400 actions= restart/5000andsc.exe failureflag SelvaraAgent 1, then starts it and waits for it to report Running. It also registers SelvaraAgent as a source in the Application event log, so the agent's service events arrive there with their text rather than as bare numbers.
The recovery action is not only there for crashes. The agent cannot restart itself: to the service control manager a process that exits on its own looks like a failure, and a process it spawns is not a service. After it has installed an update of itself it exits with a failure code, and this action is what starts the new binary. Without it a host would update itself once and stay down. The failure flag belongs to it: without it Windows runs recovery actions only for a process that dies outright, never for a service that stops and reports a failure code, which is exactly what the agent does after an update.
LocalSystem is needed because the agent replaces its own binary under
Program Files when it updates, which nothing below an administrator may
write to.
Where everything lives
| Path | Contents |
|---|---|
C:\Program Files\Selvara\selvara-agent.exe | the binary |
C:\ProgramData\Selvara\config.yaml | the configuration, readable only by SYSTEM and the administrators |
C:\Program Files\Selvara\server-url | only on a host that was moved to another server: the server it reports to now, which outranks webhook_url |
C:\ProgramData\Selvara\logs\agent.log | the log, rotated at 8 MiB with three generations kept - this is where the output goes that a Linux host writes to the journal |
service SelvaraAgent, display name Selvara Agent | the service itself, started automatically at boot |
The contents of the configuration file are the same as on Linux; the table under The configuration file above applies unchanged.
Check that it works
sc query SelvaraAgent
Get-Content C:\ProgramData\Selvara\logs\agent.log -Tail 50 -Wait
Get-EventLog -LogName Application -Source SelvaraAgent -Newest 20
& "C:\Program Files\Selvara\selvara-agent.exe" -version
In the dashboard the system turns online within a cycle or two, exactly as a Linux host does.
To watch a collection cycle directly, run the binary with -foreground: it
then runs in the console instead of under the service control manager, which
is the quickest way to see what it reports. Stop the service first, or the two
report as the same system.
Updating and uninstalling
Run the same command without -ApiKey to replace the binary and keep the
configuration:
& ([scriptblock]::Create((irm https://<your-dashboard>/api/agent/install.ps1))) -Endpoint https://<your-dashboard>
& ([scriptblock]::Create((irm https://<your-dashboard>/api/agent/install.ps1))) -Uninstall
Uninstalling stops and deletes the service and removes both
C:\Program Files\Selvara and C:\ProgramData\Selvara, the configuration
and its webhook secret with them. By hand it is the same three steps:
Stop-Service SelvaraAgent
sc.exe delete SelvaraAgent
Remove-Item -Recurse -Force "C:\Program Files\Selvara", "C:\ProgramData\Selvara"
The system stays registered in the dashboard with all its history until you delete it there, and it will go offline - which, unless you silence or delete it, is an alert. The same applies as on Linux: see Uninstalling above.
The Windows binary is not signed
Nothing in the build signs it, so Windows treats it as software from an
unknown publisher. SmartScreen warns when someone downloads the .exe from
the dashboard and starts it by hand, and Defender or another endpoint product
may quarantine it - in which case the service will not start and the log
stays empty. Allowing it is a decision for whoever runs the host. Signing
would need a code-signing certificate, which this project does not have.