Skip to main content

Agent Settings

Almost everything the agent needs it finds on the host itself — see Agent Collectors. What it cannot find is the short list on this page: admin tokens, database logins, a metrics URL behind a non-default path, the URLs you want checked. Those are entered in the dashboard and fetched by the agent.

Nothing here goes into config.yaml on the host. A value set in the dashboard wins over the same value in the file, needs no shell access, and can be set once for a whole cluster.

Where they are set​

ScopeWhereApplies to
Systemthe system's page → Configuration tabthat one host
Clusterthe cluster's page → Configuration tabevery system in that cluster

Both forms show only the fields that are relevant: the ones belonging to a service the agent actually detected, plus the host-wide ones. A host without Garage never shows a Garage token field.

And only the ones its platform can use. A setting that means nothing on an operating system is not offered there: a Windows system is not asked for units_watch, certificates_paths or zfs_pools, and a Linux system is not asked for services_watch, eventlog_channels, tasks_watch or certstore_stores. The platform is what the host reported, so a system whose agent has never reported still shows everything — the dashboard does not know yet — and so does every cluster form, because a cluster's members may run both. A value already stored for a setting that is no longer shown stays stored; it is only no longer offered.

The system scope wins. The agent receives the cluster's values first and the system's own on top, key by key. A field on a system's form shows where its effective value comes from, so an inherited value is visible as inherited rather than looking unset.

Clearing a field removes the value at that scope. Clearing a system value does not blank the setting — it falls back to the cluster's.

Administrators can save, and managers can for a system at a customer granted to them or a cluster whose every member is. Everybody else sees the fields and the values (secrets masked) but cannot change them. Every save records a settings.changed event on the system or cluster, naming the keys that were written and who wrote them, so a collector that started failing can be put next to the change that caused it. See Event Rules.

How a setting reaches the agent​

Every system carries a configuration version. Saving on a system bumps that system's version; saving on a cluster bumps the version of every system in it.

The dashboard answers each report with the version it currently holds for that host. When that differs from the version the agent has applied, the agent fetches /api/agent/settings on its next cycle — so at the default interval a change lands within about 30 seconds. Independently of that it refreshes every five minutes, which also covers a dashboard that was unreachable when the change was made.

When the values arrive, the agent reconciles its collectors: one whose inputs changed is rebuilt, one that now has what it was missing is started, one that is no longer needed is closed. Nothing restarts and no other collector is disturbed. The one visible effect is that a rebuilt collector's rates are missing from the next report, because a rate needs two samples.

You can watch this from either end. On the host, journalctl -u selvara-agent logs Applied settings revision N. In the dashboard, the Agents page has a Configuration column reading Current or Pending, with the applied and desired version numbers on hover — see Agent Rollout.

Secrets​

Fields marked secret below are encrypted at rest with the instance's encryption key, the same one that protects the SMTP password. They are:

  • Masked in the UI as ••••••••, in both the value and the inherited value. Re-saving a form with the mask still in the field leaves the stored secret untouched, so you can edit the field next to it without wiping a token.
  • Never returned in the clear by the admin API. The only reader that decrypts them is the settings endpoint the agent calls, and it only ever returns the values belonging to the system whose signed token it just verified.
  • Written to the host in the clear, in memory, for as long as the collector needs them. A token that grants more than read access to a service should not be used here; create a dedicated one.

If your instance's encryption key changes, stored secrets can no longer be decrypted. They are skipped rather than sent as garbage, and the affected collectors fall back to needs_config. Re-enter them.

Reference​

Text and secret fields take a single value. List fields take one entry per line (commas work too). The boolean field is on when set to true.

ServiceKeyWhat it is forSecret
Garagegarage_admin_urlAdmin API base URL, when detection found the wrong one or Garage does not listen on port 3903no
Garagegarage_admin_tokenToken for the v2 admin API. Nothing is collected without ityes
Patronipatroni_urlREST API base URL, when it is not on port 8008 or not reachable on the detected addressno
Patronipatroni_usernameOffered for a protected REST API — see the note belowno
Patronipatroni_passwordOffered for a protected REST API — see the note belowyes
etcdetcd_urlClient URL, when it is not the detected oneno
PgBouncerpgbouncer_addresshost:port of the admin interface, overriding the detected oneno
PgBouncerpgbouncer_userA login listed in PgBouncer's stats_users. Nothing is collected without user and passwordno
PgBouncerpgbouncer_passwordPassword for that loginyes
PostgreSQLpostgresql_addresshost:port, overriding the detected oneno
PostgreSQLpostgresql_userA role holding pg_monitor. Nothing is collected without user and passwordno
PostgreSQLpostgresql_passwordPassword for that roleyes
PostgreSQLpostgresql_databaseDatabase to connect to; postgres when emptyno
nginxnginx_status_urlThe stub_status URL. Needed whenever it is not http://127.0.0.1/nginx_statusno
PHP-FPMphpfpm_socketThe FPM socket path. Needed when no /run/php/php*-fpm.sock existsno
Traefiktraefik_metrics_urlThe Prometheus metrics URL. Needed when only Traefik's ping endpoint was found; /metrics is appended if missingno
Caddycaddy_admin_urlAdmin API base URL; http://127.0.0.1:2019 when emptyno
Dockerdocker_container_statsSet to true to add CPU and memory per running container. Costs one stats request per container per cycleno
CrowdSeccrowdsec_lapi_urlLocal API URL, overriding the detected addressno
CrowdSeccrowdsec_api_keyA bouncer key (cscli bouncers add selvara). Nothing is collected without ityes
ZFSzfs_poolsPools to monitor, overriding detection. One per lineno
Proxmox VEproxmox_guests_excludeVMIDs of guests left out of the backup and autostart checks, one per line: test machines, guests meant to be off, guests backed up some other way. They are still listedno
host-widechecks_urlsURLs to check from this host, one per line. The HTTP check collector exists only because of this settingno
host-wide (Linux)units_watchsystemd units to report individually, one per line. A name without a suffix is read as .service. Failed units are counted whether listed here or notno
host-wide (Linux)certificates_pathsGlob patterns for certificate files, one per line. Replaces the defaults (/etc/letsencrypt/live/*/fullchain.pem, /etc/haproxy/certs/*.pem) and runs the collector even where those match nothing. A file under <root>/live/<name>/ is still left out when <root>/renewal/<name>.conf does not exist — see Agent Collectorsno
host-wideservices_disabledDetected services switched off on this host, comma-separated. Set on the system's Services tab, not in the settings form — see belowno
host-wide (Windows)services_watchWindows services to report individually, one per line. The names the service control manager knows, not the display names. Automatic services that are not running are counted whether listed here or notno
host-wide (Windows)eventlog_channelsEvent log channels to read, one per line. Added to the defaults (Security, System, Application), which are always readno
host-wide (Windows)tasks_watchScheduled tasks to report the last run of, one per line, written the way Task Scheduler addresses them (\Folder\Task). A name without a folder is a task in the root folder. Failed tasks are counted whether listed here or notno
host-wide (Windows)certstore_storesLocalMachine certificate stores to read, one per line. Added to the default (My), which is always read; WebHosting and Root are the usual additionsno

The host-wide keys belong to no detected service. checks_urls is offered on every system and cluster; the rest are offered where the host's operating system can use them, as described above. services_disabled is never offered by the form.

Switching a service off​

Detection runs a collector for everything it finds, which is not always wanted: an nginx that only exists to redirect to HTTPS, a PostgreSQL nobody is meant to log into. A system's Services tab lists every service the agent detected with a checkbox; clearing one switches that service off for this host. Administrators can save, and managers for a system at a customer granted to them; everybody else sees the list read-only.

The choice is stored as the setting services_disabled, a comma-separated list of service names, and reaches the agent like every other setting. For a service in that list the agent

  • runs no collector and reports no needs_config for it,
  • reports its collector state as disabled,
  • keeps detecting it, so it stays on the Services tab and can be switched back on.

The dashboard then shows no badge for it, offers none of its fields on the Configuration tab, does not count it towards the Collector error and Needs configuration tiles, and the systems list's service filter no longer matches the host on it.

The setting can be stored for a cluster as well. As with every setting, a value stored on the system replaces the cluster's as a whole — the two lists are not merged — and the Services tab says when the list shown is inherited from the cluster.

An agent older than the one this dashboard ships ignores the setting and keeps collecting; it takes effect once the host has updated — see Agent Rollout.

Services that do nothing until a setting is supplied​

These are detected, shown with an amber badge, and collect nothing at all until the value is there:

ServiceMissingReason shown
Garagegarage_admin_tokenadmin token missing (garage_admin_token)
PgBouncerpgbouncer_user and pgbouncer_passwordstats login missing (pgbouncer_user, pgbouncer_password)
PostgreSQLpostgresql_user and postgresql_passwordmonitoring login missing (postgresql_user, postgresql_password)
CrowdSeccrowdsec_api_keybouncer API key missing (crowdsec_api_key)

Three more depend on the host's layout rather than on a credential, and ask only when detection came up short:

ServiceMissingReason shown
nginxnginx_status_url, when the default stub_status URL did not answerstub_status not reachable at the default URL (nginx_status_url)
Traefiktraefik_metrics_url, when only the ping endpoint was foundmetrics endpoint not found (traefik_metrics_url)
PHP-FPMphpfpm_socket, when no socket matched the globFPM socket not found (phpfpm_socket)

A note on the Patroni credentials​

patroni_username and patroni_password are offered by the form, but the Patroni collector in this version sends no authentication with its requests. A Patroni REST API that requires a login therefore cannot be read, and filling these fields will not change that. Patroni's REST API only protects its write endpoints by default; the read endpoints the collector uses are normally open.

Where these are not​

Two things that look like agent settings and are not:

  • Thresholds and what alerts on a metric are the dashboard's, not the agent's. The agent reports numbers; what counts as too high lives in Alert Rules.
  • The webhook URL, the webhook secret and the collection interval are in config.yaml on the host, because the agent needs them before it can ask the dashboard anything. See Agent Installation. The one exception is the server address during a move, which the dashboard hands out with the settings — see Moving to a new server.