Agent Collectors
Collectors are detected, not configured
There is no list of collectors to maintain. Every five minutes the agent
looks at the host — the process names in /proc, the listening sockets in
/proc/net/tcp and /proc/net/tcp6, the network interfaces, the images of
the running containers — and runs a collector for everything it finds. A
service that appears on the host is picked up at the next look, without a
restart and without touching config.yaml. That describes a Linux host; a
Windows one is looked at the same way but has less to find, which
Windows hosts sets out.
A collectors: list in config.yaml is ignored. Dashboards before
agent 2.0 wrote the full list into every host's file whether the host ran
those services or not; the agent now reads the field only so it can log,
once at start, that it is ignoring it.
What you do configure is the handful of things detection cannot find out: tokens, database logins, URLs to check. Those live in the dashboard — see Agent Settings.
How detection picks an address
Every probe runs concurrently under a shared three-second deadline. A service bound to one address is looked for on that address. A service bound to the wildcard is tried on loopback first, then on the host's own addresses with WireGuard addresses before private ones before public ones. Whatever answers becomes the address the collector uses.
Versions come from the service itself where it reports one (haproxy -v,
nginx -v, cscli version, Patroni's REST API, the Docker daemon). For
services that only run in a container — Traefik, Garage, Patroni,
PostgreSQL, CrowdSec, etcd, HAProxy, nginx, PgBouncer, Caddy — the image
tag of the running container fills the gap.
The states
Every report carries, per detected service, what its collector managed to do. This is what the badges and the Info tab show.
| State | Meaning | What to do |
|---|---|---|
ok | the collector ran in the last cycle | nothing |
error | the collector failed; the error text is shown with it | read the error — it names the endpoint, the status or the file that failed. A collector that could not even be built is retried at the next refresh |
needs_config | the service was detected but a setting is missing; the reason names the setting | enter it on the system's or the cluster's Configuration tab. The agent applies it within a cycle. See Agent Settings |
none | the service was detected but no collector ran for it | this agent is older than the collector for that service; let it update (Agent Rollout) |
disabled | the service was switched off for this host on the system's Services tab; no collector runs and no setting is asked for | nothing. Detection keeps running, so the service stays listed there and can be switched back on. See Agent Settings |
Timing
All collectors run in parallel with an eight-second budget for the whole
set. A collector that overruns loses its turn and reports
timed out after 8s; it is also skipped while its previous run is still
going. Everything else goes out on time — one slow database never costs you
the heartbeat.
Anything described below as a rate compares two consecutive collections, so it is missing from the first report after a start and after a collector was rebuilt because one of its settings changed.
Host collectors
These seven always run on a Linux host. There is nothing to detect and nothing to configure. A Windows host runs a different set, and that is the larger of the two differences between the platforms — see Windows hosts.
cpu
Needs nothing. Reports busy percentage over the interval across all cores, logical CPU count, iowait and steal shares, and the 1/5/15-minute load averages.
memory
Needs nothing. Reports RAM used as a percentage plus total, used and available bytes; swap used as a percentage plus total and used bytes.
disk
Needs nothing.
Reports per mount point: space used as a percentage, total, used and
free bytes, and inode usage where the filesystem counts inodes. Per block
device behind a reported mount: read and write rates.
Leaves out pseudo filesystems (tmpfs, overlay, squashfs, proc and the
like), loop devices, /snap, partitions under 100 MiB, repeated bind
mounts of one device, and ZFS datasets other than pool roots — those belong
to the zfs collector.
network
Needs nothing.
Reports in and out byte rates per interface and totalled over the
physical ones, an error-and-drop rate per interface, and TCP socket counts
by state (established, syn_recv, time_wait, close_wait, listen).
Leaves out loopback and the virtual links of Docker and libvirt
(veth*, br-*, docker*, virbr*), which are neither reported nor
counted in the totals.
security
Needs /etc/passwd, /etc/group and who on PATH.
Reports user and group counts, active sessions and distinct logged-in
users, and one-shot markers for user and group creation and deletion,
changes to passwd/group, and logins and logouts seen in who. Emits
host.user_added and host.user_removed events from the passwd diff, and
host.group_added, host.group_removed, host.group_member_added and
host.group_member_removed from the group diff (on Windows, from the local
groups and their members). The first collection after the agent starts only
takes stock, so nothing is reported for changes made while it was stopped.
Note the change markers never appear in the first cycle after a start —
there is nothing to compare against yet.
host
Needs nothing, and every individual check is optional: timedatectl
for clock sync, apt-check (update-notifier) or apt-get for pending
updates, /proc/vmstat for OOM kills. A host without one of those simply
omits that series.
Reports whether the clock is NTP-synchronised, whether a reboot is
required (with the packages that asked for it), how many updates are
pending and how many of those are security updates, and OOM kills since
boot plus their rate.
Clock skew (clock_skew) is not collected on the host. Every report
since agent 2.11.0 carries the moment it was sent by the host's clock, and
the dashboard stores how many seconds that is off its own clock of receiving
it — either way, so a clock ahead and a clock behind read the same. It is
measured on delivery, so a report that waited in the outbox is not mistaken
for a clock that is behind. Whether the time service calls itself
synchronised is still shown, but it drops to no for a moment on hosts whose
clock is right, and the default rule watches the skew instead.
journal
Needs journalctl. A host without it is decided once and then skipped
silently — no metrics, and no error every cycle.
Reports counts of sshd's Accepted lines and sudo's COMMAND lines read
this cycle. The agent also sends each of them as a host.login or host.sudo
event; the dashboard discards those on arrival and keeps only the counts.
Note it reads forward from the cursor of its previous read, so each line
is seen exactly once whatever the interval, and a restarted agent neither
replays the journal nor misses what happened while it was down. At most 50
entries per collection become events; the counts still cover every entry.
Fails with journalctl failed: …. A cursor whose journal file was
rotated away fails the seek every time, so the collector resets to now and
reads again on the next cycle.
Service collectors
Each of these runs when — and only when — the service was detected.
docker
Detected when the Docker socket (/var/run/docker.sock, or
docker.socket in config.yaml) exists and answers /info.
Needs access to that socket. On a swarm member the role is read as
manager or worker, and the cluster id is a fingerprint of the manager
addresses — identical on every member, which is what groups them into one
cluster in the dashboard.
Reports container counts by state (running, stopped, paused, total,
restarting, unhealthy); per container its running state with image, health,
restart count, start time, exit code and whether it was OOM-killed; image,
volume and network counts and the bytes they occupy; and on swarm managers
the node and service inventory — nodes total, ready, down, draining,
managers and managers reachable, per-node readiness, service count and
desired/running/missing replicas per service.
Optional the setting docker_container_stats adds CPU and memory per
running container. It costs one stats request per container per cycle; on a
host with many containers that is the one setting here that can make the
collector slow enough to hit the eight-second budget.
Emits container.restart_loop when a container's restart count grows by
three or more within ten minutes (at most once per container per thirty
minutes) and container.died when a container that was running exits
non-zero.
Fails with the Docker API's error when the socket is not readable.
patroni
Detected when GET /patroni answers with a role or a Patroni version —
on the patroni_url setting or patroni.url in the config if set,
otherwise on every address listening on port 8008, otherwise on loopback.
Needs the REST API reachable without authentication: the collector
sends none. The detected scope becomes the host's cluster
(patroni/<scope>), and a Patroni member's cluster wins over swarm
membership.
Reports whether the node is running and its state, whether it is the
leader, its timeline, whether the cluster is unlocked (no member holds the
leader lock), how many members exist and how many are running or streaming,
and the replication lag of every replica and sync standby.
Emits patroni.promoted, patroni.demoted and, while leader,
patroni.timeline_changed.
Fails with the HTTP error or status from the REST API.
etcd
Detected when an address listening on port 2379 answers /health
healthy.
Needs the client URL reachable over plain HTTP. etcd_url overrides the
detected address. The metrics page adds detail; its absence does not fail
the collector.
Reports health (with the reason when unhealthy), whether the member sees
a leader, leader changes since the process started, backend database size
against its quota in bytes and percent, pending and failed raft proposals,
and the WAL fsync p99 — which is a lifetime figure, not a last-interval one.
pgbouncer
Detected when something listens on port 6432.
Needs pgbouncer_user and pgbouncer_password for a login listed in
PgBouncer's stats_users. Without both the state is needs_config.
pgbouncer_address overrides the detected host:port. The connection is
made with sslmode=disable to PgBouncer's own pgbouncer admin database,
and that database is excluded from every figure.
Reports per pool: active and waiting clients, active, idle and used
server connections, and the longest wait of a queued client; totals across
pools; per database: query and transaction rates and the average query time
of the last stats period; and client slots used and free against
max_client_conn.
Fails with the connection or authentication error from PgBouncer.
postgresql
Detected when something listens on port 5432, otherwise when a
postgres or postmaster process is running.
Needs postgresql_user and postgresql_password for a role holding
pg_monitor. Without both the state is needs_config. postgresql_address
and postgresql_database (default postgres) override the defaults; the
connection uses sslmode=prefer.
Reports connections total and active against max_connections, whether
the node is in recovery, the number of standbys and each one's replication
lag in bytes on a primary, replay delay on a standby, transaction-id age
against the wraparound limit, cache hit ratio, transactions open longer than
60 seconds, the size of every database the role may connect to, and rates
for transactions, deadlocks and temp file bytes.
Fails with the connection, authentication or permission error; a role
without pg_monitor shows up here.
garage
Detected when /health answers at all — 200 or 503, either proves
Garage is listening — on the garage_admin_url setting or
garage.admin_url in the config, otherwise on every address listening on
port 3903, otherwise on loopback.
Needs garage_admin_token for the v2 admin API. Without it the state is
needs_config, and Garage is the service most often found in that state.
The layout query adds the per-node series; losing it does not fail the
collector.
Reports cluster health, known, connected and storage nodes (and how many
are OK), partitions with quorum and fully replicated, the layout version,
and per node: whether it is up, its zone, address, whether it is draining,
its capacity, and the usage of its data partition where it reports one.
Also a Garage node with neither a swarm nor a Patroni cluster gets its
cluster identity from the layout (garage/<digest of the node ids>), and
the node matching this host gains the detail zone <zone>.
haproxy
Detected when a haproxy process is running.
Needs the stats socket at /run/haproxy/admin.sock (stats socket in
haproxy.cfg), readable by root. Without the socket there is nothing to
read and the collector errors.
Reports per frontend: up/down, current sessions, session rate and 5xx
rate; per backend: up/down, sessions, queued requests, servers up against
servers configured, 5xx rate; per server: up/down with its status, check
status, address and weight.
Also the addresses of every backend server are reported as the service's
targets, which is how the dashboard places this host with the cluster it
serves.
wireguard
Detected when an interface named wg* exists.
Needs wg on PATH and CAP_NET_ADMIN, which the installed unit
grants.
Reports per peer: seconds since its last handshake, whether it is stale
(no handshake within 180 seconds), and receive and transmit rates; per
interface: peer count and stale peer count. A peer that never completed a
handshake has no handshake age. Peers are labelled by interface and first
allowed IP, or by public key prefix when they have no allowed IPs.
A stale peer shows as not connected, in grey, and no default rule alerts on
it: a road-warrior client that is switched off looks exactly like one that
cannot reach the server. For a site-to-site tunnel that must stay up, add a
rule on wg_peer_stale above 0 scoped to that system.
traefik
Detected when a /metrics page carrying traefik_ series answers on an
address listening on port 8082 or on loopback, or when /ping answers
on port 8083.
Needs Prometheus metrics enabled in Traefik. If detection found only the
ping endpoint there is nothing to read, and the state is needs_config
until traefik_metrics_url is set. A URL set there gets /metrics
appended if it does not already end in it.
Reports per entrypoint: request and 5xx rates and open connections; per
service: request and 5xx rates and the health of each backend server;
whether the running configuration is the last one loaded (with the last
success and failure times); and days left on each certificate Traefik
serves.
nginx
Detected when an nginx process is running. The address is the
stub_status URL — nginx_status_url, nginx.status_url in the config, or
http://127.0.0.1/nginx_status — if it answers with Active connections.
Needs a stub_status location. nginx is detected from the process, so a
host where the default URL does not answer with HTTP 200 and a body
containing Active connections shows nginx in needs_config until
nginx_status_url points at a working one.
Reports active connections, connections reading, writing and waiting,
and rates for accepted connections, requests and dropped connections —
dropped meaning accepted but not handled, which is what a resource limit
looks like.
Setup on the host. A server block of its own on loopback, so the page is reachable from the agent and from nowhere else:
cat > /etc/nginx/conf.d/selvara-status.conf <<'EOF'
server {
listen 127.0.0.1:80;
server_name 127.0.0.1;
location = /nginx_status {
stub_status;
allow 127.0.0.1;
deny all;
access_log off;
}
}
EOF
nginx -t && systemctl reload nginx
curl -s http://127.0.0.1/nginx_status
The last line must print Active connections: …. On Debian and Ubuntu the
file can go under sites-available with a link from sites-enabled instead;
nginx reads either. If port 80 on loopback is already taken by another
server block, listen on a free port such as 127.0.0.1:8080 and set
nginx_status_url to http://127.0.0.1:8080/nginx_status on the system's
Configuration tab — see Agent Settings.
The agent looks again at its next detection, within five minutes, and the Collectors page clears with the report after it. To see it at once:
systemctl restart selvara-agent
phpfpm
Detected when a PHP-FPM master runs on the host itself. The agent reads
the configuration the master names in its process title
(php-fpm: master process (/etc/php/8.2/fpm/php-fpm.conf)) and the files it
includes, and takes each pool's listen and pm.status_path from there —
so a socket in any directory, with any name and owner, is found, and so is a
pool listening on a TCP port. Of several pools it reads the first with a
pm.status_path, else the first; the detail on the system page lists them
all. A master inside a container is not counted: its socket and
configuration are the container's, and a Docker host with a PHP container no
longer reports PHP-FPM. Only when no master names its configuration, or the
host's cannot be read, is the first match of /run/php/php*-fpm.sock tried.
Needs pm.status_path in the pool (/fpm-status is assumed for a socket
set by hand), and permission to connect to the pool's socket. The page is read
over FastCGI straight from the socket — no web server is involved. When no
pool and no socket was found the state is needs_config until
phpfpm_socket is set, as a path or host:port. A socket under /home or
/root stays out of reach: the unit hides both directories.
Reports active, idle and total worker processes with the pool name and
process manager, the listen queue, how often the pool hit pm.max_children
since it started, and rates for that same event, accepted requests and slow
requests.
Fails with php-fpm socket <path>: permission denied; the agent needs the socket's group (listen.group) or an ACL entry (listen.acl_groups) when the
socket refuses the agent. Connecting to a Unix socket needs write access to
it, and a pool socket is usually www-data:www-data 0660. The agent runs as
root but without CAP_DAC_OVERRIDE, so it is refused like any account
outside the group. The agent fixes that itself: when it finds a pool socket
whose group it is not in, it writes a drop-in with that group to its unit
and restarts (see
Agent Installation), so a
pool added later is picked up by the next detection, within five minutes.
The error only stays when that is not possible — no systemd, or a unit
started by hand.
Setup on the host. In the pool's configuration — on Debian and Ubuntu
/etc/php/<version>/fpm/pool.d/www.conf, elsewhere often
/etc/php-fpm.d/www.conf:
pm.status_path = /fpm-status
; Only when the agent cannot join the socket's group by itself (no
; systemd). Either give the agent's group an ACL entry on the socket ...
listen.acl_groups = www-data
; ... or make the socket's group one the agent is in:
; listen.group = www-data
Then reload PHP-FPM and check the socket:
systemctl reload php8.2-fpm # the unit is named after the version
ls -l /run/php/php*-fpm.sock
systemctl show selvara-agent -p SupplementaryGroups
cat /etc/systemd/system/selvara-agent.service.d/socket-groups.conf
On a host with several PHP versions or pools, set the status page in every pool at once. The agent watches the first pool socket it finds, so it is simplest when all of them answer. On Debian and Ubuntu:
# pm.status_path in every pool that does not have it yet
for f in /etc/php/*/fpm/pool.d/*.conf; do
grep -qE '^\s*pm\.status_path' "$f" || echo 'pm.status_path = /fpm-status' >> "$f"
done
# check each version's configuration and reload it
for d in /etc/php/*/fpm; do
v=$(basename "$(dirname "$d")")
php-fpm$v -t && systemctl reload php$v-fpm
done
The line is appended at the end of each file, which lands in the right pool
as long as a file holds one pool — the Debian default, www.conf. A file
with several [pool] sections needs the line in each section by hand.
The status page is reachable through the pool socket only; it becomes public
only if the web server hands /fpm-status to PHP, which the usual
location ~ \.php$ block does not.
The agent reads the page on its next cycle, and the Collectors page clears with the next full report, within five minutes. To see it at once:
systemctl restart selvara-agent
listen.acl_groups replaces the owner and group permissions of the socket
with an ACL, so list every group that needs the socket — the web server's
included — not just the agent's.
nfs
Detected when /etc/exports has at least one real entry (role
server), or an nfs/nfs4 mount exists (role client). A host that is
both gets both.
Needs nothing.
Reports the number of exports and the number of clients allowed on each;
per NFS mount whether a stat answered within two seconds and how long it
took — a hung mount shows as not OK with the timeout as its latency.
Also export clients and mount servers are reported as targets;
wildcards and netgroups are left out.
zfs
Detected when zpool list names at least one pool.
Needs zpool and zfs on PATH. The pools monitored come from the
zfs_pools setting, then zfs.pools in the config, then detection.
Reports per pool: size, allocated and free bytes, fragmentation and
capacity percentages, whether the state is ONLINE, the read, write and
checksum error counters from zpool status, and whether a scrub is running.
Per dataset, recursively: used and available bytes and used percentage.
proxmox
Detected when pvesh is installed and /etc/pve/.members names the
node — a Proxmox VE node with its cluster file system mounted. The host is
then shown as Proxmox VE with the pve-manager version instead of Debian,
and a node in a cluster joins a Proxmox VE cluster on the Clusters page.
Needs nothing: no API token, no user. The agent runs as root on the
node, and pvesh — the command-line side of the API the web interface
uses — answers root without a login. One pvesh call is made per
collection, for the resource list; quorum, guest configs, replication and
backups are read from the files Proxmox keeps them in.
Reports for this node only — every node sees the whole cluster through
the API, and reporting it from each would show every guest on every node's
page:
- per guest (VM or container, keyed by its VMID): running or not, name, type, CPU, memory, network in and out, whether it starts on boot, the age of its newest backup and the error of a backup that failed since. Templates are left out.
- the number of guests, running guests, guests set to start on boot that are stopped, and guests without a backup in the last 14 days; the age of the oldest newest backup.
- per storage the node can reach: available or not, used percentage and bytes, capacity, type and whether it is shared.
- per LVM thin pool, when the node has an LVM-thin storage: data and metadata used, every five minutes. A thin pool whose metadata runs full stops every guest on it however much data space is left, and the storage list shows only the data side. See the note below on how it is read.
- per replication job of a guest on this node: whether it is failing, the first line of its error, the failure count and the time since its last sync.
- on a cluster: whether it is quorate, nodes online and in total, and HA
resources in the
errororfencestate. A standalone node reports none of these. - whether a subscription is active, checked every six hours.
A backup that fails for a guest is a proxmox.backup_failed event naming
the guest and the error. Any other task on the node that fails — a
migration, a start, a backup job that failed before reaching a guest — is a
proxmox.task_failed event with the task's own error. A task that ends
with warnings is not failed: vzdump counts every file that changed while it
read it as a warning, and the backup is complete.
Backups are read from the task logs of the backup jobs on this node
(/var/log/pve/tasks), not from the backup storages: listing a Proxmox
Backup Server with a long retention takes minutes and often times out. The
logs record the outcome per guest, so every vzdump run counts, scheduled or
started by hand, whatever storage it wrote to. What they do not show is a
backup this node did not make — one taken on another node before a guest
migrated here counts from the first backup after the move. A guest that is
meant to have no backup, or is meant to be off, goes into the
proxmox_guests_exclude setting; it stays in the list but no longer counts
toward guests without backup, the oldest backup or autostart guests
stopped.
Replication is read from /etc/pve/replication.cfg and the state file
the replication runner keeps on each node. The API endpoint that joins the
two takes a lock on the job list even to read it, and the agent's read-only
sandbox refuses that lock.
Note on the thin pools: the percentages come from the kernel's device
mapper, which answers only to CAP_SYS_ADMIN, and the agent's unit does not
have it. The agent therefore runs that one command, lvs --nolocking,
through systemd-run as a short transient unit with
CAP_SYS_ADMIN and little else — a read-only system, no network, no new
privileges. See Agent Installation.
Note a VM's own disk usage is not reported: Proxmox reads it only through the QEMU guest agent and shows 0 without it. The storage the disk lives on is what is watched.
mdraid
Detected when /proc/mdstat lists at least one array: Linux software
RAID, which is what every Synology NAS and many plain servers run on.
Needs nothing; the kernel's own file is read.
Reports per array: whether it is healthy, how many members are marked
failed and how many slots are not up, the RAID level and up/slots count,
and the progress of a running rebuild, resync, check or reshape. An array
losing its redundancy is a critical raid.degraded event, and regaining it
raid.recovered. The RAID degraded default rule alerts after a minute.
Note a RAID 0 or linear array has no redundancy to lose and only turns
unhealthy when it goes inactive. On Synology, DSM builds its system and swap
partitions (md0, md1) as RAID 1 across every bay, empty ones included,
so there a missing slot is normal and only a member marked failed counts;
the data arrays are judged like any other.
fail2ban
Detected when a fail2ban-server process is running.
Needs fail2ban-client and its socket at
/var/run/fail2ban/fail2ban.sock, readable by root.
Reports the number of active jails, and currently failed attempts and
currently banned addresses both per jail and summed.
Fails with whatever fail2ban-client said. A refused socket is the
usual cause — the agent must be able to read it, which is one of the reasons
the unit runs as root.
Note these are current figures, not averages. A dashboard card reading
0.465 failed attempts means it is reading averaged chart data rather than
the current value.
certificates
Detected when files match /etc/letsencrypt/live/*/fullchain.pem or
/etc/haproxy/certs/*.pem.
Needs read access to those files. The certificates_paths setting (one
glob per line) replaces the default patterns, and also runs the
collector on hosts where the defaults match nothing.
Reports days left until expiry per certificate, negative once expired,
with subject, issuer, expiry date, file and SANs. Certificates are labelled
by common name, with the file name appended when two files share one.
Leaves out certificates certbot no longer manages. A file under
<root>/live/<name>/ is reported only while <root>/renewal/<name>.conf
exists. certbot delete and hand cleanups often leave the live directory
behind; nothing renews it, and its expiry would otherwise raise an alert
about a certificate no one serves. This holds for the default patterns and
for every pattern in certificates_paths alike. A renewal directory that
cannot be read does not count as missing, so the certificate is reported.
The series of a certificate that is left out, or whose file is gone, age out
of the dashboard after seven days without a report.
Note only the first CERTIFICATE block of each file is read, which in
fullchain and HAProxy bundles is the leaf — the one whose expiry matters. A
file that cannot be parsed is skipped rather than failing the collector.
On the dashboard, a system's Certificates card hides certificates that
expired more than 30 days ago and says how many it hid.
systemd
Detected when /run/systemd/system exists.
Needs systemctl.
Reports the number of failed units and up to twenty of their names
always; for units named in the units_watch setting, whether each is active
and its state; and for timers, how long since each last fired and with what
result. A name in units_watch without a suffix is read as a .service.
Timers that never fired and template instances are left out.
Emits systemd.unit_failed when a unit joins the failed set and
systemd.unit_recovered when it leaves it. Units already failed when the
agent starts are recorded, not announced.
caddy
Detected when a caddy process is running, or the admin API answers
/config/ — on caddy_admin_url, caddy.admin_url in the config, or
http://127.0.0.1:2019.
Needs the admin API reachable. The request-rate series additionally need
the metrics option on Caddy's servers; without it you get the upstreams
and nothing else, which is not an error.
Reports rates for requests, 5xx responses and request errors, requests
in flight, and per reverse-proxy upstream whether it is healthy, plus
healthy and total upstream counts.
Fails with Get "http://127.0.0.1:2019/metrics": EOF (or "connection
reset by peer") when Caddy runs in a container with port 2019 published.
Caddy binds its admin API to localhost:2019 inside the container, so
Docker's proxy accepts the connection on the host and finds nothing to hand
it to. ss -ltnp | grep 2019 then shows docker-proxy rather than caddy.
Setup on the host for Caddy in a container: have the admin API listen on every interface inside the container, keep the published port on the host's loopback, and allow the host names the agent sends:
{
admin 0.0.0.0:2019 {
origins 127.0.0.1:2019 localhost:2019
}
servers {
metrics
}
}
# docker compose
ports:
- "127.0.0.1:2019:2019" # loopback only: the admin API can change the config
Reload Caddy and check from the host with
curl -s http://127.0.0.1:2019/metrics | head. Never publish 2019 on a
public address: the admin API accepts configuration changes.
crowdsec
Detected when a crowdsec process is running, or a running container's
image repository contains crowdsec and not bouncer — crowdsecurity/crowdsec
counts, …/traefik-crowdsec-bouncer does not. For a process the address is
the first address bound to port 8080, else http://127.0.0.1:8080; for a
container it is the host port published for the container's 8080, loopback
when published on the wildcard.
Needs crowdsec_api_key — a bouncer key, from
cscli bouncers add selvara. Without it the state is needs_config.
crowdsec_lapi_url overrides the detected address. The collector reads
GET /v1/decisions with the key in X-Api-Key and a five-second timeout.
Reports whether the Local API answered, the number of active decisions,
and decisions broken down by origin (crowdsec, CAPI, cscli, lists), by
remediation type (ban, captcha) and by the ten scenarios with the most
decisions. Anything without an origin, type or scenario is counted as
unknown.
Fails with the status and the name of the setting to fix when the key is
rejected. crowdsec_lapi_up goes to 0 alongside the error, so the failure
is alertable and not only visible as a red badge.
The one collector with no service behind it
http_checks
Runs when the checks_urls setting holds at least one URL, one per
line. Nothing is detected; this is entirely yours.
Needs nothing but reachability from the host.
Reports per URL: whether it answered below 400, the HTTP status, request
latency including DNS, TCP and TLS, and — for HTTPS — days left on the
served certificate with subject, issuer and expiry.
Note each URL is fetched on a fresh connection with a seven-second
timeout and at most five redirects. A URL that fails is a series, not a
collector error: the check reports 0 and the collector stays ok. That is
what makes these alertable per URL rather than as one lump.
snmp_poller
Runs when the dashboard has assigned network devices to this host on its SNMP tab. Nothing is detected. Needs UDP port 161 from this host to each device. Reports nothing about this host itself; each device is reported as a system of its own. See SNMP Devices.
Windows hosts
A Windows host runs the same agent, reports to the same webhook, applies the same settings and updates itself from the same rollout. What differs is what there is to see, and the difference is large enough to be worth reading before an existing alert rule is pointed at one. Installing the agent there is on Agent Installation.
Detection finds fewer services
Only the probes that work over HTTP or from the listening table run on
Windows: Patroni, etcd, PgBouncer, PostgreSQL, Garage, Traefik, nginx and
Caddy. Everything else under Service collectors above is found through a
unix socket, a path under /proc or /run, or a Linux command-line tool, and
is therefore not detected on a Windows host at all: Docker, HAProxy,
WireGuard, PHP-FPM, NFS, ZFS, fail2ban, CrowdSec and systemd.
That is the answer to "why does this Windows host report no Docker". The agent
is not failing and the collector is not in error — the service simply cannot
be found from where the agent stands, so nothing about it is reported. The
file-based certificates collector is absent for the same reason; the
certstore collector below reports the same tls_days_left series from the
machine certificate store instead.
The host collectors on Windows
Fourteen collectors always run: cpu, memory, disk, network,
security, host, services, eventlog, defender, firewall,
certstore, tasks, shares and bitlocker. There is nothing to detect for
any of them. Four take a list from the dashboard — which services, which event
logs, which tasks, which certificate stores — and run without it as well.
The first six are the ones a Linux host runs, reading Windows' own counters in
place of /proc. Three of the things they report do not carry over; see
Three things that do not carry over below. There is no journal
collector — eventlog takes its place.
The eight that are new with Windows follow.
services
Needs the service control manager, which every Windows host has. The
services_watch setting names the services that get a series of their own;
the failure count needs no configuration.
Reports services_failed, the number of services set to start
automatically that are not running, with up to twenty of their names in the
metadata; and per service named in services_watch, service_active as 1 or
0 with its state, start type, process id and display name.
Emits service.started and service.stopped for a watched service.
Services already running when the agent starts are recorded, not announced.
Note the names are the ones the service control manager knows, not the
display names the Services console shows, and they are matched without regard
to case. A watched service the host does not have still gets a series, reading
not running, so a service that was uninstalled or renamed is visible instead
of quietly vanishing from the dashboard. A delayed automatic start counts as
automatic: a service set to start by itself is expected to be running either
way.
eventlog
Detected when wevtutil.exe is on the host. A host without it is decided
once and then skipped silently — no metrics, and no error every cycle.
Needs read access to the channels it reads, which for Security means
LocalSystem or membership of Event Log Readers. The installed service runs
as LocalSystem and has it. eventlog_channels replaces the three channels
read by default — Security, System and Application.
Reports eventlog_errors per channel, counting the records that channel
itself called critical or an error this cycle, and eventlog_logins,
eventlog_failed_logins and eventlog_privileged summed across the channels
read.
Emits host.user_added and host.user_removed from the account records,
host.audit_log_cleared when the security audit log is cleared,
host.unexpected_shutdown when the host stopped without shutting down —
once, although Windows records such a stop twice, as Kernel-Power 41 and as
6008 — and host.rebooted when it booted or a shutdown was requested.
The logon records also become host.login and host.login_failed events on
the agent's side; the dashboard discards them on arrival, and the counts above
are what remains of them.
Fails with the error per channel, named by channel. A channel that cannot
be read does not stop the others: what the rest yielded is still reported, and
the failure is what turns the collector red.
Note each channel is read forward from the record the collector last saw,
so a record is seen exactly once whatever the interval; reading starts at the
agent's own start, so a restarted agent does not replay the channel. At most
200 records per channel per read and at most 50 events per collection become
events; the counts still cover every record read.
defender
Detected when Get-MpComputerStatus exists. A Server Core host without
the feature, or an edition that does not carry Defender, is decided once and
reports nothing rather than failing every cycle.
Needs nothing — the cmdlet answers without elevation.
Reports defender_enabled as 1 or 0 with the running mode, whether
real-time protection is on and whether tamper protection is on;
defender_signature_age, the age in days of the older of the two definition
sets, because the host is only as current as whichever was updated last; and
defender_scan_age, the age in days of the more recent of the quick and full
scans, with both in the metadata.
Emits host.defender_disabled and host.defender_enabled. The state at
the agent's start is recorded, not announced.
Note passive mode counts as protected. Defender steps aside when
another antivirus registers itself: the service keeps running and scans on
demand but not in real time, which is how the two are meant to coexist. The
mode says which way the host is protected. Reading passive as "off" would
alarm on every host with a third-party product installed. A definition set or
a scan that never happened is left out rather than reported as a number. The
status is read every five minutes beside the collection, not inside it.
firewall
Detected when Get-NetFirewallProfile exists. A host without it is
decided once and reports nothing rather than failing every cycle.
Needs nothing.
Reports firewall_enabled as 1 or 0 per profile — domain, private,
public — with the default inbound action in the metadata as block, allow
or default.
Emits host.firewall_disabled and host.firewall_enabled per profile.
The state at the agent's start is recorded, not announced.
Note the value reported is the one in force, Group Policy included. A
profile nobody ever configured reports its inbound action as default, which
is the common case rather than an oddity: Windows still blocks unsolicited
inbound traffic there, so it is not an open profile. The profiles are read
every five minutes beside the collection.
certstore
Needs read access to the machine certificate stores. certstore_stores
adds stores to the one read on every host, My; WebHosting and Root
are the usual additions.
Reports tls_days_left per certificate, negative once expired, with
subject, issuer, expiry date, store, SANs and thumbprint. From My only the
certificates the store holds a private key for are reported — the ones the
host itself serves with. Intermediates and certificates other software put
there have no key in the store, and their expiry is nobody's to act on. A
store added through certstore_stores is reported in full. Certificates are
labelled by common name, with the thumbprint appended when two share one.
Fails with the store path and the error, for a store that could not be
opened. The stores that could be read still report; the failure is what turns
the collector red. A store that is merely empty is no error, and a certificate
that cannot be parsed is skipped rather than failing the collector.
Note Root is deliberately not read by default: a machine's root store
carries dozens of CA certificates, expired ones among them, and they would
bury the certificate the host actually serves. A certificate held in two of
the configured stores is one series, not two.
tasks
Needs PowerShell and the task scheduler. tasks_watch names the tasks
that get an age of their own; the failure count needs no configuration.
Reports tasks_failed, the number of tasks whose last run ended badly,
with up to twenty of their names in the metadata; and per task named in
tasks_watch, task_last_run_age in seconds with the run's outcome as the
hexadecimal status Windows documents it by.
Emits task.failed and task.recovered. Tasks already failing when the
agent starts are recorded, not announced.
Note a task is written the way Task Scheduler addresses it,
\Folder\Task; a name without a folder means one in the root folder.
Disabled tasks are left out entirely. Tasks under \Microsoft\ are left
out of the failure count — Windows' own housekeeping fails and retries on
its own schedule and is nobody's alarm. A result the scheduler wrote about the
task rather than the program — running right now, never run yet, no runs left
— is not a failure either, or a busy host would be permanently in alarm. The
library is walked every five minutes beside the collection, because asking the
task service about every task costs far more than a collection may spend.
shares
Needs the SMB server (Get-SmbShare).
Reports smb_shares, the number of shares this host publishes, and
smb_share_sessions per share with the path it points at.
Emits host.share_added and host.share_removed. A share nobody expected
is worth knowing about, and so is one that went away: whatever reached it
stopped working at that moment. The shares a host already had when the agent
started are its configuration, not a change to it, and raise nothing.
Note the administrative shares are neither counted nor reported: the drive
roots, ADMIN$ and IPC$, everything whose name ends in a dollar sign.
Nobody created them and they exist on every host. The shares are read every
five minutes beside the collection.
bitlocker
Detected when Get-BitLockerVolume exists. An edition without BitLocker,
or a Server Core host without the feature, is decided once and reports nothing
rather than failing every cycle. A host that has the cmdlet and refused the
read is not that case: that is an error, it is reported, and it is tried
again — the service runs with privileges a hand-run command did not have.
Reports bitlocker_protected as 1 or 0 per mount point, and
bitlocker_encrypted_percent per mount point for an encryption still running.
Emits host.bitlocker_off when a volume that was protected no longer is.
A volume that stops being listed was detached rather than unlocked and raises
nothing.
Note the status and method in the metadata are the raw integers the
BitLocker API returns, not words. The tables that name those values could
not be confirmed against a host, so they are passed through untranslated
rather than guessed at. The volumes are read every fifteen minutes beside the
collection — whether a volume is encrypted changes about as often as the
machine is rebuilt.
Three things that do not carry over
A rule written against a Linux fleet can be wrong on a Windows host in three ways. None of them is a fault in the agent; all three are what the platform does or does not have.
Swap is the commit charge, not the page file. swap and swap_used on
Windows report how much of the commit limit is committed, which counts every
reservation whether or not a page was ever written to disk. Sixty per cent
is an ordinary idle Windows host and does not mean it is swapping. An alert
rule that fires above, say, 50 % swap on Linux will fire permanently here.
Only drive letters are reported. disk covers the fixed drives and
nothing else: removable drives, mapped network drives and optical drives are
left out, because a mapped drive reported here would be this host's disk on
the dashboard as well as the file server's. A volume mounted into a folder
rather than given a letter is invisible to the collector, so a host that
keeps its data on one will not show that space anywhere.
There is no load average, and no iowait or steal. Windows has no run
queue of the kind load_1m names, and its CPU times have no iowait or steal
buckets. Those series are therefore absent on a Windows host rather than
reported as zero — load_1m, load_5m, load_15m, cpu_iowait and
cpu_steal are simply not in the report. An alert rule that watches one of
them will never fire on a Windows host and never tell you it is not watching.
cpu itself, the busy percentage, is reported as everywhere.
Two smaller absences belong here as well. updates_pending comes from a
Windows Update search, which takes minutes and therefore runs every six
hours rather than every cycle, so the figure can be that old. And there are
no oom_kills_total or oom_kills_rate: Windows trims working sets and fails
allocations instead of killing a process to reclaim memory, so there is no
counter to read and a zero would claim more than the agent knows.
The first cycle after a start is thin
Six of the checks cost far more than a collection may spend and run on their own schedule beside it: Defender, the firewall, the scheduled tasks, the shares, BitLocker and the Windows Update count. The first report after a start carries nothing for them — not a zero, nothing — and their values appear from the next cycle. A host that has just been installed and shows none of them is not misconfigured; give it a cycle.
Reading the Security event log is privileged
This is the one that catches people out while testing. The Security channel is privileged: it takes LocalSystem, or membership of Event Log Readers. The installed service runs as LocalSystem and has no trouble with it.
An agent started by hand in an ordinary console does not have either, and the
eventlog collector then shows as failing with a message like
Security: exit status 5: Zugriff verweigert
— the wording follows the host's display language. The other channels are still read and still reported; it is the Security one that fails, and it turns the whole collector red.
So judge that collector from the installed service, not from a hand-run binary. If you do want to run the agent by hand against the Security channel, put the account you run it as into Event Log Readers first.
Reading collector state
A system's Info tab has two places to look.
Detected services shows one badge per service the agent found, tinted by
collector state — green for ok, red for error, amber for
needs_config, plain for none. A service in disabled gets no badge; it
is listed on the Services tab instead. Hovering a badge gives the version, the
state in words and the error text.
Collectors lists every collector that reported in the last cycle, with
Running or Failing and the error. This card is built from the
collector_ok series each collector writes every cycle, so it covers the
host collectors too, not only detected services. A collector silent for ten
minutes while the host kept reporting drops off the list — it no longer
runs, and showing its last error would be showing a collector that is not
there.
The Collectors page (/collectors, in the sidebar; the Needs
configuration and Collector error tiles on the overview lead there too) answers the question
across every host at once. It has a tab for collectors waiting for
configuration and one for collectors in error, and groups them by service
and then by the reason the agent gave — hosts that fail the same way usually
need the same setting, and if they form a cluster it can be set on the
cluster once. Each host links to its Configuration tab.
The Agents page rolls this up across all hosts: one row per system with an error count and a "needs configuration" count, both hoverable for the service names, and a filter for With collector problems. See Agent Rollout.
On the host itself, journalctl -u selvara-agent -f logs
Collector X failing: … once when a collector starts failing and
Collector X recovered when it stops. Those transitions also become
collector.failed and collector.recovered events, so they can notify —
see Event Rules.
What "needs config" means in practice
needs_config is not a fault. It means the agent found a service, knows
exactly which collector belongs to it, and is missing one value it cannot
discover: a token, a login, a socket path, a metrics URL. Nothing is
collected for that service until you supply it, and nothing else is
affected.
The state carries the reason, and the reason names the setting key —
admin token missing (garage_admin_token),
stats login missing (pgbouncer_user, pgbouncer_password),
metrics endpoint not found (traefik_metrics_url). Enter it on the
system's Configuration tab, or on the cluster's when it is the same
value for every member, and the agent applies it within a cycle without a
restart.
Full reference: Agent Settings.