Skip to main content

Agent Collectors

Collectors are detected, not configured​

There is no list of collectors to maintain. Every five minutes the agent looks at the host — the process names in /proc, the listening sockets in /proc/net/tcp and /proc/net/tcp6, the network interfaces, the images of the running containers — and runs a collector for everything it finds. A service that appears on the host is picked up at the next look, without a restart and without touching config.yaml. That describes a Linux host; a Windows one is looked at the same way but has less to find, which Windows hosts sets out.

A collectors: list in config.yaml is ignored. Dashboards before agent 2.0 wrote the full list into every host's file whether the host ran those services or not; the agent now reads the field only so it can log, once at start, that it is ignoring it.

What you do configure is the handful of things detection cannot find out: tokens, database logins, URLs to check. Those live in the dashboard — see Agent Settings.

How detection picks an address​

Every probe runs concurrently under a shared three-second deadline. A service bound to one address is looked for on that address. A service bound to the wildcard is tried on loopback first, then on the host's own addresses with WireGuard addresses before private ones before public ones. Whatever answers becomes the address the collector uses.

Versions come from the service itself where it reports one (haproxy -v, nginx -v, cscli version, Patroni's REST API, the Docker daemon). For services that only run in a container — Traefik, Garage, Patroni, PostgreSQL, CrowdSec, etcd, HAProxy, nginx, PgBouncer, Caddy — the image tag of the running container fills the gap.

The states​

Every report carries, per detected service, what its collector managed to do. This is what the badges and the Info tab show.

StateMeaningWhat to do
okthe collector ran in the last cyclenothing
errorthe collector failed; the error text is shown with itread the error — it names the endpoint, the status or the file that failed. A collector that could not even be built is retried at the next refresh
needs_configthe service was detected but a setting is missing; the reason names the settingenter it on the system's or the cluster's Configuration tab. The agent applies it within a cycle. See Agent Settings
nonethe service was detected but no collector ran for itthis agent is older than the collector for that service; let it update (Agent Rollout)
disabledthe service was switched off for this host on the system's Services tab; no collector runs and no setting is asked fornothing. Detection keeps running, so the service stays listed there and can be switched back on. See Agent Settings

Timing​

All collectors run in parallel with an eight-second budget for the whole set. A collector that overruns loses its turn and reports timed out after 8s; it is also skipped while its previous run is still going. Everything else goes out on time — one slow database never costs you the heartbeat.

Anything described below as a rate compares two consecutive collections, so it is missing from the first report after a start and after a collector was rebuilt because one of its settings changed.

Host collectors​

These seven always run on a Linux host. There is nothing to detect and nothing to configure. A Windows host runs a different set, and that is the larger of the two differences between the platforms — see Windows hosts.

cpu​

Needs nothing. Reports busy percentage over the interval across all cores, logical CPU count, iowait and steal shares, and the 1/5/15-minute load averages.

memory​

Needs nothing. Reports RAM used as a percentage plus total, used and available bytes; swap used as a percentage plus total and used bytes.

disk​

Needs nothing. Reports per mount point: space used as a percentage, total, used and free bytes, and inode usage where the filesystem counts inodes. Per block device behind a reported mount: read and write rates. Leaves out pseudo filesystems (tmpfs, overlay, squashfs, proc and the like), loop devices, /snap, partitions under 100 MiB, repeated bind mounts of one device, and ZFS datasets other than pool roots — those belong to the zfs collector.

network​

Needs nothing. Reports in and out byte rates per interface and totalled over the physical ones, an error-and-drop rate per interface, and TCP socket counts by state (established, syn_recv, time_wait, close_wait, listen). Leaves out loopback and the virtual links of Docker and libvirt (veth*, br-*, docker*, virbr*), which are neither reported nor counted in the totals.

security​

Needs /etc/passwd, /etc/group and who on PATH. Reports user and group counts, active sessions and distinct logged-in users, and one-shot markers for user and group creation and deletion, changes to passwd/group, and logins and logouts seen in who. Emits host.user_added and host.user_removed events from the passwd diff, and host.group_added, host.group_removed, host.group_member_added and host.group_member_removed from the group diff (on Windows, from the local groups and their members). The first collection after the agent starts only takes stock, so nothing is reported for changes made while it was stopped. Note the change markers never appear in the first cycle after a start — there is nothing to compare against yet.

host​

Needs nothing, and every individual check is optional: timedatectl for clock sync, apt-check (update-notifier) or apt-get for pending updates, /proc/vmstat for OOM kills. A host without one of those simply omits that series. Reports whether the clock is NTP-synchronised, whether a reboot is required (with the packages that asked for it), how many updates are pending and how many of those are security updates, and OOM kills since boot plus their rate.

Clock skew (clock_skew) is not collected on the host. Every report since agent 2.11.0 carries the moment it was sent by the host's clock, and the dashboard stores how many seconds that is off its own clock of receiving it — either way, so a clock ahead and a clock behind read the same. It is measured on delivery, so a report that waited in the outbox is not mistaken for a clock that is behind. Whether the time service calls itself synchronised is still shown, but it drops to no for a moment on hosts whose clock is right, and the default rule watches the skew instead.

journal​

Needs journalctl. A host without it is decided once and then skipped silently — no metrics, and no error every cycle. Reports counts of sshd's Accepted lines and sudo's COMMAND lines read this cycle. The agent also sends each of them as a host.login or host.sudo event; the dashboard discards those on arrival and keeps only the counts. Note it reads forward from the cursor of its previous read, so each line is seen exactly once whatever the interval, and a restarted agent neither replays the journal nor misses what happened while it was down. At most 50 entries per collection become events; the counts still cover every entry. Fails with journalctl failed: …. A cursor whose journal file was rotated away fails the seek every time, so the collector resets to now and reads again on the next cycle.

Service collectors​

Each of these runs when — and only when — the service was detected.

docker​

Detected when the Docker socket (/var/run/docker.sock, or docker.socket in config.yaml) exists and answers /info. Needs access to that socket. On a swarm member the role is read as manager or worker, and the cluster id is a fingerprint of the manager addresses — identical on every member, which is what groups them into one cluster in the dashboard. Reports container counts by state (running, stopped, paused, total, restarting, unhealthy); per container its running state with image, health, restart count, start time, exit code and whether it was OOM-killed; image, volume and network counts and the bytes they occupy; and on swarm managers the node and service inventory — nodes total, ready, down, draining, managers and managers reachable, per-node readiness, service count and desired/running/missing replicas per service. Optional the setting docker_container_stats adds CPU and memory per running container. It costs one stats request per container per cycle; on a host with many containers that is the one setting here that can make the collector slow enough to hit the eight-second budget. Emits container.restart_loop when a container's restart count grows by three or more within ten minutes (at most once per container per thirty minutes) and container.died when a container that was running exits non-zero. Fails with the Docker API's error when the socket is not readable.

patroni​

Detected when GET /patroni answers with a role or a Patroni version — on the patroni_url setting or patroni.url in the config if set, otherwise on every address listening on port 8008, otherwise on loopback. Needs the REST API reachable without authentication: the collector sends none. The detected scope becomes the host's cluster (patroni/<scope>), and a Patroni member's cluster wins over swarm membership. Reports whether the node is running and its state, whether it is the leader, its timeline, whether the cluster is unlocked (no member holds the leader lock), how many members exist and how many are running or streaming, and the replication lag of every replica and sync standby. Emits patroni.promoted, patroni.demoted and, while leader, patroni.timeline_changed. Fails with the HTTP error or status from the REST API.

etcd​

Detected when an address listening on port 2379 answers /health healthy. Needs the client URL reachable over plain HTTP. etcd_url overrides the detected address. The metrics page adds detail; its absence does not fail the collector. Reports health (with the reason when unhealthy), whether the member sees a leader, leader changes since the process started, backend database size against its quota in bytes and percent, pending and failed raft proposals, and the WAL fsync p99 — which is a lifetime figure, not a last-interval one.

pgbouncer​

Detected when something listens on port 6432. Needs pgbouncer_user and pgbouncer_password for a login listed in PgBouncer's stats_users. Without both the state is needs_config. pgbouncer_address overrides the detected host:port. The connection is made with sslmode=disable to PgBouncer's own pgbouncer admin database, and that database is excluded from every figure. Reports per pool: active and waiting clients, active, idle and used server connections, and the longest wait of a queued client; totals across pools; per database: query and transaction rates and the average query time of the last stats period; and client slots used and free against max_client_conn. Fails with the connection or authentication error from PgBouncer.

postgresql​

Detected when something listens on port 5432, otherwise when a postgres or postmaster process is running. Needs postgresql_user and postgresql_password for a role holding pg_monitor. Without both the state is needs_config. postgresql_address and postgresql_database (default postgres) override the defaults; the connection uses sslmode=prefer. Reports connections total and active against max_connections, whether the node is in recovery, the number of standbys and each one's replication lag in bytes on a primary, replay delay on a standby, transaction-id age against the wraparound limit, cache hit ratio, transactions open longer than 60 seconds, the size of every database the role may connect to, and rates for transactions, deadlocks and temp file bytes. Fails with the connection, authentication or permission error; a role without pg_monitor shows up here.

garage​

Detected when /health answers at all — 200 or 503, either proves Garage is listening — on the garage_admin_url setting or garage.admin_url in the config, otherwise on every address listening on port 3903, otherwise on loopback. Needs garage_admin_token for the v2 admin API. Without it the state is needs_config, and Garage is the service most often found in that state. The layout query adds the per-node series; losing it does not fail the collector. Reports cluster health, known, connected and storage nodes (and how many are OK), partitions with quorum and fully replicated, the layout version, and per node: whether it is up, its zone, address, whether it is draining, its capacity, and the usage of its data partition where it reports one. Also a Garage node with neither a swarm nor a Patroni cluster gets its cluster identity from the layout (garage/<digest of the node ids>), and the node matching this host gains the detail zone <zone>.

haproxy​

Detected when a haproxy process is running. Needs the stats socket at /run/haproxy/admin.sock (stats socket in haproxy.cfg), readable by root. Without the socket there is nothing to read and the collector errors. Reports per frontend: up/down, current sessions, session rate and 5xx rate; per backend: up/down, sessions, queued requests, servers up against servers configured, 5xx rate; per server: up/down with its status, check status, address and weight. Also the addresses of every backend server are reported as the service's targets, which is how the dashboard places this host with the cluster it serves.

wireguard​

Detected when an interface named wg* exists. Needs wg on PATH and CAP_NET_ADMIN, which the installed unit grants. Reports per peer: seconds since its last handshake, whether it is stale (no handshake within 180 seconds), and receive and transmit rates; per interface: peer count and stale peer count. A peer that never completed a handshake has no handshake age. Peers are labelled by interface and first allowed IP, or by public key prefix when they have no allowed IPs.

A stale peer shows as not connected, in grey, and no default rule alerts on it: a road-warrior client that is switched off looks exactly like one that cannot reach the server. For a site-to-site tunnel that must stay up, add a rule on wg_peer_stale above 0 scoped to that system.

traefik​

Detected when a /metrics page carrying traefik_ series answers on an address listening on port 8082 or on loopback, or when /ping answers on port 8083. Needs Prometheus metrics enabled in Traefik. If detection found only the ping endpoint there is nothing to read, and the state is needs_config until traefik_metrics_url is set. A URL set there gets /metrics appended if it does not already end in it. Reports per entrypoint: request and 5xx rates and open connections; per service: request and 5xx rates and the health of each backend server; whether the running configuration is the last one loaded (with the last success and failure times); and days left on each certificate Traefik serves.

nginx​

Detected when an nginx process is running. The address is the stub_status URL — nginx_status_url, nginx.status_url in the config, or http://127.0.0.1/nginx_status — if it answers with Active connections. Needs a stub_status location. nginx is detected from the process, so a host where the default URL does not answer with HTTP 200 and a body containing Active connections shows nginx in needs_config until nginx_status_url points at a working one. Reports active connections, connections reading, writing and waiting, and rates for accepted connections, requests and dropped connections — dropped meaning accepted but not handled, which is what a resource limit looks like.

Setup on the host. A server block of its own on loopback, so the page is reachable from the agent and from nowhere else:

cat > /etc/nginx/conf.d/selvara-status.conf <<'EOF'
server {
listen 127.0.0.1:80;
server_name 127.0.0.1;

location = /nginx_status {
stub_status;
allow 127.0.0.1;
deny all;
access_log off;
}
}
EOF
nginx -t && systemctl reload nginx
curl -s http://127.0.0.1/nginx_status

The last line must print Active connections: …. On Debian and Ubuntu the file can go under sites-available with a link from sites-enabled instead; nginx reads either. If port 80 on loopback is already taken by another server block, listen on a free port such as 127.0.0.1:8080 and set nginx_status_url to http://127.0.0.1:8080/nginx_status on the system's Configuration tab — see Agent Settings.

The agent looks again at its next detection, within five minutes, and the Collectors page clears with the report after it. To see it at once:

systemctl restart selvara-agent

phpfpm​

Detected when a PHP-FPM master runs on the host itself. The agent reads the configuration the master names in its process title (php-fpm: master process (/etc/php/8.2/fpm/php-fpm.conf)) and the files it includes, and takes each pool's listen and pm.status_path from there — so a socket in any directory, with any name and owner, is found, and so is a pool listening on a TCP port. Of several pools it reads the first with a pm.status_path, else the first; the detail on the system page lists them all. A master inside a container is not counted: its socket and configuration are the container's, and a Docker host with a PHP container no longer reports PHP-FPM. Only when no master names its configuration, or the host's cannot be read, is the first match of /run/php/php*-fpm.sock tried. Needs pm.status_path in the pool (/fpm-status is assumed for a socket set by hand), and permission to connect to the pool's socket. The page is read over FastCGI straight from the socket — no web server is involved. When no pool and no socket was found the state is needs_config until phpfpm_socket is set, as a path or host:port. A socket under /home or /root stays out of reach: the unit hides both directories. Reports active, idle and total worker processes with the pool name and process manager, the listen queue, how often the pool hit pm.max_children since it started, and rates for that same event, accepted requests and slow requests. Fails with php-fpm socket <path>: permission denied; the agent needs the socket's group (listen.group) or an ACL entry (listen.acl_groups) when the socket refuses the agent. Connecting to a Unix socket needs write access to it, and a pool socket is usually www-data:www-data 0660. The agent runs as root but without CAP_DAC_OVERRIDE, so it is refused like any account outside the group. The agent fixes that itself: when it finds a pool socket whose group it is not in, it writes a drop-in with that group to its unit and restarts (see Agent Installation), so a pool added later is picked up by the next detection, within five minutes. The error only stays when that is not possible — no systemd, or a unit started by hand.

Setup on the host. In the pool's configuration — on Debian and Ubuntu /etc/php/<version>/fpm/pool.d/www.conf, elsewhere often /etc/php-fpm.d/www.conf:

pm.status_path = /fpm-status

; Only when the agent cannot join the socket's group by itself (no
; systemd). Either give the agent's group an ACL entry on the socket ...
listen.acl_groups = www-data
; ... or make the socket's group one the agent is in:
; listen.group = www-data

Then reload PHP-FPM and check the socket:

systemctl reload php8.2-fpm # the unit is named after the version
ls -l /run/php/php*-fpm.sock
systemctl show selvara-agent -p SupplementaryGroups
cat /etc/systemd/system/selvara-agent.service.d/socket-groups.conf

On a host with several PHP versions or pools, set the status page in every pool at once. The agent watches the first pool socket it finds, so it is simplest when all of them answer. On Debian and Ubuntu:

# pm.status_path in every pool that does not have it yet
for f in /etc/php/*/fpm/pool.d/*.conf; do
grep -qE '^\s*pm\.status_path' "$f" || echo 'pm.status_path = /fpm-status' >> "$f"
done

# check each version's configuration and reload it
for d in /etc/php/*/fpm; do
v=$(basename "$(dirname "$d")")
php-fpm$v -t && systemctl reload php$v-fpm
done

The line is appended at the end of each file, which lands in the right pool as long as a file holds one pool — the Debian default, www.conf. A file with several [pool] sections needs the line in each section by hand.

The status page is reachable through the pool socket only; it becomes public only if the web server hands /fpm-status to PHP, which the usual location ~ \.php$ block does not.

The agent reads the page on its next cycle, and the Collectors page clears with the next full report, within five minutes. To see it at once:

systemctl restart selvara-agent

listen.acl_groups replaces the owner and group permissions of the socket with an ACL, so list every group that needs the socket — the web server's included — not just the agent's.

nfs​

Detected when /etc/exports has at least one real entry (role server), or an nfs/nfs4 mount exists (role client). A host that is both gets both. Needs nothing. Reports the number of exports and the number of clients allowed on each; per NFS mount whether a stat answered within two seconds and how long it took — a hung mount shows as not OK with the timeout as its latency. Also export clients and mount servers are reported as targets; wildcards and netgroups are left out.

zfs​

Detected when zpool list names at least one pool. Needs zpool and zfs on PATH. The pools monitored come from the zfs_pools setting, then zfs.pools in the config, then detection. Reports per pool: size, allocated and free bytes, fragmentation and capacity percentages, whether the state is ONLINE, the read, write and checksum error counters from zpool status, and whether a scrub is running. Per dataset, recursively: used and available bytes and used percentage.

proxmox​

Detected when pvesh is installed and /etc/pve/.members names the node — a Proxmox VE node with its cluster file system mounted. The host is then shown as Proxmox VE with the pve-manager version instead of Debian, and a node in a cluster joins a Proxmox VE cluster on the Clusters page. Needs nothing: no API token, no user. The agent runs as root on the node, and pvesh — the command-line side of the API the web interface uses — answers root without a login. One pvesh call is made per collection, for the resource list; quorum, guest configs, replication and backups are read from the files Proxmox keeps them in. Reports for this node only — every node sees the whole cluster through the API, and reporting it from each would show every guest on every node's page:

  • per guest (VM or container, keyed by its VMID): running or not, name, type, CPU, memory, network in and out, whether it starts on boot, the age of its newest backup and the error of a backup that failed since. Templates are left out.
  • the number of guests, running guests, guests set to start on boot that are stopped, and guests without a backup in the last 14 days; the age of the oldest newest backup.
  • per storage the node can reach: available or not, used percentage and bytes, capacity, type and whether it is shared.
  • per LVM thin pool, when the node has an LVM-thin storage: data and metadata used, every five minutes. A thin pool whose metadata runs full stops every guest on it however much data space is left, and the storage list shows only the data side. See the note below on how it is read.
  • per replication job of a guest on this node: whether it is failing, the first line of its error, the failure count and the time since its last sync.
  • on a cluster: whether it is quorate, nodes online and in total, and HA resources in the error or fence state. A standalone node reports none of these.
  • whether a subscription is active, checked every six hours.

A backup that fails for a guest is a proxmox.backup_failed event naming the guest and the error. Any other task on the node that fails — a migration, a start, a backup job that failed before reaching a guest — is a proxmox.task_failed event with the task's own error. A task that ends with warnings is not failed: vzdump counts every file that changed while it read it as a warning, and the backup is complete.

Backups are read from the task logs of the backup jobs on this node (/var/log/pve/tasks), not from the backup storages: listing a Proxmox Backup Server with a long retention takes minutes and often times out. The logs record the outcome per guest, so every vzdump run counts, scheduled or started by hand, whatever storage it wrote to. What they do not show is a backup this node did not make — one taken on another node before a guest migrated here counts from the first backup after the move. A guest that is meant to have no backup, or is meant to be off, goes into the proxmox_guests_exclude setting; it stays in the list but no longer counts toward guests without backup, the oldest backup or autostart guests stopped.

Replication is read from /etc/pve/replication.cfg and the state file the replication runner keeps on each node. The API endpoint that joins the two takes a lock on the job list even to read it, and the agent's read-only sandbox refuses that lock.

Note on the thin pools: the percentages come from the kernel's device mapper, which answers only to CAP_SYS_ADMIN, and the agent's unit does not have it. The agent therefore runs that one command, lvs --nolocking, through systemd-run as a short transient unit with CAP_SYS_ADMIN and little else — a read-only system, no network, no new privileges. See Agent Installation.

Note a VM's own disk usage is not reported: Proxmox reads it only through the QEMU guest agent and shows 0 without it. The storage the disk lives on is what is watched.

mdraid​

Detected when /proc/mdstat lists at least one array: Linux software RAID, which is what every Synology NAS and many plain servers run on. Needs nothing; the kernel's own file is read. Reports per array: whether it is healthy, how many members are marked failed and how many slots are not up, the RAID level and up/slots count, and the progress of a running rebuild, resync, check or reshape. An array losing its redundancy is a critical raid.degraded event, and regaining it raid.recovered. The RAID degraded default rule alerts after a minute. Note a RAID 0 or linear array has no redundancy to lose and only turns unhealthy when it goes inactive. On Synology, DSM builds its system and swap partitions (md0, md1) as RAID 1 across every bay, empty ones included, so there a missing slot is normal and only a member marked failed counts; the data arrays are judged like any other.

fail2ban​

Detected when a fail2ban-server process is running. Needs fail2ban-client and its socket at /var/run/fail2ban/fail2ban.sock, readable by root. Reports the number of active jails, and currently failed attempts and currently banned addresses both per jail and summed. Fails with whatever fail2ban-client said. A refused socket is the usual cause — the agent must be able to read it, which is one of the reasons the unit runs as root. Note these are current figures, not averages. A dashboard card reading 0.465 failed attempts means it is reading averaged chart data rather than the current value.

certificates​

Detected when files match /etc/letsencrypt/live/*/fullchain.pem or /etc/haproxy/certs/*.pem. Needs read access to those files. The certificates_paths setting (one glob per line) replaces the default patterns, and also runs the collector on hosts where the defaults match nothing. Reports days left until expiry per certificate, negative once expired, with subject, issuer, expiry date, file and SANs. Certificates are labelled by common name, with the file name appended when two files share one. Leaves out certificates certbot no longer manages. A file under <root>/live/<name>/ is reported only while <root>/renewal/<name>.conf exists. certbot delete and hand cleanups often leave the live directory behind; nothing renews it, and its expiry would otherwise raise an alert about a certificate no one serves. This holds for the default patterns and for every pattern in certificates_paths alike. A renewal directory that cannot be read does not count as missing, so the certificate is reported. The series of a certificate that is left out, or whose file is gone, age out of the dashboard after seven days without a report. Note only the first CERTIFICATE block of each file is read, which in fullchain and HAProxy bundles is the leaf — the one whose expiry matters. A file that cannot be parsed is skipped rather than failing the collector. On the dashboard, a system's Certificates card hides certificates that expired more than 30 days ago and says how many it hid.

systemd​

Detected when /run/systemd/system exists. Needs systemctl. Reports the number of failed units and up to twenty of their names always; for units named in the units_watch setting, whether each is active and its state; and for timers, how long since each last fired and with what result. A name in units_watch without a suffix is read as a .service. Timers that never fired and template instances are left out. Emits systemd.unit_failed when a unit joins the failed set and systemd.unit_recovered when it leaves it. Units already failed when the agent starts are recorded, not announced.

caddy​

Detected when a caddy process is running, or the admin API answers /config/ — on caddy_admin_url, caddy.admin_url in the config, or http://127.0.0.1:2019. Needs the admin API reachable. The request-rate series additionally need the metrics option on Caddy's servers; without it you get the upstreams and nothing else, which is not an error. Reports rates for requests, 5xx responses and request errors, requests in flight, and per reverse-proxy upstream whether it is healthy, plus healthy and total upstream counts. Fails with Get "http://127.0.0.1:2019/metrics": EOF (or "connection reset by peer") when Caddy runs in a container with port 2019 published. Caddy binds its admin API to localhost:2019 inside the container, so Docker's proxy accepts the connection on the host and finds nothing to hand it to. ss -ltnp | grep 2019 then shows docker-proxy rather than caddy.

Setup on the host for Caddy in a container: have the admin API listen on every interface inside the container, keep the published port on the host's loopback, and allow the host names the agent sends:

{
admin 0.0.0.0:2019 {
origins 127.0.0.1:2019 localhost:2019
}
servers {
metrics
}
}
# docker compose
ports:
- "127.0.0.1:2019:2019" # loopback only: the admin API can change the config

Reload Caddy and check from the host with curl -s http://127.0.0.1:2019/metrics | head. Never publish 2019 on a public address: the admin API accepts configuration changes.

crowdsec​

Detected when a crowdsec process is running, or a running container's image repository contains crowdsec and not bouncer — crowdsecurity/crowdsec counts, …/traefik-crowdsec-bouncer does not. For a process the address is the first address bound to port 8080, else http://127.0.0.1:8080; for a container it is the host port published for the container's 8080, loopback when published on the wildcard. Needs crowdsec_api_key — a bouncer key, from cscli bouncers add selvara. Without it the state is needs_config. crowdsec_lapi_url overrides the detected address. The collector reads GET /v1/decisions with the key in X-Api-Key and a five-second timeout. Reports whether the Local API answered, the number of active decisions, and decisions broken down by origin (crowdsec, CAPI, cscli, lists), by remediation type (ban, captcha) and by the ten scenarios with the most decisions. Anything without an origin, type or scenario is counted as unknown. Fails with the status and the name of the setting to fix when the key is rejected. crowdsec_lapi_up goes to 0 alongside the error, so the failure is alertable and not only visible as a red badge.

The one collector with no service behind it​

http_checks​

Runs when the checks_urls setting holds at least one URL, one per line. Nothing is detected; this is entirely yours. Needs nothing but reachability from the host. Reports per URL: whether it answered below 400, the HTTP status, request latency including DNS, TCP and TLS, and — for HTTPS — days left on the served certificate with subject, issuer and expiry. Note each URL is fetched on a fresh connection with a seven-second timeout and at most five redirects. A URL that fails is a series, not a collector error: the check reports 0 and the collector stays ok. That is what makes these alertable per URL rather than as one lump.

snmp_poller​

Runs when the dashboard has assigned network devices to this host on its SNMP tab. Nothing is detected. Needs UDP port 161 from this host to each device. Reports nothing about this host itself; each device is reported as a system of its own. See SNMP Devices.

Windows hosts​

A Windows host runs the same agent, reports to the same webhook, applies the same settings and updates itself from the same rollout. What differs is what there is to see, and the difference is large enough to be worth reading before an existing alert rule is pointed at one. Installing the agent there is on Agent Installation.

Detection finds fewer services​

Only the probes that work over HTTP or from the listening table run on Windows: Patroni, etcd, PgBouncer, PostgreSQL, Garage, Traefik, nginx and Caddy. Everything else under Service collectors above is found through a unix socket, a path under /proc or /run, or a Linux command-line tool, and is therefore not detected on a Windows host at all: Docker, HAProxy, WireGuard, PHP-FPM, NFS, ZFS, fail2ban, CrowdSec and systemd.

That is the answer to "why does this Windows host report no Docker". The agent is not failing and the collector is not in error — the service simply cannot be found from where the agent stands, so nothing about it is reported. The file-based certificates collector is absent for the same reason; the certstore collector below reports the same tls_days_left series from the machine certificate store instead.

The host collectors on Windows​

Fourteen collectors always run: cpu, memory, disk, network, security, host, services, eventlog, defender, firewall, certstore, tasks, shares and bitlocker. There is nothing to detect for any of them. Four take a list from the dashboard — which services, which event logs, which tasks, which certificate stores — and run without it as well.

The first six are the ones a Linux host runs, reading Windows' own counters in place of /proc. Three of the things they report do not carry over; see Three things that do not carry over below. There is no journal collector — eventlog takes its place.

The eight that are new with Windows follow.

services​

Needs the service control manager, which every Windows host has. The services_watch setting names the services that get a series of their own; the failure count needs no configuration. Reports services_failed, the number of services set to start automatically that are not running, with up to twenty of their names in the metadata; and per service named in services_watch, service_active as 1 or 0 with its state, start type, process id and display name. Emits service.started and service.stopped for a watched service. Services already running when the agent starts are recorded, not announced. Note the names are the ones the service control manager knows, not the display names the Services console shows, and they are matched without regard to case. A watched service the host does not have still gets a series, reading not running, so a service that was uninstalled or renamed is visible instead of quietly vanishing from the dashboard. A delayed automatic start counts as automatic: a service set to start by itself is expected to be running either way.

eventlog​

Detected when wevtutil.exe is on the host. A host without it is decided once and then skipped silently — no metrics, and no error every cycle. Needs read access to the channels it reads, which for Security means LocalSystem or membership of Event Log Readers. The installed service runs as LocalSystem and has it. eventlog_channels replaces the three channels read by default — Security, System and Application. Reports eventlog_errors per channel, counting the records that channel itself called critical or an error this cycle, and eventlog_logins, eventlog_failed_logins and eventlog_privileged summed across the channels read. Emits host.user_added and host.user_removed from the account records, host.audit_log_cleared when the security audit log is cleared, host.unexpected_shutdown when the host stopped without shutting down — once, although Windows records such a stop twice, as Kernel-Power 41 and as 6008 — and host.rebooted when it booted or a shutdown was requested. The logon records also become host.login and host.login_failed events on the agent's side; the dashboard discards them on arrival, and the counts above are what remains of them. Fails with the error per channel, named by channel. A channel that cannot be read does not stop the others: what the rest yielded is still reported, and the failure is what turns the collector red. Note each channel is read forward from the record the collector last saw, so a record is seen exactly once whatever the interval; reading starts at the agent's own start, so a restarted agent does not replay the channel. At most 200 records per channel per read and at most 50 events per collection become events; the counts still cover every record read.

defender​

Detected when Get-MpComputerStatus exists. A Server Core host without the feature, or an edition that does not carry Defender, is decided once and reports nothing rather than failing every cycle. Needs nothing — the cmdlet answers without elevation. Reports defender_enabled as 1 or 0 with the running mode, whether real-time protection is on and whether tamper protection is on; defender_signature_age, the age in days of the older of the two definition sets, because the host is only as current as whichever was updated last; and defender_scan_age, the age in days of the more recent of the quick and full scans, with both in the metadata. Emits host.defender_disabled and host.defender_enabled. The state at the agent's start is recorded, not announced. Note passive mode counts as protected. Defender steps aside when another antivirus registers itself: the service keeps running and scans on demand but not in real time, which is how the two are meant to coexist. The mode says which way the host is protected. Reading passive as "off" would alarm on every host with a third-party product installed. A definition set or a scan that never happened is left out rather than reported as a number. The status is read every five minutes beside the collection, not inside it.

firewall​

Detected when Get-NetFirewallProfile exists. A host without it is decided once and reports nothing rather than failing every cycle. Needs nothing. Reports firewall_enabled as 1 or 0 per profile — domain, private, public — with the default inbound action in the metadata as block, allow or default. Emits host.firewall_disabled and host.firewall_enabled per profile. The state at the agent's start is recorded, not announced. Note the value reported is the one in force, Group Policy included. A profile nobody ever configured reports its inbound action as default, which is the common case rather than an oddity: Windows still blocks unsolicited inbound traffic there, so it is not an open profile. The profiles are read every five minutes beside the collection.

certstore​

Needs read access to the machine certificate stores. certstore_stores adds stores to the one read on every host, My; WebHosting and Root are the usual additions. Reports tls_days_left per certificate, negative once expired, with subject, issuer, expiry date, store, SANs and thumbprint. From My only the certificates the store holds a private key for are reported — the ones the host itself serves with. Intermediates and certificates other software put there have no key in the store, and their expiry is nobody's to act on. A store added through certstore_stores is reported in full. Certificates are labelled by common name, with the thumbprint appended when two share one. Fails with the store path and the error, for a store that could not be opened. The stores that could be read still report; the failure is what turns the collector red. A store that is merely empty is no error, and a certificate that cannot be parsed is skipped rather than failing the collector. Note Root is deliberately not read by default: a machine's root store carries dozens of CA certificates, expired ones among them, and they would bury the certificate the host actually serves. A certificate held in two of the configured stores is one series, not two.

tasks​

Needs PowerShell and the task scheduler. tasks_watch names the tasks that get an age of their own; the failure count needs no configuration. Reports tasks_failed, the number of tasks whose last run ended badly, with up to twenty of their names in the metadata; and per task named in tasks_watch, task_last_run_age in seconds with the run's outcome as the hexadecimal status Windows documents it by. Emits task.failed and task.recovered. Tasks already failing when the agent starts are recorded, not announced. Note a task is written the way Task Scheduler addresses it, \Folder\Task; a name without a folder means one in the root folder. Disabled tasks are left out entirely. Tasks under \Microsoft\ are left out of the failure count — Windows' own housekeeping fails and retries on its own schedule and is nobody's alarm. A result the scheduler wrote about the task rather than the program — running right now, never run yet, no runs left — is not a failure either, or a busy host would be permanently in alarm. The library is walked every five minutes beside the collection, because asking the task service about every task costs far more than a collection may spend.

shares​

Needs the SMB server (Get-SmbShare). Reports smb_shares, the number of shares this host publishes, and smb_share_sessions per share with the path it points at. Emits host.share_added and host.share_removed. A share nobody expected is worth knowing about, and so is one that went away: whatever reached it stopped working at that moment. The shares a host already had when the agent started are its configuration, not a change to it, and raise nothing. Note the administrative shares are neither counted nor reported: the drive roots, ADMIN$ and IPC$, everything whose name ends in a dollar sign. Nobody created them and they exist on every host. The shares are read every five minutes beside the collection.

bitlocker​

Detected when Get-BitLockerVolume exists. An edition without BitLocker, or a Server Core host without the feature, is decided once and reports nothing rather than failing every cycle. A host that has the cmdlet and refused the read is not that case: that is an error, it is reported, and it is tried again — the service runs with privileges a hand-run command did not have. Reports bitlocker_protected as 1 or 0 per mount point, and bitlocker_encrypted_percent per mount point for an encryption still running. Emits host.bitlocker_off when a volume that was protected no longer is. A volume that stops being listed was detached rather than unlocked and raises nothing. Note the status and method in the metadata are the raw integers the BitLocker API returns, not words. The tables that name those values could not be confirmed against a host, so they are passed through untranslated rather than guessed at. The volumes are read every fifteen minutes beside the collection — whether a volume is encrypted changes about as often as the machine is rebuilt.

Three things that do not carry over​

A rule written against a Linux fleet can be wrong on a Windows host in three ways. None of them is a fault in the agent; all three are what the platform does or does not have.

Swap is the commit charge, not the page file. swap and swap_used on Windows report how much of the commit limit is committed, which counts every reservation whether or not a page was ever written to disk. Sixty per cent is an ordinary idle Windows host and does not mean it is swapping. An alert rule that fires above, say, 50 % swap on Linux will fire permanently here.

Only drive letters are reported. disk covers the fixed drives and nothing else: removable drives, mapped network drives and optical drives are left out, because a mapped drive reported here would be this host's disk on the dashboard as well as the file server's. A volume mounted into a folder rather than given a letter is invisible to the collector, so a host that keeps its data on one will not show that space anywhere.

There is no load average, and no iowait or steal. Windows has no run queue of the kind load_1m names, and its CPU times have no iowait or steal buckets. Those series are therefore absent on a Windows host rather than reported as zero — load_1m, load_5m, load_15m, cpu_iowait and cpu_steal are simply not in the report. An alert rule that watches one of them will never fire on a Windows host and never tell you it is not watching. cpu itself, the busy percentage, is reported as everywhere.

Two smaller absences belong here as well. updates_pending comes from a Windows Update search, which takes minutes and therefore runs every six hours rather than every cycle, so the figure can be that old. And there are no oom_kills_total or oom_kills_rate: Windows trims working sets and fails allocations instead of killing a process to reclaim memory, so there is no counter to read and a zero would claim more than the agent knows.

The first cycle after a start is thin​

Six of the checks cost far more than a collection may spend and run on their own schedule beside it: Defender, the firewall, the scheduled tasks, the shares, BitLocker and the Windows Update count. The first report after a start carries nothing for them — not a zero, nothing — and their values appear from the next cycle. A host that has just been installed and shows none of them is not misconfigured; give it a cycle.

Reading the Security event log is privileged​

This is the one that catches people out while testing. The Security channel is privileged: it takes LocalSystem, or membership of Event Log Readers. The installed service runs as LocalSystem and has no trouble with it.

An agent started by hand in an ordinary console does not have either, and the eventlog collector then shows as failing with a message like

Security: exit status 5: Zugriff verweigert

— the wording follows the host's display language. The other channels are still read and still reported; it is the Security one that fails, and it turns the whole collector red.

So judge that collector from the installed service, not from a hand-run binary. If you do want to run the agent by hand against the Security channel, put the account you run it as into Event Log Readers first.

Reading collector state​

A system's Info tab has two places to look.

Detected services shows one badge per service the agent found, tinted by collector state — green for ok, red for error, amber for needs_config, plain for none. A service in disabled gets no badge; it is listed on the Services tab instead. Hovering a badge gives the version, the state in words and the error text.

Collectors lists every collector that reported in the last cycle, with Running or Failing and the error. This card is built from the collector_ok series each collector writes every cycle, so it covers the host collectors too, not only detected services. A collector silent for ten minutes while the host kept reporting drops off the list — it no longer runs, and showing its last error would be showing a collector that is not there.

The Collectors page (/collectors, in the sidebar; the Needs configuration and Collector error tiles on the overview lead there too) answers the question across every host at once. It has a tab for collectors waiting for configuration and one for collectors in error, and groups them by service and then by the reason the agent gave — hosts that fail the same way usually need the same setting, and if they form a cluster it can be set on the cluster once. Each host links to its Configuration tab.

The Agents page rolls this up across all hosts: one row per system with an error count and a "needs configuration" count, both hoverable for the service names, and a filter for With collector problems. See Agent Rollout.

On the host itself, journalctl -u selvara-agent -f logs Collector X failing: … once when a collector starts failing and Collector X recovered when it stops. Those transitions also become collector.failed and collector.recovered events, so they can notify — see Event Rules.

What "needs config" means in practice​

needs_config is not a fault. It means the agent found a service, knows exactly which collector belongs to it, and is missing one value it cannot discover: a token, a login, a socket path, a metrics URL. Nothing is collected for that service until you supply it, and nothing else is affected.

The state carries the reason, and the reason names the setting key — admin token missing (garage_admin_token), stats login missing (pgbouncer_user, pgbouncer_password), metrics endpoint not found (traefik_metrics_url). Enter it on the system's Configuration tab, or on the cluster's when it is the same value for every member, and the agent applies it within a cycle without a restart.

Full reference: Agent Settings.