Skip to main content

Alert Rules

An alert rule is a standing question about a measured value: is this metric on the wrong side of this number, and has it stayed there long enough to be real? When the answer turns yes, an alert is raised and notifications go out. When it turns no again, the alert resolves by itself.

Rules live under Settings → Alert rules, which the settings navigation offers to administrators and managers. A manager writes rules for the customers and clusters it reaches in full, and sees the rules that apply to every system read-only, behind an eye icon — they raise alerts on the manager's systems too. See Users and Permissions for the details.

What a rule says​

PartMeaning
NameFree text. It becomes the first half of the alert message, so write it as the sentence you want to read in the mail: Disk above 85%.
MetricWhich measurement to watch. The dropdown is grouped by area and offers only what at least one system currently reports, plus the metric of the rule you are editing.
ConditionAbove or Below.
ThresholdThe number to compare against.
Duration (seconds)How long the condition must hold before an alert is raised.

A rule with no scope applies to every system. That is what most default rules do, and for things like disks and reachability it is what you want.

Narrowing a rule​

A rule can be limited in three ways from the dialog, and they combine — a rule with a customer and a service applies to the systems of that customer on which the agent found that service:

  • Customer — only that customer's systems.
  • Cluster — only systems currently in that cluster. Because membership is reported by the agents, a node added to the Patroni cluster tomorrow is covered without anyone touching the rule.
  • Service — only systems where the agent detected that service. This is detection, not declaration: a rule about patroni_lag scoped to Patroni never looks at a host that does not run it.

The engine also supports pinning a rule to a single system, and a rule stored that way is evaluated against that system alone. The rule dialog does not offer that choice; in practice you narrow to a customer, a cluster or a service.

The list shows the resulting scope in one column, or All systems when a rule is unnarrowed.

Metrics, series and what disk really means​

A metric type names a kind of measurement: cpu, memory, docker_containers_unhealthy. A series is one instance of that measurement on one host, and where a host has several, the instance is written after a colon:

disk:/ the root filesystem
disk:/mnt/garage a second mount on the same host
network_in_rate:eth0 one interface
f2b_failed:sshd one fail2ban jail
tls_days_left:example.com one certificate

A rule is always written against the base type — you select disk, never disk:/mnt/garage. The engine then evaluates it once per series. On a host with three mounts, one rule Disk above 85 % watches all three independently and can raise three separate alerts, each naming its own mount. The same holds for interfaces, jails, containers, HAProxy backends and certificates.

That is also why alerts carry the full series in their title: Disk (/mnt/garage) Above 85. There is no way to write a rule for one specific mount from the rule dialog; scope it to the system instead, or accept that it covers every mount on the matching hosts.

The metric areas offered are: host, disks, network, Docker, Swarm, database, proxy and web, storage, virtualization, security, checks, and agent. Which types exist under each depends on what the agents collect on their hosts — see Agent Installation.

How an alert is raised, acknowledged and resolved​

Raised. During an evaluation run, the average of the series over the rule's duration is compared against the threshold. If it is on the wrong side and no alert for that system and series is open yet, an alert is created with status ACTIVE, and notifications go out over every channel the recipients have enabled.

Levels. Rules on the same metric and in the same direction are levels of one alert, not separate alerts. Each still judges the series over its own duration, and the alert stands for the strictest level crossed: a disk at 96 % has one alert, from Disk above 95%, not a second one from Disk above 85%. When a stricter level is crossed while the alert is open, the alert escalates — it moves to that rule, turns active again even if it was acknowledged, and is mailed like a new one. It never steps back down: a disk falling from 96 % to 90 % keeps its above 95% alert, with the current value showing 90, so a value wandering around a threshold is not mailed at every crossing. It resolves once no level is crossed any more. The same holds for a database node whose disk passes Patroni node disk above 75% and then Disk above 85%. An above and a below rule on the same metric are two alerts.

Nothing is raised for a system that is currently silenced — covered by an open maintenance window, or belonging to an inactive customer. Alerts that were already open are still updated and still resolve while the silence lasts.

Acknowledged. Acknowledge on the alerts page says "seen, being handled". It stops nothing and changes no threshold; it moves the alert out of the active list and records who acknowledged it and when. An acknowledged alert stays open: it resolves by itself like an active one, and while it is open the same series raises no second alert. Administrators and operators may do this, viewers may not — the API refuses it, not just the button.

Reopened. A series that goes back over the threshold within 30 minutes of its alert resolving does not raise a new alert: the old one opens again, in the state it had — an acknowledged alert stays acknowledged — and nobody is mailed, unless it is back at a stricter level than it had, which is an escalation. Memory hovering at 90 % dips under the threshold for one window and is back the next; without this, every dip meant a resolved alert, a fresh one, a new mail and a lost acknowledgement. The same applies to an alert resolved by hand while its value stayed where it was. After 30 minutes a relapse is a new incident, with a new alert and a new mail.

Resolved. An alert resolves in four ways:

  1. The window average comes back to the right side of the threshold. The duration is the damping here: a value hovering at the threshold has to stay on one side for a full window to flip the alert. This is the only way that mails: once the alert has stayed resolved for 30 minutes, the recipients are told it is over — see Notifications.
  2. Somebody resolves it by hand on the alerts page (again: not viewers).
  3. The series stopped existing. If the host itself is still reporting but the series behind an open alert has gone quiet for 15 minutes — the disk was unmounted, the collector no longer runs — the alert is resolved rather than left open forever. An offline host is left alone, because its reachability alert is already saying what is going on.
  4. The rule no longer watches it: it was deleted, disabled, or narrowed so that the system is outside its scope. The alert is resolved on the next evaluation run.

Deleted. Administrators and managers can delete an alert outright with the bin icon on its card, in any state. Resolved alerts are also deleted automatically once they have been resolved for longer than the retention under Settings → Data retention — 90 days unless an administrator sets otherwise, and never with 0. Open alerts are never deleted automatically. Deleting a system takes its alerts along.

Every tab of the alerts page lists all its alerts, newest first: the first fifty come with the page and older ones load as you scroll to the end.

Opening a single alert. /alerts?alertId=<id> points at one alert: the tab for its state is selected, the card is outlined and scrolled into view, and the × on the chip beside the page title clears the highlight. The alert is read by id when it is not on the first page of its tab, so a resolved or long-past one still shows. This is the link an alert mail carries.

Series that nobody has reported for seven days are forgotten entirely, so a removed disk or a renamed interface stops being judged. Collector-state series are dropped after an hour of silence from a host that is otherwise reporting.

A certificate that expired more than 30 days ago is not judged by any rule on tls_days_left or traefik_cert_days_left, and an alert open on it is resolved without a mail. Such a certificate was replaced or given up — old entries in a Windows certificate store are the usual case — and would otherwise keep an alert open for as long as the entry exists. Until then an expired certificate alerts like one about to expire.

Reachability itself is a rule like any other: the heartbeat metric is 1 while reports arrive, and a Below 1 rule on it fires when they stop. Separately, a system that was online and has not reported for two minutes is marked OFFLINE and a system.offline event is recorded — that is the event path, not the alert path. A system whose offline alerts are off is marked STANDBY instead, with no event; see the next section.

Hosts that are allowed to go quiet​

A workstation is switched off at the end of the day. Left alone it would raise an alert every evening and resolve it every morning, which is how people learn to ignore the one evening it mattered.

Each system therefore carries its own answer to alert when this stops reporting, on its Configuration tab:

SettingWhat happens
AutomaticThe platform decides. A host reporting a Windows desktop edition (not Windows Server) stays quiet; everything else alerts, including a host that has not reported its platform yet.
AlwaysAlerts, whatever the platform — a workstation that is supposed to stay on.
NeverStays quiet, whatever the platform.

It covers both paths: no system.offline event, and no missing heartbeat fed to a Below 1 rule. Its restarts and shutdowns — host.rebooted, host.unexpected_shutdown — are recorded on the timeline but notify nobody, whatever the event rules say. The rule itself is untouched and keeps watching every other host, so there is no need for a second rule or a rule per customer.

It also decides what the host is shown as. A quiet host whose offline alerts are on is OFFLINE; one whose offline alerts are off is STANDBY, shown as Powered off with a grey dot. A powered-off host is not counted by the Offline attention tile and is left out of the overview's share of systems online, and when it reports again it turns ONLINE without a system.online event. Changing the setting on a host that is already quiet moves it between the two on the next evaluation run. The last report stays visible either way. See Customers and Systems.

A system that has never reported raises nothing either, whatever its setting. This is worth stating because the missing heartbeat is synthesised rather than measured: there is no series to average, so the rule's duration cannot hold it back, and a system created in the dashboard used to raise the alarm within seconds of being added. Creating the system days before anyone installs the agent on it is quiet.

Two behaviours that surprise people​

1. Alerts are evaluated when something calls the evaluator​

There is no timer inside the application. Evaluation happens when

POST /api/cron/evaluate-alerts

is called, and that call is what also marks stale systems offline, places proxies and storage hosts with their cluster, and — at most once an hour — prunes forgotten series and expired events. If the endpoint is never called, no alert is ever raised, however wrong the numbers look on screen. If it is called every 30 seconds, alerts appear within about 30 seconds of a rule's duration elapsing.

Set that schedule up once when you install; see Installation and Configuration. If a CRON_SECRET is configured, the call must carry it as a bearer token.

A consequence worth knowing: the duration is not a countdown. The engine does not remember that a value has been high for four of your five minutes. Each run averages the last duration seconds and judges that average. A rule with a duration of one hour is answered one hour after the first report, and then on every run after it.

2. The current value and the judged value are different numbers​

The number shown on the alert, the stat cards and the badges is the live value from the current-value table — one row per system and series, the last thing the agent reported.

The number the rule is judged on is the average over the duration, computed from the history table.

They rarely match exactly, and they are not meant to. Reading a chart and expecting it to cross the threshold at the moment the alert appears will mislead you twice over: the chart's longer ranges are drawn from pre-aggregated rollups whose buckets are themselves averages, so a spike that lasted 40 seconds inside a five-minute bucket is drawn much smaller than it was — while the rule, averaging over its own window, may well have seen it. Short spikes that did not last the duration are exactly what the duration is there to swallow.

If you want "the value right now", read the stat card. If you want "was this bad for long enough", trust the alert.

What a fresh instance already watches​

On first migration, if no alert rules exist yet, Selvara seeds the following rules, all enabled. They are seeded once — deleting one does not bring it back, and an instance that already had rules is left alone.

These apply to every system:

RuleMetricConditionDuration
System offlineheartbeatbelow 1300 s
CPU above 90%cpuabove 90600 s
Memory above 90%memoryabove 90600 s
Disk above 85%diskabove 85300 s
Disk above 95%diskabove 9560 s
Inodes above 90%disk_inodesabove 90300 s
Container unhealthydocker_containers_unhealthyabove 0120 s
Swarm service below desired replicasdocker_service_replicas_missingabove 0120 s
Swarm node downdocker_swarm_nodes_downabove 0120 s
Watched unit not activeunit_activebelow 160 s
HAProxy backend downhaproxy_backend_upbelow 160 s
etcd without leaderetcd_has_leaderbelow 160 s
etcd database above 80% of quotaetcd_db_used_percentabove 80300 s
Patroni cluster unlockedpatroni_cluster_unlockedabove 060 s
Patroni node downpatroni_runningbelow 160 s
Replication lag above 50 MBpatroni_lagabove 52428800300 s
PgBouncer client waiting longer than 10 spgbouncer_maxwaitabove 10300 s
PostgreSQL connections above 90%pg_connections_percentabove 90120 s
PostgreSQL transaction ID age above 50%pg_xid_age_percentabove 503600 s
Garage cluster not healthygarage_healthybelow 1120 s
Garage node downgarage_node_upbelow 1120 s
Garage node above 85%garage_node_data_used_percentabove 85300 s
ZFS pool not onlinezfs_pool_healthybelow 1120 s
ZFS pool above 85%zfs_pool_capabove 85300 s
Certificate expires within 14 daystls_days_leftbelow 143600 s
Traefik certificate expires within 14 daystraefik_cert_days_leftbelow 143600 s
Traefik backend downtraefik_service_server_upbelow 1120 s
Traefik configuration reload failedtraefik_config_reload_okbelow 1120 s
URL check failinghttp_okbelow 1120 s
NFS mount unreachablenfs_mount_okbelow 1120 s
Clock off by more than 3 minutesclock_skewabove 180900 s
RAID degradedraid_healthybelow 160 s
Proxmox cluster without quorumpve_cluster_quoratebelow 160 s
Proxmox HA resource in errorpve_ha_errorsabove 0120 s
Proxmox guest with autostart stoppedpve_guests_onboot_stoppedabove 0300 s
Proxmox backup older than 5 dayspve_backup_oldest_ageabove 432000900 s
Proxmox storage above 85%pve_storage_used_percentabove 85300 s
Proxmox storage unavailablepve_storage_activebelow 1300 s
LVM thin pool above 90%pve_thinpool_data_percentabove 90300 s
LVM thin pool metadata above 80%pve_thinpool_metadata_percentabove 80300 s
Proxmox replication failingpve_replication_okbelow 1300 s

One rule is narrowed to hosts on which the agent detected the service, and warns sooner than the rules above, because a database node has less room to spare:

RuleServiceMetricConditionDuration
Patroni node disk above 75%Patronidiskabove 75300 s

It follows Patroni rather than PostgreSQL: a Synology NAS or a mail server runs an embedded PostgreSQL of its own and is no database node.

The set is meant to mail only about faults somebody has to fix. These measurements have no default rule on purpose, and are still shown on the system page:

MetricWhy not
reboot_requiredA pending reboot is not an incident.
systemd_failed_unitsStays above 0 for months on a host with an old failed oneshot. A unit that has to run is watched by name, through Watched unit not active.
haproxy_server_upIn front of Patroni only the leader is up in the write backend and only the replicas in the read one. HAProxy backend down alerts once a backend has no server left.
collector_okA collector timing out or lacking a permission is the agent's trouble, not the host's. See the Collectors page.
docker_containers_restartingA snapshot every 30 seconds that catches a restart or not by chance. A crash loop arrives as the container.restart_loop event instead; see Event Rules.
interface_upAn unplugged switch port is not a fault.
pgbouncer_clients_waiting_totalA client queueing for a moment is how a pool works; PgBouncer client waiting longer than 10 s says when one waits too long.
cpu, memory per serviceHigh load on a database, storage or Docker host is normal. The general rules at 90 % apply to them.

On a database node, the service rule and the general disk rules are levels of one alert; see How an alert is raised.

The upgrade to 2.31.0 replaced every existing rule with this set once. The rules an earlier seed script had recreated on every start — [CRITICAL], [Docker Manager], [SECURITY] and the like — had applied to every host since rules lost their system type, and several watched metrics no agent reports. Changes to accounts and groups are events, not metrics; see Event Rules.

The backup rule allows five days, which covers nodes backed up every second or third day. Where some customers back up nightly and should be warned sooner, add a rule with a shorter threshold narrowed to them (see Narrowing a rule); a single guest that is backed up some other way goes into proxmox_guests_exclude — see Agent Settings.

A rule about Patroni or Garage costs nothing on hosts that report no such metric, scoped or not: with no series of that type, there is nothing to judge. Disable the ones you do not want rather than deleting them if you might want them back.

The disk and certificate rules also drive two of the overview's attention tiles. Disk nearly full counts the systems with an active or acknowledged alert on a disk series, Certificate expiring those with one on tls_days_left or traefik_cert_days_left; neither tile has a threshold of its own. To be warned about disks only at 95 %, disable Disk above 85% — the tile follows. A database node is counted from 75 %, through Patroni node disk above 75%. See Customers and Systems.

Who may do what​

ActionADMINMANAGEROPERATORVIEWER
See alertsevery customergranted customersgranted customersgranted customers
See a ruleevery rulerules whose scope it reachesrules whose scope it reachesrules whose scope it reaches
Create, edit, enable, disable, delete a ruleyesevery scope it names granted; never unscopednono
Acknowledge or resolve an alertyesgranted customersgranted customersno
Delete an alertyesgranted customersnono

Which customers an account has been granted is set per account; see Users and Permissions. An alert at a customer the account cannot reach answers as though it did not exist.

Related: Event Rules for things that happen rather than things that are measured, Maintenance Windows for planned silence, and Notifications for where an alert ends up.