Skip to main content

Event Rules

Metrics answer how much. Events answer what happened. A Patroni node was promoted at 03:14. A host rebooted. Somebody logged in over SSH. A container went into a restart loop. A collector broke. A setting was changed in the dashboard. None of those are numbers with a threshold — they are moments, and they are recorded as they occur.

Every event is stored on the timeline. An event rule decides which of those events also send you a notification.

Where events come from​

Two sources, marked as such on every entry:

  • Agent — the agent on a host notices something locally (a failover, a unit failing, a user account added, a container dying) and ships the event alongside its regular metrics report. Each event carries an id from the agent, so a report that sat in the outbox and arrives twice is still stored once.
  • Dashboard — Selvara itself notices something while processing reports or while you use it: a system that stopped reporting, a system that came back, an agent that changed version, a host whose boot time moved, a settings change you made.

Events are recorded whatever else is going on: under an open maintenance window, for an inactive customer, for a host in the middle of an incident. Only the notification is ever held back. The timeline is meant to be the truthful record of what happened, and it stays that way.

Events are kept for 90 days, after which they are pruned by the same cron call that evaluates alerts.

The timeline​

Events in the navigation shows everything, newest first, 50 at a time with Load more at the bottom. Each entry shows the kind, a severity badge, the system (and its customer and cluster where there is one), the source, the raw message and, expandable, the structured detail the reporter attached.

The filters are:

FilterWhat it does
SearchCase-insensitive substring of the event's message
KindOne event kind, from the kinds that actually occur, grouped by area
Minimum severityWarning or Critical; a minimum of warning includes criticals
RangeLast 24 hours, 7 days, 30 days, 90 days, or all time

With no range chosen the timeline looks back 30 days. Opening the timeline from a system or a cluster — from its detail page, or from the link in an event notification — adds a scope chip at the top; clear it to see everything again.

What an event rule does​

A rule matches events and, for each event any enabled rule matches, sends one notification to the usual recipients over the usual channels (Notifications). It consists of:

PartMeaning
NameFree text, for the rule list.
Event kindsTick boxes, grouped by area.
Minimum severityinfo, warning or critical. An event below it never matches.
ScopeAll systems, or one customer, one cluster or one system.

An empty kinds list means every kind. The dialog says so under the boxes, but it is worth repeating because it is set by accident: unticking the last box does not disable the rule, it widens it to everything at or above the minimum severity. If you want a narrow rule, tick the kinds you want and leave the severity where it belongs.

A rule has at most one scope. Scoping to a customer matches events on that customer's systems; scoping to a cluster matches events attributed to that cluster as well as events on systems currently in it.

Several rules may match the same event — it still notifies once.

An event is not notified while its system is silenced: an open maintenance window covering it, or an inactive customer. The event is still recorded.

Creating, editing, enabling, disabling and deleting event rules is open to administrators and managers — a manager only where the scope is one it reaches in full: a system or a customer it has been granted, or a cluster whose every member sits at a customer it has been granted. A rule with no scope notifies about every customer in the instance, so that one is administrator-only, and so is moving an existing rule onto a scope its editor could not have written in the first place.

The rule list is narrowed the same way and shows only the rules the account reaches. Operators and viewers change nothing, and the settings navigation offers this page to administrators and managers only. See Users and Permissions.

Event kinds and severities​

Severity is info, warning or critical, ordered that way. Each kind carries a default severity, and a reporter may raise or lower it for a particular occurrence.

System​

KindDefault severityMeaning
system.offlinecriticalThe host was online and has not reported for more than two minutes, and its offline alerts are on. A host with offline alerts off becomes Powered off instead and records nothing. See Events an alert already reports.
system.onlineinfoA host that was offline is reporting again.

Host​

KindDefault severityMeaning
host.rebootedwarningThe host's boot time changed between two reports.
host.user_addedwarningA user account appeared.
host.user_removedwarningA user account disappeared.
host.group_addedwarningA local group appeared.
host.group_removedwarningA local group disappeared.
host.group_member_addedwarningAn account joined a group that already existed — somebody added to sudo or Administrators.
host.group_member_removedwarningAn account left a group that still exists.

Logins, failed logins and sudo commands are not events. They happen all day on every host, and failed logins on a machine facing the internet are bots that fail2ban already turns away; one entry each buried everything else in the timeline. The agents still count them, as metrics — see Agent Collectors.

Database​

KindDefault severityMeaning
patroni.promotedcriticalA node became the leader — a failover happened.
patroni.demotedwarningA node became a replica.
patroni.timeline_changedwarningThe Patroni timeline moved.

Container​

KindDefault severityMeaning
container.restart_loopwarningA container is restarting repeatedly.
container.diedwarningA container stopped.

Service​

KindDefault severityMeaning
systemd.unit_failedwarningA watched systemd unit entered the failed state.
systemd.unit_recoveredinfoIt came back.

Virtualization​

KindDefault severityMeaning
proxmox.backup_failedwarningThe backup of a guest on a Proxmox VE node failed; the message names the guest and carries vzdump's error.
proxmox.task_failedwarningAnother task on a Proxmox VE node failed — a migration, a start, a backup job that failed before reaching a guest; the message carries the task and its error. One that ended with warnings is not reported.

Agent​

KindDefault severityMeaning
agent.startedinfoThe agent process started.
agent.updatedinfoThe agent's reported version changed; the message carries old → new.
agent.server_switchedinfoThe agent moved to this server; recorded on the new server. See Moving to a new server.
agent.server_switch_failedwarningThe agent was told to move but the new server did not accept it; recorded on the old server, with the reason.
collector.failedwarningA collector stopped delivering and reported an error.
collector.recoveredinfoIt is delivering again.
service.appearedinfoThe agent detected a service on the host it had not seen before.
service.disappearedwarningA service it had been watching is gone.

Settings​

KindDefault severityMeaning
settings.changedinfoAn agent setting was changed in the dashboard, for a system or for a cluster — or, for the agents' server address, for the whole instance.

Kinds are validated when a rule is saved: a rule can only name kinds from this catalogue.

Events an alert already reports​

Two kinds have an alert rule counterpart: system.offline and the heartbeat metric, raid.degraded and raid_healthy. While an enabled below 1 rule on that metric watches the system, the event is recorded but not mailed, because the alert mails when it opens and again when it resolves. On a system no such rule watches, the event notifies as any other.

What a fresh instance already notifies on​

On first migration, if no event rules exist, Selvara seeds two:

NameKindsMinimum severityScopeEnabled
Critical events(empty — every kind)criticalall systemsyes
Security and failuressee belowwarningall systemsyes

Critical events covers a failover (patroni.promoted) and a cleared security log, and a degraded RAID array or a host going offline where no alert rule reports it already. Security and failures names the warnings somebody has to act on:

  • an account joining an admin group: host.group_member_added — mailed only for root, sudo, wheel, admin, adm, docker, lxd, Administrators, Remote Desktop Users, Domain Admins and Enterprise Admins, in English or German. Joining any other group is recorded only.
  • protection switched off: host.firewall_disabled, host.defender_disabled
  • a host that crashed: host.unexpected_shutdown — except on a host whose offline alerts are off, a Windows desktop by default, which is switched off every day. See Customers and Systems.
  • a backup that failed: proxmox.backup_failed
  • a container in a crash loop: container.restart_loop

Everything else is recorded on the timeline without a mail, because it happens in normal operation:

Left outWhy
host.user_added, host.user_removed, host.group_added, host.group_removed, host.group_member_removedEveryday administration.
host.rebootedPlanned reboots count, and Windows records one up to three times. A crash arrives as host.unexpected_shutdown, a host that stays down as the System offline alert.
host.share_addedCreating a share is administration, not a fault.
host.bitlocker_offWindows Update suspends BitLocker for firmware and cumulative updates.
proxmox.task_failedAny failed task of any user, a closed console included.
task.failedAny Windows scheduled task outside \Microsoft\, including OneDrive's per-user updaters, which fail routinely.
systemd.unit_failed, interface.down, collector.failed, container.died, service.stoppedCovered by an alert rule, or too frequent to mean anything by themselves.

The seed happens once. Deleting a rule does not bring it back, and an instance that already had event rules is left untouched. If no rule exists at all, no event ever notifies — the rule list says as much when it is empty.

Instances upgraded to 2.31.0 received whichever of the two rules they did not have under that name; the upgrade to 2.34.0 set Security and failures to the kinds above.

Choosing rules that are worth having​

A few shapes that work:

  • Keep Critical events as the floor and add narrower rules on top.
  • systemd.unit_failed and collector.failed at warning, scoped to the customers whose hosts you actually operate, turn "something quietly stopped" into a message.
  • container.restart_loop at warning catches the failure mode that never crosses a metric threshold because the container keeps coming back.

Related: Alert Rules for thresholds on measurements, Notifications for where these messages go, and Maintenance Windows for switching them off on purpose.