Event Rules
Metrics answer how much. Events answer what happened. A Patroni node was promoted at 03:14. A host rebooted. Somebody logged in over SSH. A container went into a restart loop. A collector broke. A setting was changed in the dashboard. None of those are numbers with a threshold — they are moments, and they are recorded as they occur.
Every event is stored on the timeline. An event rule decides which of those events also send you a notification.
Where events come from
Two sources, marked as such on every entry:
- Agent — the agent on a host notices something locally (a failover, a unit failing, a user account added, a container dying) and ships the event alongside its regular metrics report. Each event carries an id from the agent, so a report that sat in the outbox and arrives twice is still stored once.
- Dashboard — Selvara itself notices something while processing reports or while you use it: a system that stopped reporting, a system that came back, an agent that changed version, a host whose boot time moved, a settings change you made.
Events are recorded whatever else is going on: under an open maintenance window, for an inactive customer, for a host in the middle of an incident. Only the notification is ever held back. The timeline is meant to be the truthful record of what happened, and it stays that way.
Events are kept for 90 days, after which they are pruned by the same cron call that evaluates alerts.
The timeline
Events in the navigation shows everything, newest first, 50 at a time with Load more at the bottom. Each entry shows the kind, a severity badge, the system (and its customer and cluster where there is one), the source, the raw message and, expandable, the structured detail the reporter attached.
The filters are:
| Filter | What it does |
|---|---|
| Search | Case-insensitive substring of the event's message |
| Kind | One event kind, from the kinds that actually occur, grouped by area |
| Minimum severity | Warning or Critical; a minimum of warning includes criticals |
| Range | Last 24 hours, 7 days, 30 days, 90 days, or all time |
With no range chosen the timeline looks back 30 days. Opening the timeline from a system or a cluster — from its detail page, or from the link in an event notification — adds a scope chip at the top; clear it to see everything again.
What an event rule does
A rule matches events and, for each event any enabled rule matches, sends one notification to the usual recipients over the usual channels (Notifications). It consists of:
| Part | Meaning |
|---|---|
| Name | Free text, for the rule list. |
| Event kinds | Tick boxes, grouped by area. |
| Minimum severity | info, warning or critical. An event below it never matches. |
| Scope | All systems, or one customer, one cluster or one system. |
An empty kinds list means every kind. The dialog says so under the boxes, but it is worth repeating because it is set by accident: unticking the last box does not disable the rule, it widens it to everything at or above the minimum severity. If you want a narrow rule, tick the kinds you want and leave the severity where it belongs.
A rule has at most one scope. Scoping to a customer matches events on that customer's systems; scoping to a cluster matches events attributed to that cluster as well as events on systems currently in it.
Several rules may match the same event — it still notifies once.
An event is not notified while its system is silenced: an open maintenance window covering it, or an inactive customer. The event is still recorded.
Creating, editing, enabling, disabling and deleting event rules is open to administrators and managers — a manager only where the scope is one it reaches in full: a system or a customer it has been granted, or a cluster whose every member sits at a customer it has been granted. A rule with no scope notifies about every customer in the instance, so that one is administrator-only, and so is moving an existing rule onto a scope its editor could not have written in the first place.
The rule list is narrowed the same way and shows only the rules the account reaches. Operators and viewers change nothing, and the settings navigation offers this page to administrators and managers only. See Users and Permissions.
Event kinds and severities
Severity is info, warning or critical, ordered that way. Each kind carries
a default severity, and a reporter may raise or lower it for a particular
occurrence.
System
| Kind | Default severity | Meaning |
|---|---|---|
system.offline | critical | The host was online and has not reported for more than two minutes, and its offline alerts are on. A host with offline alerts off becomes Powered off instead and records nothing. See Events an alert already reports. |
system.online | info | A host that was offline is reporting again. |
Host
| Kind | Default severity | Meaning |
|---|---|---|
host.rebooted | warning | The host's boot time changed between two reports. |
host.user_added | warning | A user account appeared. |
host.user_removed | warning | A user account disappeared. |
host.group_added | warning | A local group appeared. |
host.group_removed | warning | A local group disappeared. |
host.group_member_added | warning | An account joined a group that already existed — somebody added to sudo or Administrators. |
host.group_member_removed | warning | An account left a group that still exists. |
Logins, failed logins and sudo commands are not events. They happen all day on every host, and failed logins on a machine facing the internet are bots that fail2ban already turns away; one entry each buried everything else in the timeline. The agents still count them, as metrics — see Agent Collectors.
Database
| Kind | Default severity | Meaning |
|---|---|---|
patroni.promoted | critical | A node became the leader — a failover happened. |
patroni.demoted | warning | A node became a replica. |
patroni.timeline_changed | warning | The Patroni timeline moved. |
Container
| Kind | Default severity | Meaning |
|---|---|---|
container.restart_loop | warning | A container is restarting repeatedly. |
container.died | warning | A container stopped. |
Service
| Kind | Default severity | Meaning |
|---|---|---|
systemd.unit_failed | warning | A watched systemd unit entered the failed state. |
systemd.unit_recovered | info | It came back. |
Virtualization
| Kind | Default severity | Meaning |
|---|---|---|
proxmox.backup_failed | warning | The backup of a guest on a Proxmox VE node failed; the message names the guest and carries vzdump's error. |
proxmox.task_failed | warning | Another task on a Proxmox VE node failed — a migration, a start, a backup job that failed before reaching a guest; the message carries the task and its error. One that ended with warnings is not reported. |
Agent
| Kind | Default severity | Meaning |
|---|---|---|
agent.started | info | The agent process started. |
agent.updated | info | The agent's reported version changed; the message carries old → new. |
agent.server_switched | info | The agent moved to this server; recorded on the new server. See Moving to a new server. |
agent.server_switch_failed | warning | The agent was told to move but the new server did not accept it; recorded on the old server, with the reason. |
collector.failed | warning | A collector stopped delivering and reported an error. |
collector.recovered | info | It is delivering again. |
service.appeared | info | The agent detected a service on the host it had not seen before. |
service.disappeared | warning | A service it had been watching is gone. |
Settings
| Kind | Default severity | Meaning |
|---|---|---|
settings.changed | info | An agent setting was changed in the dashboard, for a system or for a cluster — or, for the agents' server address, for the whole instance. |
Kinds are validated when a rule is saved: a rule can only name kinds from this catalogue.
Events an alert already reports
Two kinds have an alert rule counterpart: system.offline and the heartbeat
metric, raid.degraded and raid_healthy. While an enabled below 1 rule on
that metric watches the system, the event is recorded but not mailed, because
the alert mails when it opens and again when it resolves. On a system no such
rule watches, the event notifies as any other.
What a fresh instance already notifies on
On first migration, if no event rules exist, Selvara seeds two:
| Name | Kinds | Minimum severity | Scope | Enabled |
|---|---|---|---|---|
| Critical events | (empty — every kind) | critical | all systems | yes |
| Security and failures | see below | warning | all systems | yes |
Critical events covers a failover (patroni.promoted) and a cleared
security log, and a degraded RAID array or a host going offline where no alert
rule reports it already. Security and failures names the warnings somebody has
to act on:
- an account joining an admin group:
host.group_member_added— mailed only forroot,sudo,wheel,admin,adm,docker,lxd, Administrators, Remote Desktop Users, Domain Admins and Enterprise Admins, in English or German. Joining any other group is recorded only. - protection switched off:
host.firewall_disabled,host.defender_disabled - a host that crashed:
host.unexpected_shutdown— except on a host whose offline alerts are off, a Windows desktop by default, which is switched off every day. See Customers and Systems. - a backup that failed:
proxmox.backup_failed - a container in a crash loop:
container.restart_loop
Everything else is recorded on the timeline without a mail, because it happens in normal operation:
| Left out | Why |
|---|---|
host.user_added, host.user_removed, host.group_added, host.group_removed, host.group_member_removed | Everyday administration. |
host.rebooted | Planned reboots count, and Windows records one up to three times. A crash arrives as host.unexpected_shutdown, a host that stays down as the System offline alert. |
host.share_added | Creating a share is administration, not a fault. |
host.bitlocker_off | Windows Update suspends BitLocker for firmware and cumulative updates. |
proxmox.task_failed | Any failed task of any user, a closed console included. |
task.failed | Any Windows scheduled task outside \Microsoft\, including OneDrive's per-user updaters, which fail routinely. |
systemd.unit_failed, interface.down, collector.failed, container.died, service.stopped | Covered by an alert rule, or too frequent to mean anything by themselves. |
The seed happens once. Deleting a rule does not bring it back, and an instance that already had event rules is left untouched. If no rule exists at all, no event ever notifies — the rule list says as much when it is empty.
Instances upgraded to 2.31.0 received whichever of the two rules they did not have under that name; the upgrade to 2.34.0 set Security and failures to the kinds above.
Choosing rules that are worth having
A few shapes that work:
- Keep Critical events as the floor and add narrower rules on top.
systemd.unit_failedandcollector.failedat warning, scoped to the customers whose hosts you actually operate, turn "something quietly stopped" into a message.container.restart_loopat warning catches the failure mode that never crosses a metric threshold because the container keeps coming back.
Related: Alert Rules for thresholds on measurements, Notifications for where these messages go, and Maintenance Windows for switching them off on purpose.