Customers and Systems
Everything Selvara monitors hangs off two things: a customer, which groups servers and owns the scope of most settings, and a system, which is one monitored server running the agent.
The model
- A customer is the group. It has a name, a short code, optional contact details and notes, and an active switch.
- A system is one server. It belongs to exactly one customer.
- A system additionally carries a region and a list of tags of its own. These are free text on the server, not a grouping with settings attached — they exist so you can find hosts again.
- A cluster is not created by hand. Agents report that they belong to a Patroni scope, a Docker Swarm, a Garage cluster or a Proxmox VE cluster, and the cluster appears by itself.
Alert rules, event rules and maintenance windows can all be scoped to a customer, to a cluster or to a single system. None of them can be scoped to a region or a tag.
Customers
Creating one
The customers page lists one compact card per customer — the customers this account may reach, which for an administrator is all of them and for everybody else is what they have been granted (see Users and Permissions). A card carries the name, the number of active alerts, the number of systems and how many of them are online, offline and powered off. The search box above the cards matches the name and the customer code, although the code is not shown on the card. Add customer opens a dialog with:
| Field | Meaning |
|---|---|
| Name | What people read: the company or project name. |
| Customer code | A short identifier, at most 20 characters. The form lower-cases it and replaces spaces with hyphens as you type. It must be unique; a code already in use is rejected. |
| Contact | Optional name of the person to call. |
| Contact email | Optional. It is stored with the customer; it is not a notification recipient (see Notifications). |
| Notes | Free internal text, up to 2000 characters. |
| Active | On by default. See below. |
The code is the short label that travels outside the dashboard. It appears
next to the customer name in the header of the customer's page, in the
customer picker of the system dialog, in the overview's recent-alerts list,
and — the reason to keep it short and recognisable — in every alert email as
the Customer line, and in event mails as part of the facts. Pick something
you will recognise at 3 a.m.
Administrators and managers create customers; a manager is granted the customer they create, and edits and deletes only customers granted to them.
The active switch
Switching a customer to inactive is the permanent version of a maintenance window. While it is off:
- every system, metric, alert and event it already has is kept,
- its agents keep reporting and their data keeps being stored, so charts and the timeline stay complete,
- no new alert is raised for any of its systems and no notification is sent,
- alerts that were already open keep being updated and still resolve normally, so the quiet ends with a clean slate rather than with alerts frozen mid-incident,
- the customer disappears from the customers page unless Show inactive is ticked.
This is the same silence a maintenance window produces, and it is answered in the same place in the code — see Maintenance Windows.
Deleting one
A customer can only be deleted once it has no systems left. The dialog says how many are still attached and the request is refused by the API as well. Delete or move the servers first.
The customer detail page
Opening a customer gives you four stat cards (total, online, offline, powered off) and these tabs:
- Systems — the servers of this customer, in the same list as the systems page (see Finding things in the systems list) with the same filters and actions, minus the customer column and the customer filter. Add system there preselects this customer.
- Metrics — CPU, memory and disk for every system of the customer drawn on one chart each, over a selectable range. Hosts can be hidden from the legend, and hiding one hides it in all three charts, which is how you set a known noisy box aside.
- Alerts — the customer's currently active alerts.
- Users — administrators only: who may reach this customer, as a checkbox per account. Administrators are not listed, because they reach every customer already. See Users and Permissions.
The alert count in the tab header updates live.
Systems
Adding a server
Add system asks for:
| Field | Meaning |
|---|---|
| Name | The display name in every list. |
| Customer | Required. A system cannot exist without one. |
| Region | Free text, optional. |
| Tags | Free text labels, optional. |
There is no hostname to type. The agent reports the host's own name with every report, and that is what the dashboard shows under the name and in the details. Until the agent has reported once, the details say so instead of showing a hostname.
On save the dashboard shows the agent key once together with a ready-made install command of the form:
curl -sSL https://selvara.example.com/api/agent/install.sh | sudo bash -s -- \
--api-key <the key> \
--endpoint https://selvara.example.com
The key is stored hashed and is never shown again. Copy it now; if you lose it, use the key icon in the row to regenerate it — which invalidates the old key, so the agent on that host has to be reconfigured. The full walk-through is in Agent Installation.
Until the agent reports for the first time, the system sits in the list with no last-seen time. From the first report on, the agent fills in what it found by itself: the detected services, the host facts, and the cluster it belongs to.
Adding, editing, deleting a system and regenerating its key are open to administrators, and to managers for systems at a customer granted to them.
Network devices
A switch, router or firewall is added on the SNMP tab of the host that polls it, not with Add system: see SNMP Devices. In the list it carries an SNMP badge and a second line naming its poller, and the kind filter shows servers or network devices alone. It has no agent key, and its customer is always its poller's.
Region and tags
Both live on the server, not on the customer, and both are free text:
- Region is one value — where the box stands. Only the surrounding whitespace is trimmed; the capitalisation you type is kept, because it is a name people read.
- Tags are a list. Each is trimmed and folded to lower case, blanks and
repeats are dropped.
Prod,prodandPRODare therefore one tag, which is what stops the same label from becoming two entries in the filter.
Both fields offer the values already in use as suggestions while you type. The suggestions are asked of the server rather than read off the rows the browser happens to be showing, so a region used only by a host further down the list is still offered — and does not quietly acquire a second spelling.
Whether it alerts when it goes quiet
A system's Configuration tab decides whether this host raises the alarm
when it stops reporting. The default, automatic, lets the platform answer:
a host reporting a Windows desktop edition (not Windows Server) stays quiet,
because a workstation is switched off at the end of the day; everything else
alerts. Always and Never settle it for this one host. A host that stays
quiet is also not mailed about for restarting or shutting down: its
host.rebooted and host.unexpected_shutdown events stay on the timeline.
The same setting decides what the host is shown as once it has not reported for two minutes:
| Offline alerts | Status | Shown as |
|---|---|---|
| on | OFFLINE | Offline, red |
| off | STANDBY | Powered off, grey |
A powered-off host raises no event and no alert, is not counted by the
Offline attention tile, and is left out of the overview's share of systems
online — that share is online systems divided by all systems minus the
powered-off ones, so a row of workstations switched off for the night does
not drag it down. When it reports again it turns ONLINE without a
system.online event, because it never went offline as far as anyone was
told.
Changing the setting on a host that is already quiet moves it between
OFFLINE and STANDBY on the next health check rather than at its next
outage. See Alert Rules for what exactly is held back.
Finding things in the systems list
The table has one row per system:
| Column | Shows |
|---|---|
| Status | a dot coloured by status — green online, red offline, grey powered off |
| Name | the name, with the cluster and the host's role in it on a second line, a Maintenance badge while an open maintenance window covers it, and a warning icon when its agent version differs from the one this dashboard ships |
| Customer | the customer's name, linking to its page |
| Alerts | the number of active alerts |
| Last seen | how long ago the agent last reported |
| Actions | see below |
The actions at the end of a row:
- Show details (the eye icon), for everybody who can see the system, opens a dialog with the hostname, the operating system and platform, the cluster, the region, the tags, the detected services, the agent version with a hint when it is outdated, and the last report both relative and as a date and time. A button in the dialog opens the system's page.
- Manage API key (the key icon), Edit and Delete are there for administrators, and for managers on systems at a customer granted to them.
The filter bar above the table combines a free-text search over name and hostname with these dropdowns:
| Filter | Matches |
|---|---|
| Status | Online, Offline, Powered off (STANDBY), Degraded, Unknown |
| Service | Systems where the agent detected that service and it is not switched off on the host |
| Customer | One customer |
| Region | One region, from the values in use |
| Tag | One tag, from the values in use |
| Kind | Servers or network devices |
| Alerts | With or without active alerts |
| Attention | One of the attention states below |
Every filter and the search live in the URL (/systems?status=OFFLINE&q=db).
Going back to the list, with the browser or a back link, shows it filtered
the way it was left, and a filtered list can be bookmarked or sent to
someone. The overview tiles link straight to /systems?attention=….
Reset clears all of them. The same holds for the customers list, the
agents list and the event timeline.
A system's, cluster's or customer's back link leads to the page you came from: Back to the customer when you opened the system from a customer, Back to alerts from the alerts page, with that page's filters kept. A page opened directly, from a mail or a bookmark, leads back to its list.
Clusters
Clusters are discovered, not configured. When an agent reports that its host is part of one, the cluster is created on first sight:
| Kind | Reported by |
|---|---|
| PostgreSQL cluster (Patroni) | the Patroni scope |
| Docker Swarm | the Swarm identifier |
| Garage cluster | the Garage node set |
| Proxmox VE cluster | the cluster name and a hash of its corosync key |
The name starts as the identifier the service itself uses — a Patroni scope reads well, a Swarm fingerprint less so — and can be renamed in the dashboard; the identifier underneath stays as reported. A Proxmox cluster starts under its own name; the hash beside it keeps two customers' clusters that share a name apart.
Hosts that stand beside a cluster rather than in it are joined automatically:
if an HAProxy's backends or an NFS server's clients are systems of one cluster,
that host is placed with the cluster in the role haproxy or nfs. Membership
an agent reported for itself always wins over this.
Deleting a cluster only forgets the grouping and its settings. The members stay, and their agents recreate the cluster with the next report as long as the service behind it still runs. Renaming and deleting a cluster is open to administrators, and to managers granted every member's customer.
Clusters matter beyond the display: an alert rule, an event rule and a maintenance window can each be scoped to a cluster, which then covers whichever systems currently belong to it.
The overview and the attention row
The dashboard home page shows four counters (customers, systems by status, active alerts, share of systems online), a row of attention tiles, the customers — one line each, by name, with their systems as status dots — and the five newest active alerts. The share of systems online leaves powered-off systems out of the count it divides by.
Every attention tile is a link to the systems list filtered to that state, except Collector error and Needs configuration, which open the Collectors page (see Agent Collectors): it says what each host lacks rather than only which hosts. A tile showing zero stays visible but muted — "nothing" and "not loaded" must not look alike. All of it is read from the current state, never from history.
| Tile | A system is counted when |
|---|---|
| Offline | its status is OFFLINE — a powered-off (STANDBY) system is not counted |
| Collector error | the agent reports a detected service whose collector is in the error state |
| Needs configuration | a detected service's collector is waiting for a setting, such as a token or a URL (needs_config) |
| Disk nearly full | it has an active or acknowledged alert on a disk series |
| Container unhealthy | its current docker_containers_unhealthy count is above 0 |
| Certificate expiring | it has an active or acknowledged alert on a tls_days_left or traefik_cert_days_left series |
| Agent outdated | it has reported an agent version and that version differs from the one this dashboard ships |
| Under maintenance | an open maintenance window covers it |
The first seven are faults or near-faults, in red or amber. Under maintenance is blue on purpose: that system is silent because somebody decided so, and it must not read as a failure.
Disk nearly full and Certificate expiring have no threshold of their own: they follow the alert rules, so the tile lights up exactly when an alert on those series is open, and goes dark when it resolves. Acknowledging the alert does not clear the tile; resolving it does. A fresh instance ships disk rules at 85 % and 95 % — to be warned only at 95 %, disable the 85 % rule. Container unhealthy reads the current value and needs no rule.
A system is marked OFFLINE when it was online and has not reported for more
than two minutes, and its offline alerts are on. That also writes a
system.offline event, which is why the tile and the timeline agree. With
offline alerts off it becomes STANDBY instead — see
Whether it alerts when it goes quiet.