Skip to main content

Maintenance Windows

A maintenance window is a period you announce in advance so that planned work does not wake anybody. While it is open, the systems it covers raise nothing and notify nobody — but everything that happens to them is still recorded.

Windows live under Settings → Maintenance windows.

What a window suppresses, and what it does not​

While a window is open, for every system it covers:

  • No new alert is raised. A threshold crossed during the window creates nothing; there is no alert to acknowledge afterwards and no mail about it.
  • No event notifies. The reboot, the failover, the unit that failed and came back — all of them are matched against your event rules as usual and then held back.

What still happens, exactly as on any other day:

  • Metrics keep arriving and are stored. Charts have no gap over a maintenance window.
  • Events are recorded. The reboot you performed is on the timeline, marked with its real time. That is the point: the window changes who gets woken, not what is true.
  • Status still changes. A host that stops reporting is still marked OFFLINE, and the system.offline event is written — it just does not notify.
  • Alerts that were already open keep being updated and still resolve.

That last one is worth dwelling on. Opening a window does not close the alerts that are already active; it stops new ones. An alert that was open when the window started continues to have its current value refreshed, and if the value comes back inside its threshold for a full rule duration, it resolves normally — quietly, since resolutions do not notify either way. The effect is that the quiet ends with a clean slate rather than with a set of alerts frozen mid-incident that you have to sort out by hand.

If you want an already-open alert gone before the work starts, acknowledge or resolve it on the alerts page first.

Scopes​

A window has exactly one scope:

ScopeCovers
All systemsEverything this instance monitors.
CustomerEvery system belonging to that customer.
ClusterEvery system currently in that cluster, plus events attributed to the cluster itself.
SystemOne server.

Cluster and customer scopes follow membership as it stands at the moment of evaluation, so a node that joins the cluster during the window is covered without anyone editing the window.

Creating one​

New window asks for a name, the scope and, when the scope is not all systems, the target; a start and an end; and an optional note. The start defaults to the next five-minute mark and the end to two hours later. The times are entered and displayed in your browser's local time zone.

The name is what you will read in the list afterwards — Kernel update database cluster is useful, maintenance is not. The note is the place for a ticket number or a phone number.

The end must be after the start. A window whose start is in the future is only upcoming: it does nothing until its start time arrives, and then becomes active without anyone doing anything.

Administrators and managers may create, end and delete windows — a manager only where the scope is one it reaches in full: a system or a customer it has been granted, or a cluster whose every member sits at a customer it has been granted. A window over all systems silences every customer in the instance, which is more than any grant can add up to, so that one is administrator-only.

Operators and viewers create nothing: the API refuses them too, not just the interface, and the settings navigation offers this page to administrators and managers only. See Users and Permissions.

Repeating windows​

A window that recurs — the patch night on the second Tuesday, the backup hour every night — is created once. Repeat offers, worded for the start you picked:

  • Every day
  • Every week on the start's weekday
  • Every month on the same date, e.g. the 8th. A month without that date — the 31st in April — uses its last day.
  • Every month on the same weekday, e.g. the second Tuesday. A start in a month's last week, a fifth Tuesday, means the last Tuesday of every month.

Each occurrence starts at the same wall-clock time and lasts as long as the first. The time is kept in the time zone of the browser that created the window, so a window at 02:00 stays at 02:00 across the change to and from summer time. An occurrence must be shorter than its period — at most a day when daily, a week when weekly, 28 days when monthly — so occurrences never overlap.

Repeat until is optional: occurrences that would start after that day do not happen. Left empty, the window repeats until it is deleted.

In the list a repeating window shows how it repeats under its name, and its start and end are those of the open occurrence or, between occurrences, of the next one; it sits under Upcoming between occurrences and under Past once its last occurrence is over. End now on a repeating window ends only the occurrence that is open — the next one starts as planned. To stop the series, delete it.

The list​

Windows are grouped into Active, Upcoming and Past (last 30 days), the past section hidden until you ask for it. Each row shows the name, the scope and its target, start, end, the note and who created it. Active rows carry a badge and a tint.

The list holds the windows the account reaches, so a window naming another customer's system — and the name of whoever opened it — is not in it. The instance-wide window is the exception and stays visible to everybody: it silences their systems too, and a host marked Maintenance with nothing on screen to explain it is worse than seeing a window you did not open.

An active window has an End now button, which sets its end to this moment. Use it when the work finished early — the alternative is a quiet estate for another two hours. Deleting a window removes it from the list entirely; ending it keeps the record of what was silenced and when.

Systems covered by an open window are also marked in two other places: a Maintenance badge next to the name in the systems list, and the blue Under maintenance tile on the overview, which links to those systems. That tile is blue on purpose — the system is quiet because somebody decided so, and it must not read as a fault.

The permanent version: an inactive customer​

Switching a customer to inactive produces exactly the same silence, for all its systems, for as long as the switch stays off: no alert is raised, no notification is sent, everything is still recorded, and alerts already open keep updating and resolving. There is no end time — it lasts until somebody switches the customer back on.

Use a window for work with a beginning and an end. Use the inactive switch for a customer you are no longer on call for, or one whose infrastructure you are keeping data on but no longer watching. See Customers and Systems for what else the switch changes.

Practical notes​

  • Create the window before you start. A window that opens at 21:00 does nothing about the alert that fired at 20:58.
  • Scope as tightly as the work deserves. All systems while you reboot one database node means a genuine outage elsewhere goes unreported for the duration.
  • Give it a generous end. Ending early is one click; a window that expired while the cluster was still resyncing produces the alert storm you were trying to avoid.
  • After the window closes, the next evaluation run judges the current state afresh. Anything still wrong is raised then — so the alerts you get after a window are the ones that are still true, not a replay of the ones you missed.