Reports, uptime and SLA

What the numbers mean, where they come from, and what a client can see for themselves — because a report you cannot explain is one you cannot charge for.

On this page
  1. How uptime is calculated
  2. Downtime and MTTR
  3. SLA: the promise, and the budget it implies
  4. How far back a report can go
  5. What the client sees

How uptime is calculated#

Uptime is the share of checks in the period that succeeded. Not an estimate from incident durations — the checks themselves, counted. A period with 43,200 checks and six failures reads 99.99%, and the six are in the incident list underneath, so the figure and its cause are on the same page.

Checks taken inside a maintenance window count as neither up nor down. They are excluded from both sides of the fraction, and the report says so in words when any were excluded — a percentage that quietly ignores an hour is worse than a lower percentage.

Downtime and MTTR#

Downtime
The total length of the incidents in the period. An incident that is still open has no length yet and is shown as ongoing rather than counted at zero.
MTTR
The mean time to resolve, over the incidents that were resolved inside the period. An incident that started last month and closed this one is not smuggled into the average.
Comparison
Every headline figure is shown against the previous period of the same length. Where there is nothing to compare against, the line is absent rather than reading "0%", which is a claim.

SLA: the promise, and the budget it implies#

A target percentage can be set on a plan, and overridden for one client. Once it exists, the report carries a block comparing what was promised with what was delivered, and an error budget: the amount of downtime the target itself allows in the period.

99.5% over thirty days is about three and a half hours. The bar on the report fills with the downtime actually used, so a mostly-empty bar is good news without needing a caption. It is the number worth showing a client who asks why they are paying for monitoring in a month when nothing broke.

If a target is missed, the product can calculate the credit due under the policy you set, and issue it as a credit line on an invoice. It never issues one silently — somebody presses the button.

How far back a report can go#

Check history is retained for as long as the plan says — 90 days on the starting plan. Ask for a longer period and the report is produced for the period that actually has data, with a note on its face saying where it starts and why. A report that silently answers a different question from the one asked is the failure mode here, so it is made impossible to miss.

What the client sees#

Their portal
A login of their own showing their sites, their uptime, their incidents and their invoices — under your name, logo and colour, and nothing belonging to any other client.
A PDF
The uptime report as a document, with the SLA block, per-site availability and the incident list. It can be generated on demand, delivered on a schedule — the schedule needs e-mail, see Alerts — or produced by the client themselves from their portal, choosing the period and which of their sites to include. Either way the document is stored, so opening it again returns the same file rather than recomputing it.
A public link
A read-only page for a period, at an address that carries its own token — for the person at the client who does not want a login.
CSV
The same figures, for a spreadsheet.