Alert rules
An alert rule turns events into alerts. Each enabled rule compiles a CEL expression; the alerting engine evaluates every event on the bus against every rule of the event’s tenant. A match opens an alert — deduplicated by a key, optionally delayed by pendingFor, labelled, titled, and handed to an escalation policy — and an OK/resolve event closes it again. Rules are edited in Alerting → Alert rules (Alarm-Regeln) (with a built-in tester), through /api/v1/alert-rules (generic resource CRUD, objects:read/config:write, If-Match on PUT), or as bundle kind AlertRule. Every create/update/delete recompiles the tenant’s rules immediately.
How the engine runs
Section titled “How the engine runs”- One goroutine consumes the event queue; the queue is blocking, events are never dropped under load.
- Events that reach the rules:
state_changeandflapping_start/flapping_endfrom the check pipeline,ingressfrom every event source,heartbeat_missedfrom heartbeats,incident_updatewithaction: "created"fromPOST /api/v1/incidents, andsystemevents from the AI service. Lifecycle events the engine or API emit themselves (alert_opened,ack,notification,escalation,downtime,silence,config, engine-createdincident_update) are fan-out only and never re-enter the rules. - A 5-second ticker runs the timed work:
pendingForpromotion, heartbeat rules and heartbeat resources, auto-close, re-arming of suppressed alerts, and snooze wake-ups. - Disabled rules are skipped at load time. A rule that fails to compile is rejected on save with
422 np:validation/alert-rule(the test endpoints reportnp:validation/rule; a bundle apply stops with422 np:bundle/apply).
AlertRule fields
Section titled “AlertRule fields”| Field | Type / default | Semantics |
|---|---|---|
name |
string | Unique per tenant; used in default dedup keys and in alert_opened.rule |
disabled |
bool, false |
Skip the rule |
match |
CEL string | Required unless heartbeat is set; mutually exclusive with it |
heartbeat |
{source, expectEvery} |
Silence detection for an event source (name or id) or a heartbeat resource — see Heartbeat rules |
pendingFor |
duration, 0 |
Delay opening until the condition has been pending this long |
dedupKey |
Go template, empty | Overrides the default dedup key — see Dedup key |
severity |
critical | warning | info | ok, empty |
Alert severity; empty = the triggering event’s severity |
title |
Go template, empty | Alert title; defaults described under Title template |
autoCloseAfter |
duration, 0 |
Open and acked alerts of this rule older than this are flipped to expired (checked every 5 s); auto-close tickets are closed too; no alert_resolved event is emitted |
resolveOnOk |
bool, true |
When false, a clear event that does not match the rule does not resolve the alert (a clear event that does match still resolves) |
escalationPolicy |
policy name or id | Chain started when a new alert opens (unless suppressed) — see Escalation policies |
groupId |
alert-group name | Stored only; see Alert groups |
setLabels |
map | Merged over the event labels onto the alert (rule labels win) |
incident |
bool, false |
Every alert this rule opens gets its own incident (createdBy: rule:<name>); it auto-resolves when the last member alert resolves |
Durations are Go strings (30s, 5m, 24h); JSON also accepts a bare integer of seconds.
kind: AlertRulemetadata: { name: host-down-critical }spec: match: 'event.type == "state_change" && event.stateType == "hard" && (event.state == "CRITICAL" || event.state == "DOWN")' severity: critical title: "{{ .event.object }} is {{ .event.state }}" pendingFor: 2m autoCloseAfter: 24h escalationPolicy: ops-page setLabels: { np.sound: np_klaxon, team: sre } incident: trueThe CEL environment
Section titled “The CEL environment”Rules see exactly one variable, event, in a sandboxed CEL environment (no I/O, optimizer on, cost limit 10 000). Evaluation semantics:
- A missing key, field or index is a legitimate no-match — event shapes vary by type, so
event.payload.subjecton astate_changeevent just does not match. - Any other runtime error is logged and the event is neither a match nor a clear.
- A non-boolean result is a no-match.
| Path | Value |
|---|---|
event.type |
event type string: ingress, state_change, flapping_start, flapping_end, heartbeat_missed, incident_update, system |
event.severity |
critical | warning | info | ok |
event.ts |
RFC 3339 timestamp string |
event.objectId |
object id, or "" |
event.source |
sourceId — the EventSource id for ingress events, the heartbeat id for heartbeat_missed, otherwise "" |
event.labels.<k>, event.labels["k.v"] |
labels from the payload (NormEvent labels, object labels on state changes); {} when absent. Use the bracket form for keys with dots |
event.payload.<k> |
the raw payload map. For ingress events the inner archived body’s keys are hoisted to the top level (NormEvent keys win), so event.payload.subject, event.payload.from, event.payload.body (mail) or any field of a webhook/MQTT/SMS JSON body are addressable directly |
event.object, event.host, event.kind |
object name, host name, host/service — state_change only |
event.fromState, event.state |
OK, WARNING, CRITICAL, UNKNOWN, UP, DOWN, UNREACHABLE on state changes; for every other event state is derived from severity: CRITICAL, WARNING, OK or INFO |
event.stateType |
soft | hard (state_change) |
event.attempt |
check attempt number (state_change) |
event.output, event.summary |
check output / event summary (summary falls back to output, else "") |
event.metric |
dominant perfdata label (state_change) |
event.dedupKey |
the NormEvent’s dedup key (ingress) |
Examples that are known to work:
event.type == "state_change" && event.stateType == "hard" && (event.state == "CRITICAL" || event.state == "DOWN")event.type == "state_change" && event.kind == "service" && event.labels.env == "prod" && event.state != "OK"event.payload.subject.matches('(?i)feuer|brand') # mail subject, case-insensitive regexevent.payload.from == '[email protected]' # mail headerevent.payload.body.contains('Zone 12') # mail bodyevent.labels.source == 'mail-line' && event.severity == 'critical'event.summary.matches('FEUER.*Halle')event.labels.topic == 'factory/hall3/fire' # MQTT topicevent.labels["espa.priority"] == "1" # ESPA priority (dotted key)event.type == "incident_update" && event.payload.action == "created"event.type == "heartbeat_missed"event.type == "ingress" && event.source == "0199a0b1-…" # a specific source by idCEL string functions you will use most: ==, contains(), startsWith(), endsWith(), matches() (RE2 regex), in for list membership, &&/||/!, has(event.labels.env) to test presence without triggering a no-match.
Open and clear semantics
Section titled “Open and clear semantics”A matching event is a clear when event.severity == "ok", or event.state is OK/UP, or payload.resolve == true. A clear resolves the open or acked alert with the same dedup key (and discards a pending draft), stops its chain, emits alert_resolved, and may auto-resolve a rule-created incident. Any other match opens:
- The alert draft is upserted by dedup key. If an open/acked alert with that key exists, it is refreshed instead — severity can only rise, title and payload are replaced by the newest event, the event id is appended (last 50 kept) — and no new chain starts.
- A genuinely new alert emits
alert_opened{alertId, title, severity, rule, labels}, opens its incident whenincident: true, then passes the suppression gate (downtime, silence, flapping, parent host down — see Reliability). Not suppressed →StartChainwithescalationPolicy; suppressed → anotificationevent withstatus: suppressedand the reason, and the chain starts later if suppression lifts while the alert is still open.
Resolve-on-OK: a clear event that does not match the rule also resolves — the engine recomputes the dedup key for that event and resolves whatever is open under it — unless resolveOnOk: false. This is what lets a rule that only matches problem states (event.state != "OK") close its alerts on recovery.
Dedup key
Section titled “Dedup key”The dedup key decides whether an event refreshes an existing alert or opens a new one. Open and acked alerts are unique per (tenant, dedupKey).
Default, when dedupKey is empty:
<objectId>/<ruleName>if the event has an object (one alert per object and rule);- else
<ruleName>/<event.dedupKey>if the normalized event carries a dedup key (webhook mapping, Alertmanager fingerprint, mail Message-ID, trap, ESPA-X call id …); - else
<ruleName>/<sourceId>(one alert per rule and source).
A custom dedupKey is a Go text/template with data {{ .event.* }} (the CEL view, lowercase), {{ .object.id }} and {{ .rule.name }}, rendered with missingkey=zero. Examples: {{ .rule.name }}/{{ .event.labels.host }}, {{ .event.labels.topic }}, fire/{{ .event.payload.zone }}. A template that fails to parse yields <ruleName>/badtemplate (every event folds into one alert — check the tester). Heartbeat rules always use heartbeat/<ruleName>.
Title template
Section titled “Title template”title is a Go template whose only data is {{ .event.* }} — the same view the CEL expression sees. Correct forms: {{ .event.summary }}, {{ .event.object }} is {{ .event.state }}, {{ .event.payload.subject }}, {{ .event.labels.host }}: {{ .event.output }}. If the template is empty, fails, or renders empty, the title falls back to the event summary, then <object> is <state>, then the rule name.
Pending, auto-close, labels and severity
Section titled “Pending, auto-close, labels and severity”pendingFor— the first matching event stores a pending draft keyed by dedup key; further matches refresh it (newest title/payload win); every 5 s drafts whose first match is at leastpendingForold are opened. A clear event for the same dedup key before that deletes the draft and nothing opens. The pending map is in memory: a restart forgets drafts, and they are rebuilt from the next matching event.autoCloseAfter— every 5 s, open and acked alerts of the rule withopenedAtolder than the duration becomeexpired(not resolved, no event; the escalation engine treatsexpiredlike resolved and marks remaining timers done). Useful for “informational” rules whose alerts never get an explicit clear.setLabels— merged over the event labels; this is where you steer outputs:np.sound,np.volume,np.overrideSilentfor the alarm app,np.ttsto override the spoken text of voice calls, and any label your escalation actions, silences (selector) or outgoing webhooks (selector) should see. Alert labels are also what the correlator clusters on.severity— empty means “inherit from the event”. Because a refreshed alert’s severity can only rise, a rule without a fixed severity can escalate a warning alert to critical when a critical event with the same dedup key arrives.incident— one incident per alert; see Alerts and incidents.
Escalation hookup
Section titled “Escalation hookup”escalationPolicy names (or ids) an escalation policy. The chain starts when a new alert opens and is not suppressed; step offsets count from openedAt. Refreshes of an existing alert never restart the chain; an acknowledgement ends it; a snooze restarts it from step 0 at the wake-up time. A rule without a policy still opens visible alerts (UI, API, SSE, outgoing webhooks) — nobody is paged.
Heartbeat rules
Section titled “Heartbeat rules”Instead of match, a rule can watch for silence of an event source:
kind: AlertRulemetadata: { name: sensor-gateway-silent }spec: heartbeat: source: sensor-gateway # EventSource name or id — or a Heartbeat resource name expectEvery: 10m severity: warningThe engine records the last time it saw any event carrying a sourceId (every ingress event counts as a beat). heartbeat.source may be the event source’s name or id; a name is resolved per evaluation, so renaming the source and the rule together keeps working. It may also name a Heartbeat resource — its beats (POST /heartbeats/{name}/beat) then count as “seen”, which is what the demo seed’s demo-heartbeat-rule relies on. A heartbeat rule arms after the first event/beat; from then on, if nothing arrived for longer than expectEvery, it opens an alert with dedup key heartbeat/<ruleName> and title No event from "<source>" for <duration> (expected every <expectEvery>); as soon as events resume, the alert is resolved by that key. The event last-seen map is in memory — after a restart an event-source rule re-arms only once the source sends again (heartbeat-resource beats are persisted).
Testing rules
Section titled “Testing rules”Both test endpoints are side-effect free and need only alerts:read:
| Endpoint | Body | Result |
|---|---|---|
POST /api/v1/alert-rules:test |
{"rule": {…AlertRule…}, "demoEvents": [Event…]?, "from"?, "to"?} — without demoEvents and from, zero events are evaluated |
{"matched": n, "wouldOpen": [alert drafts, one per dedup key], "sampleViews": [≤ 5 CEL views]} |
POST /api/v1/alert-rules/{name}:test |
optional {"demoEvents"?, "from"?, "to"?}; default window = the last 24 h of stored events (max 1000) |
same |
matched counts every matching event, including clear events; wouldOpen is deduplicated by dedup key; sampleViews shows the exact event object the CEL expression saw — the fastest way to discover field names. Compile errors return 422 np:validation/rule.
curl -s -X POST "$NP/api/v1/alert-rules:test" -H "Authorization: Bearer $TOK" -H 'Content-Type: application/json' -d '{ "rule": {"name":"mail-fire","match":"event.payload.subject.matches(\"(?i)feuer|brand\")","severity":"critical", "title":"{{ .event.payload.subject }}"}, "demoEvents": [{"type":"ingress","severity":"warning","sourceId":"mail-1", "payload":{"summary":"FEUER Halle 3","labels":{"source":"ops-mailbox"},In the UI, every rule row has a Test rule (Regel testen) button that evaluates the stored rule against the last 24 hours and lists what would open; the edit dialog has a test panel with editable demo-event JSON. Escalation policies have their own simulator (:simulate) — see Escalation policies.
Alert groups (configuration only)
Section titled “Alert groups (configuration only)”/api/v1/alert-groups (bundle kind AlertGroup, Alerting → Groups) stores {name, groupBy: [label keys], window, aggregate: count|min|max|avg|sum|median, valuePath, minCount} and a rule can reference one via groupId. No runtime code evaluates alert groups today — they are stored, bundled and shown, nothing more. Storm handling that actually runs is the correlator (five alerts sharing a label pair within 120 s become one incident), described in Alerts and incidents. Treat alert groups as reserved configuration.
Related
Section titled “Related”- Every event field you can match on, per source type: Event sources
- What happens to an alert after it opens: Acknowledge and snooze, Escalation policies
- Suppression, re-arming, what a restart forgets: Reliability
- The event model and type catalog: Events
- REST reference:
post_alert_rules,post_alert_rules_test,post_alert_rules_name_test