~/tech-with-ugur

Recording Rules Done Right: One Dataset, Two Grafana Dashboards, and Alerts Straight to Your Own Webhook

2026-08-25 observability

Run the companion lab

helm install a monitoring stack, port-forward Grafana, admire the built-in dashboards — that’s where nearly every kube-prometheus-stack tutorial ends. It’s also the exact point where the real work begins, because none of those built-ins answer the questions your team will actually ask: which of our nodes is hottest, which pods are about to hit their limits, and who tells us — not a SaaS pager — when something crosses a line.

This lab builds that whole layer on top of the chart, declaratively, and then proves it works by breaking the cluster on purpose: recording rules that pre-compute utilization keyed by human-readable node names, two dashboards with byte-identical structure that make the case for recording rules in a single screenshot, and Grafana-managed alerts delivered to a small TypeScript server you own. Everything runs in a local kind cluster; one make up from a fresh clone brings all of it up.

The first decision happens before anything is installed: pin the chart. kube-prometheus-stack moves fast and breaks label conventions along the way; this lab pins 88.5.4 (and kindest/node:v1.36.1, Kubernetes 1.36), so the joins you’re about to see keep working months from now no matter what later chart releases rename.

The node-name problem

node-exporter runs as a DaemonSet and Prometheus scrapes each pod, so every node-level series carries an instance label — an IP:port pair. Your worst-nodes panel ends up ranking 10.244.1.3:9100 against 10.244.2.4:9100, which is useless to a human at 3 a.m.

The fix lives in a metric most people never notice: node_uname_info has a value of 1 and a nodename label with the real name. Multiply your expression by it with an on(instance) group_left(nodename) join and the name comes along for the ride. Do that once and you never want to type it again — which is precisely the argument for recording rules.

From rules/recording-rules.yaml:

apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: lab-recording-rules
  namespace: monitoring
  labels:
    release: kps          # required: the chart's Prometheus only selects rules with the release label
spec:
  groups:
    - name: lab.node.utilization
      interval: 15s
      rules:
        # node-exporter only knows scrape instances (IP:port); node_uname_info
        # carries the human node name. The join + label_replace attach it as `node`.
        - record: node:cpu_utilization:percent
          expr: |
            label_replace(
              100 * (1 - avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])))
                * on(instance) group_left(nodename) node_uname_info,
              "node", "$1", "nodename", "(.+)")
        # ...

Two details here bite everyone the first time. The release: kps label is not decoration — the chart’s Prometheus is configured to select only PrometheusRule objects carrying its own release label, and a rule without it is silently ignored. And the label_replace wrapper is what turns the joined-in nodename into a proper node label, so everything downstream can filter on node=~"lab-kps-worker.*" like it always wished it could.

Pods need a different join. Their utilization is measured against their own resource limits, and the node name comes from kube_pod_info rather than node_uname_info — same file:

    - name: lab.pod.utilization
      interval: 15s
      rules:
        # Pod usage as % of its declared limit; kube_pod_info supplies the node name.
        - record: pod:cpu_utilization_vs_limit:percent
          expr: |
            100 * sum by (namespace, pod) (rate(container_cpu_usage_seconds_total{container!=""}[5m]))
              / sum by (namespace, pod) (kube_pod_container_resource_limits{resource="cpu"})
              * on(namespace, pod) group_left(node) max by (namespace, pod, node) (kube_pod_info)
        # ...

One honest caveat baked into that design: a pod with no limits set contributes no series at all. It can’t show up on the worst-pods panels and can’t fire the pod alerts, no matter how much it consumes. That’s a property of “percent of own limit” as a definition, and it’s worth stating out loud rather than discovering later.

The proof: two dashboards, same panels, different queries

Here’s the part that makes recording rules visceral instead of theoretical. The lab ships two dashboards whose JSON is structurally byte-identical — same six panels, same grid positions, same five template variables — differing only in the queries. One asks Prometheus the raw question; the other asks for the recorded answer.

The worst-nodes CPU panel in dashboards/raw-queries.json:

topk($worst_x, 100 * (1 - avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m]))) * on(instance) group_left(nodename) node_uname_info{nodename=~"$node"})

The same panel in dashboards/recorded-metrics.json:

topk($worst_x, node:cpu_utilization:percent{node=~"$node"})

The pod version is worse in raw form — a three-way join across cadvisor, kube-state-metrics limits, and kube_pod_info — and collapses to topk($worst_x, pod:cpu_utilization_vs_limit:percent{namespace=~"$namespace"}) recorded. Open both dashboards side by side, hit Edit on any panel, and the whole argument fits in one screenshot: identical picture, and one of the two queries reads like plain English. Beyond readability, Prometheus evaluates the gnarly version once per rule interval instead of once per dashboard refresh per viewer — recording rules are also a load story.

A Grafana limitation surfaced while wiring the variables, and it’s instructive: the dashboards define $cpu_threshold / $mem_threshold variables, but Grafana cannot template panel fieldConfig thresholds (or alert conditions) from dashboard variables. The red 80% line in the panels and the 80 in the alert rules are both literals, kept in sync by hand. Of the five variables, only $worst_x and the $node / $namespace filters actually do anything — the threshold variables exist to document the intended knob that Grafana doesn’t offer.

Dashboards as code: the sidecar trick

Neither dashboard was ever clicked together in the UI. The chart’s Grafana ships with a sidecar that watches for ConfigMaps carrying a label and imports whatever they contain. The lab’s helm/values.yaml, trimmed to the interesting part:

alertmanager:
  enabled: false         # alerting is Grafana-managed in this lab
# ...
grafana:
  defaultDashboardsEnabled: false   # only the lab's two dashboards
  # ...
  sidecar:
    dashboards:
      enabled: true
      label: grafana_dashboard
      labelValue: "1"
      searchNamespace: monitoring
    alerts:
      enabled: true               # off by default in the chart — this lab's key switch
      label: grafana_alert
      labelValue: "1"
      searchNamespace: monitoring

Note alertmanager.enabled: false — more on that in a moment — and sidecar.alerts.enabled: true, which is off by default in the chart. That one line is the least-known switch in this whole setup: it gives Grafana-managed alerting the same ConfigMap-based provisioning path the dashboards use.

Getting a JSON file into a labeled ConfigMap without hand-writing YAML wrappers is a one-liner pipeline, from scripts/apply-observability.sh:

apply_labeled_configmap() {
  local name="$1" file="$2" label="$3"
  kubectl create configmap "$name" -n monitoring \
    --from-file="$(basename "$file")=$file" \
    --dry-run=client -o yaml \
    | kubectl label --local -f - "$label=1" -o yaml \
    | kubectl apply -f -
}

Dashboards stay as plain .json files you can edit and diff; alerting resources stay as plain .yaml; the script wraps and labels each one at apply time. Within a minute of kubectl apply, the sidecar has imported them.

Alerting without Alertmanager

Alertmanager is disabled in this cluster. Instead, the alerts are Grafana-managed: Grafana itself evaluates queries against Prometheus and handles delivery. The whole alerting layer is three declarative files shipped through that alerts sidecar.

Six alert rules — node CPU, node memory, node unhealthy, pod CPU, pod memory, pod unhealthy — and every one of them queries a recorded metric, never a raw join. This is the operational payoff of the rules: the alert definitions become trivially readable. The first rule in alerting/alert-rules.yaml:

apiVersion: 1
groups:
  - orgId: 1
    name: lab-utilization
    folder: lab-alerts
    interval: 10s
    rules:
      - uid: lab-node-cpu-high
        title: LabNodeCpuHigh
        condition: threshold
        for: 30s
        noDataState: OK
        execErrState: OK
        labels: { lab: kube-prometheus-recording-rules }
        annotations:
          summary: "Node {{ $labels.node }} CPU above 80%."
        data:
          - refId: query
            relativeTimeRange: { from: 300, to: 0 }
            datasourceUid: prometheus
            model: { refId: query, instant: true, expr: "node:cpu_utilization:percent" }
          - refId: threshold
            datasourceUid: __expr__
            model:
              refId: threshold
              type: threshold
              expression: query
              conditions:
                - evaluator: { type: gt, params: [80] }
      # ...

The {{ $labels.node }} in the summary is the recording rules paying rent again — the alert can name the node because the recorded metric carries the name.

Delivery is a webhook contact point plus a root notification policy, each its own file. alerting/contact-points.yaml:

apiVersion: 1
contactPoints:
  - orgId: 1
    name: webhook-app
    receivers:
      - uid: webhook-app
        type: webhook
        settings:
          url: http://webhook-app.webhook-app.svc:8080/alerts
          httpMethod: POST

and alerting/policies.yaml:

apiVersion: 1
policies:
  - orgId: 1
    receiver: webhook-app
    group_by: ["alertname"]
    group_wait: 10s
    group_interval: 30s
    repeat_interval: 4h

That URL points at the last piece: a deliberately small TypeScript server running in the cluster, whose only job is to receive Grafana’s webhook POSTs and log every alert as a structured JSON line — so you can see exactly what your code would get instead of a SaaS pager. The core of webhook-app/src/server/handler.ts:

  if (req.method === "POST" && req.url === "/alerts") {
    try {
      const body = await readBody(req);
      logger.info({ bytes: body.length }, "Receiving alert...");
      const payload: unknown = JSON.parse(body);
      logger.debug({ payload }, "Receiving alert payload.");
      for (const summary of summarizeAlerts(payload)) {
        logger.info(
          {
            alertname: summary.alertname,
            status: summary.status,
            labels: summary.labels,
          },
          "Receiving alert succeeded.",
        );
      }
      respond(res, 200, { ok: true });
    } catch (err) {
      logger.error({ err }, "Receiving alert failed.");
      respond(res, 400, { ok: false });
    }
    return;
  }

Grafana’s payload follows the Alertmanager webhook shape — an alerts array where each entry carries status, labels, and annotations — so kubectl logs on this pod during the finale prints one clean JSON line per alert, with the node or pod name sitting right in the labels.

Breaking things on purpose

A monitoring lab that never fires its alerts is a screenshot, not a proof. make verify checks the whole chain — recorded metrics carry node names, both dashboards imported, all six rules and the contact point provisioned — then deploys three fault workloads (a CPU hog, a memory hog, a crashloop pod) and stops the kubelet on one worker with a plain docker exec <node> systemctl stop kubelet, and finally polls the webhook server’s logs until all six alerts have arrived.

Implementation taught two lessons here that no tutorial mentions.

First: kind nodes share your machine’s kernel. node-exporter inside a kind “node” reports the host’s CPU and memory, so a single hog pod capped at 200m can never move a node-level metric past 80%. The verify script sizes the hog fleets from live capacity instead of hardcoding counts — and the memory math has a scar to show. An early version sized purely off a percentage of total memory; on a machine whose Docker VM already sat at 65% usage, it cheerfully added ~6.7Gi of tmpfs on top and took the kind API server down with it. The committed version sizes off the gap to the target and enforces an independent 90% ceiling, from scripts/verify.sh:

  local mem_replicas mem_capped
  read -r mem_replicas mem_capped <<< "$(awk -v tgt=85 -v ceiling=90 -v cur="$cur_mem" -v total="$mem_total_bytes" -v cap=128 'BEGIN {
    fill = 110 * 1024 * 1024
    extra_to_target = (tgt - cur) / 100 * total
    extra_to_ceiling = (ceiling - cur) / 100 * total
    extra = extra_to_target
    if (extra_to_ceiling < extra) extra = extra_to_ceiling
    if (extra < 0) extra = 0
    v = extra / fill
    if (v < 1) v = 1
    r = int(v); if (v > r) r++
    capped = 0
    if (r > cap) { r = cap; capped = 1 }
    print r, capped
  }')"

Fault injection that can kill the thing doing the observing is a genre of bug worth knowing about before you build chaos tooling for real clusters.

Second: pick your victim in advance. If the kubelet you stop happens to be on the node running Grafana, you’ve blinded yourself mid-demo. The lab pins the entire observation plane — Prometheus, Grafana, kube-state-metrics, the operator, the webhook server, and the fault pods themselves — to worker 1 via nodeSelector, so worker 2 exists purely to be sacrificed. Every run breaks the same node, Grafana stays reachable throughout, and LabNodeUnhealthy fires because kube-state-metrics (safe on worker 1) watches worker 2 go NotReady. A shortened node-monitor-grace-period of 20s in kind/cluster.yaml makes that detection take seconds instead of most of a minute.

The finale, watched from the webhook pod’s logs: the hogs push node CPU and memory over 80, the crashloop pod trips CrashLoopBackOff, worker 2 goes dark — and six distinct alertnames land as JSON, one after another, in a server you wrote yourself.

Run it

make up       # kind cluster + chart + rules/dashboards/alerting/webhook
make verify   # fires and observes all six alerts (~10-15 min; stops a kubelet — the cluster is disposable)
make down     # deletes the cluster

The lab repository has the full tree — every file shown here plus the verify script that checks the whole chain end to end. Port-forward Grafana (kubectl port-forward -n monitoring svc/kps-grafana 3000:80, password in the kps-grafana secret), open the two dashboards side by side, and enjoy the one screenshot that ends the “do we really need recording rules?” conversation.