Recording Rules Done Right: One Dataset, Two Grafana Dashboards, and Alerts Straight to Your Own Webhook
helm install a monitoring stack, port-forward Grafana, admire the
built-in dashboards — that’s where nearly every kube-prometheus-stack
tutorial ends. It’s also the exact point where the real work begins,
because none of those built-ins answer the questions your team will
actually ask: which of our nodes is hottest, which pods are about to
hit their limits, and who tells us — not a SaaS pager — when something
crosses a line.
This lab builds that whole layer on top of the chart, declaratively, and
then proves it works by breaking the cluster on purpose: recording rules
that pre-compute utilization keyed by human-readable node names, two
dashboards with byte-identical structure that make the case for
recording rules in a single screenshot, and Grafana-managed alerts
delivered to a small TypeScript server you own. Everything runs in a
local kind cluster; one make up from a
fresh clone brings all of it up.
The first decision happens before anything is installed: pin the chart.
kube-prometheus-stack moves fast and breaks label conventions along the
way; this lab pins 88.5.4 (and kindest/node:v1.36.1, Kubernetes
1.36), so the joins you’re about to see keep working months from now no
matter what later chart releases rename.
The node-name problem
node-exporter runs as a DaemonSet and Prometheus scrapes each pod, so
every node-level series carries an instance label — an IP:port pair.
Your worst-nodes panel ends up ranking 10.244.1.3:9100 against
10.244.2.4:9100, which is useless to a human at 3 a.m.
The fix lives in a metric most people never notice: node_uname_info
has a value of 1 and a nodename label with the real name. Multiply
your expression by it with an on(instance) group_left(nodename) join
and the name comes along for the ride. Do that once and you never want
to type it again — which is precisely the argument for recording rules.
From rules/recording-rules.yaml:
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: lab-recording-rules
namespace: monitoring
labels:
release: kps # required: the chart's Prometheus only selects rules with the release label
spec:
groups:
- name: lab.node.utilization
interval: 15s
rules:
# node-exporter only knows scrape instances (IP:port); node_uname_info
# carries the human node name. The join + label_replace attach it as `node`.
- record: node:cpu_utilization:percent
expr: |
label_replace(
100 * (1 - avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])))
* on(instance) group_left(nodename) node_uname_info,
"node", "$1", "nodename", "(.+)")
# ...
Two details here bite everyone the first time. The release: kps label
is not decoration — the chart’s Prometheus is configured to select only
PrometheusRule objects carrying its own release label, and a rule
without it is silently ignored. And the label_replace wrapper is what
turns the joined-in nodename into a proper node label, so everything
downstream can filter on node=~"lab-kps-worker.*" like it always
wished it could.
Pods need a different join. Their utilization is measured against their
own resource limits, and the node name comes from kube_pod_info
rather than node_uname_info — same file:
- name: lab.pod.utilization
interval: 15s
rules:
# Pod usage as % of its declared limit; kube_pod_info supplies the node name.
- record: pod:cpu_utilization_vs_limit:percent
expr: |
100 * sum by (namespace, pod) (rate(container_cpu_usage_seconds_total{container!=""}[5m]))
/ sum by (namespace, pod) (kube_pod_container_resource_limits{resource="cpu"})
* on(namespace, pod) group_left(node) max by (namespace, pod, node) (kube_pod_info)
# ...
One honest caveat baked into that design: a pod with no limits set contributes no series at all. It can’t show up on the worst-pods panels and can’t fire the pod alerts, no matter how much it consumes. That’s a property of “percent of own limit” as a definition, and it’s worth stating out loud rather than discovering later.
The proof: two dashboards, same panels, different queries
Here’s the part that makes recording rules visceral instead of theoretical. The lab ships two dashboards whose JSON is structurally byte-identical — same six panels, same grid positions, same five template variables — differing only in the queries. One asks Prometheus the raw question; the other asks for the recorded answer.
The worst-nodes CPU panel in dashboards/raw-queries.json:
topk($worst_x, 100 * (1 - avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m]))) * on(instance) group_left(nodename) node_uname_info{nodename=~"$node"})
The same panel in dashboards/recorded-metrics.json:
topk($worst_x, node:cpu_utilization:percent{node=~"$node"})
The pod version is worse in raw form — a three-way join across cadvisor,
kube-state-metrics limits, and kube_pod_info — and collapses to
topk($worst_x, pod:cpu_utilization_vs_limit:percent{namespace=~"$namespace"})
recorded. Open both dashboards side by side, hit Edit on any panel, and
the whole argument fits in one screenshot: identical picture, and one of
the two queries reads like plain English. Beyond readability, Prometheus
evaluates the gnarly version once per rule interval instead of once per
dashboard refresh per viewer — recording rules are also a load story.
A Grafana limitation surfaced while wiring the variables, and it’s
instructive: the dashboards define $cpu_threshold / $mem_threshold
variables, but Grafana cannot template panel fieldConfig thresholds
(or alert conditions) from dashboard variables. The red 80% line in the
panels and the 80 in the alert rules are both literals, kept in sync by
hand. Of the five variables, only $worst_x and the $node /
$namespace filters actually do anything — the threshold variables
exist to document the intended knob that Grafana doesn’t offer.
Dashboards as code: the sidecar trick
Neither dashboard was ever clicked together in the UI. The chart’s
Grafana ships with a sidecar that watches for ConfigMaps carrying a
label and imports whatever they contain. The lab’s helm/values.yaml,
trimmed to the interesting part:
alertmanager:
enabled: false # alerting is Grafana-managed in this lab
# ...
grafana:
defaultDashboardsEnabled: false # only the lab's two dashboards
# ...
sidecar:
dashboards:
enabled: true
label: grafana_dashboard
labelValue: "1"
searchNamespace: monitoring
alerts:
enabled: true # off by default in the chart — this lab's key switch
label: grafana_alert
labelValue: "1"
searchNamespace: monitoring
Note alertmanager.enabled: false — more on that in a moment — and
sidecar.alerts.enabled: true, which is off by default in the
chart. That one line is the least-known switch in this whole setup: it
gives Grafana-managed alerting the same ConfigMap-based provisioning
path the dashboards use.
Getting a JSON file into a labeled ConfigMap without hand-writing YAML
wrappers is a one-liner pipeline, from scripts/apply-observability.sh:
apply_labeled_configmap() {
local name="$1" file="$2" label="$3"
kubectl create configmap "$name" -n monitoring \
--from-file="$(basename "$file")=$file" \
--dry-run=client -o yaml \
| kubectl label --local -f - "$label=1" -o yaml \
| kubectl apply -f -
}
Dashboards stay as plain .json files you can edit and diff; alerting
resources stay as plain .yaml; the script wraps and labels each one at
apply time. Within a minute of kubectl apply, the sidecar has imported
them.
Alerting without Alertmanager
Alertmanager is disabled in this cluster. Instead, the alerts are Grafana-managed: Grafana itself evaluates queries against Prometheus and handles delivery. The whole alerting layer is three declarative files shipped through that alerts sidecar.
Six alert rules — node CPU, node memory, node unhealthy, pod CPU, pod
memory, pod unhealthy — and every one of them queries a recorded
metric, never a raw join. This is the operational payoff of the rules:
the alert definitions become trivially readable. The first rule in
alerting/alert-rules.yaml:
apiVersion: 1
groups:
- orgId: 1
name: lab-utilization
folder: lab-alerts
interval: 10s
rules:
- uid: lab-node-cpu-high
title: LabNodeCpuHigh
condition: threshold
for: 30s
noDataState: OK
execErrState: OK
labels: { lab: kube-prometheus-recording-rules }
annotations:
summary: "Node {{ $labels.node }} CPU above 80%."
data:
- refId: query
relativeTimeRange: { from: 300, to: 0 }
datasourceUid: prometheus
model: { refId: query, instant: true, expr: "node:cpu_utilization:percent" }
- refId: threshold
datasourceUid: __expr__
model:
refId: threshold
type: threshold
expression: query
conditions:
- evaluator: { type: gt, params: [80] }
# ...
The {{ $labels.node }} in the summary is the recording rules paying
rent again — the alert can name the node because the recorded metric
carries the name.
Delivery is a webhook contact point plus a root notification policy,
each its own file. alerting/contact-points.yaml:
apiVersion: 1
contactPoints:
- orgId: 1
name: webhook-app
receivers:
- uid: webhook-app
type: webhook
settings:
url: http://webhook-app.webhook-app.svc:8080/alerts
httpMethod: POST
and alerting/policies.yaml:
apiVersion: 1
policies:
- orgId: 1
receiver: webhook-app
group_by: ["alertname"]
group_wait: 10s
group_interval: 30s
repeat_interval: 4h
That URL points at the last piece: a deliberately small TypeScript
server running in the cluster, whose only job is to receive Grafana’s
webhook POSTs and log every alert as a structured JSON line — so you can
see exactly what your code would get instead of a SaaS pager. The core
of webhook-app/src/server/handler.ts:
if (req.method === "POST" && req.url === "/alerts") {
try {
const body = await readBody(req);
logger.info({ bytes: body.length }, "Receiving alert...");
const payload: unknown = JSON.parse(body);
logger.debug({ payload }, "Receiving alert payload.");
for (const summary of summarizeAlerts(payload)) {
logger.info(
{
alertname: summary.alertname,
status: summary.status,
labels: summary.labels,
},
"Receiving alert succeeded.",
);
}
respond(res, 200, { ok: true });
} catch (err) {
logger.error({ err }, "Receiving alert failed.");
respond(res, 400, { ok: false });
}
return;
}
Grafana’s payload follows the Alertmanager webhook shape — an alerts
array where each entry carries status, labels, and annotations —
so kubectl logs on this pod during the finale prints one clean JSON
line per alert, with the node or pod name sitting right in the labels.
Breaking things on purpose
A monitoring lab that never fires its alerts is a screenshot, not a
proof. make verify checks the whole chain — recorded metrics carry
node names, both dashboards imported, all six rules and the contact
point provisioned — then deploys three fault workloads (a CPU hog, a
memory hog, a crashloop pod) and stops the kubelet on one worker with a
plain docker exec <node> systemctl stop kubelet, and finally polls the
webhook server’s logs until all six alerts have arrived.
Implementation taught two lessons here that no tutorial mentions.
First: kind nodes share your machine’s kernel. node-exporter inside
a kind “node” reports the host’s CPU and memory, so a single hog pod
capped at 200m can never move a node-level metric past 80%. The verify
script sizes the hog fleets from live capacity instead of hardcoding
counts — and the memory math has a scar to show. An early version sized
purely off a percentage of total memory; on a machine whose Docker VM
already sat at 65% usage, it cheerfully added ~6.7Gi of tmpfs on top and
took the kind API server down with it. The committed version sizes off
the gap to the target and enforces an independent 90% ceiling, from
scripts/verify.sh:
local mem_replicas mem_capped
read -r mem_replicas mem_capped <<< "$(awk -v tgt=85 -v ceiling=90 -v cur="$cur_mem" -v total="$mem_total_bytes" -v cap=128 'BEGIN {
fill = 110 * 1024 * 1024
extra_to_target = (tgt - cur) / 100 * total
extra_to_ceiling = (ceiling - cur) / 100 * total
extra = extra_to_target
if (extra_to_ceiling < extra) extra = extra_to_ceiling
if (extra < 0) extra = 0
v = extra / fill
if (v < 1) v = 1
r = int(v); if (v > r) r++
capped = 0
if (r > cap) { r = cap; capped = 1 }
print r, capped
}')"
Fault injection that can kill the thing doing the observing is a genre of bug worth knowing about before you build chaos tooling for real clusters.
Second: pick your victim in advance. If the kubelet you stop happens
to be on the node running Grafana, you’ve blinded yourself mid-demo. The
lab pins the entire observation plane — Prometheus, Grafana,
kube-state-metrics, the operator, the webhook server, and the fault pods
themselves — to worker 1 via nodeSelector, so worker 2 exists purely
to be sacrificed. Every run breaks the same node, Grafana stays
reachable throughout, and LabNodeUnhealthy fires because
kube-state-metrics (safe on worker 1) watches worker 2 go NotReady. A
shortened node-monitor-grace-period of 20s in kind/cluster.yaml
makes that detection take seconds instead of most of a minute.
The finale, watched from the webhook pod’s logs: the hogs push node CPU
and memory over 80, the crashloop pod trips CrashLoopBackOff, worker 2
goes dark — and six distinct alertnames land as JSON, one after another,
in a server you wrote yourself.
Run it
make up # kind cluster + chart + rules/dashboards/alerting/webhook
make verify # fires and observes all six alerts (~10-15 min; stops a kubelet — the cluster is disposable)
make down # deletes the cluster
The lab repository
has the full tree — every file shown here plus the verify script that
checks the whole chain end to end. Port-forward Grafana
(kubectl port-forward -n monitoring svc/kps-grafana 3000:80, password
in the kps-grafana secret), open the two dashboards side by side, and
enjoy the one screenshot that ends the “do we really need recording
rules?” conversation.