> Per AI Blitz Plan §platform.procedures. Layer: 1. Repo: devops-infra-helm-charts. # Procedure — Modify Prometheus / VictoriaMetrics alert rules > **Layer:** Layer 1 — values diff + PR. > **Blast radius:** PagerDuty / on-call paging across the cluster's tenants. A wrong threshold pages everyone; a missing rule masks an outage. > **Approval:** observability owner + on-call lead for the affected severity. This procedure is the narrow alert-rule slice of [modify-observability-config.md](modify-observability-config.md). Use this when the entire intent of the PR is alert-rule editing — adding, removing, retuning thresholds, changing `for:` windows, or rewriting `expr:` PromQL. --- ## Where alert rules live | File | Engine | Notes | |------|--------|-------| | `helm-overrides//vmalert/custom-values.yaml` | vmalert | Most clusters use vmalert against VM cluster as the primary alerting engine. Rules typically under `config.rules` or rendered as a `VMRule` CRD. | | `helm-overrides//victoria-metrics-cluster/custom-values.yaml` | vmalert (bundled) / VM ruler | Some clusters carry rules in the cluster chart's `vmalert.config` block. | | `helm-overrides//prometheus/custom-values.yaml` | Prometheus / `kube-prometheus-stack` | Rules under `additionalPrometheusRulesMap` or `serverFiles.alerting_rules.yml`. | | `helm-overrides//kube-prometheus-stack/custom-values.yaml` | Prometheus | Same as above; older clusters use this chart name. | | `helm-overrides//mimir/custom-values.yaml` | Mimir ruler | Rules go through the ruler API; YAML in `runtimeConfig` or via a sidecar `ConfigMap`. | Confirm which cluster runs which engine before editing — sample existing rules in the override. --- ## When to use - Adding/removing one or more alert rules in the cluster's primary alerting engine. - Tuning `expr:` PromQL or `for:` windows on existing rules. - Re-routing alerts (changing `severity`, `team`, or other routing labels). - Adding/removing recording rules for downstream alert efficiency. Do **not** use this procedure for: - Touching scrape configs, retention, datasources, or sizing in the same PR — split it. See [modify-observability-config.md](modify-observability-config.md). - Bumping the alerting chart's version — see [update-chart-version.md](update-chart-version.md). - Writing app-specific alert rules that belong in the consuming service's own values — those go in the service's repo. --- ## Inputs | Input | Example | |-------|---------| | Target cluster | `k8s-supply-prd-ase1` | | Target chart | `vmalert` / `prometheus` / `kube-prometheus-stack` / `mimir` | | Rule name(s) added/changed | `HighErrorRate`, `KafkaConsumerLag` | | `expr:` PromQL | `sum(rate(http_requests_total{status=~"5.."}[5m])) by (service) > 0.05` | | `for:` window | `5m` | | Severity / routing labels | `severity: critical`, `team: supply-platform` | | Approval ticket | CMR-… | --- ## Pre-conditions - [ ] You have `promtool` and/or `vmalert` available locally. - [ ] You know the cluster's Alertmanager / vmalert notifier routing (which `severity` / `team` label routes where). - [ ] You have run the new PromQL expression against the cluster's VM/Prometheus to confirm it returns sensible values *before* writing the rule — open Grafana Explore. - [ ] You have on-call sign-off if this rule pages on-call. --- ## Steps ### 1. Locate the rules block ```bash yq e '.config.rules // .additionalPrometheusRulesMap // .serverFiles' \ helm-overrides///custom-values.yaml ``` Identify the YAML path the engine expects: - **vmalert**: `config.groups[].rules[]`. - **kube-prometheus-stack**: `additionalPrometheusRulesMap..groups[].rules[]`. - **prometheus**: `serverFiles.alerting_rules.yml.groups[].rules[]`. - **mimir ruler**: per-tenant config — confirm the cluster's tenancy setup. ### 2. Author the rule ```yaml - alert: # PascalCase, no spaces; surfaces in pages and dashboards expr: | for: # e.g. 5m, 10m labels: severity: team: # routes to the right PagerDuty service annotations: summary: "" description: "" runbook_url: ``` #### Threshold discipline - Express thresholds as `rate(...) > 0.05` rather than `> 5%`. Ensure units in the expression match the units in `summary` / `description`. - Avoid `for:` windows shorter than the scrape interval × 2. Sub-`1m` windows are flap factories. - Avoid alerting on absolute counts (`sum(http_requests_total) > 1000`); prefer rates / ratios. - For SLO-style alerts, use multi-window multi-burn-rate (Google SRE workbook). Single-window threshold alerts are a known anti-pattern. #### Routing implications Changing `severity` from `warning` to `critical` (or vice versa) **silently re-routes the page** — the rule may now wake on-call where it previously emailed (or vice versa). Same for `team:` — a typo here drops alerts into the wrong PagerDuty service. Cross-check the cluster's Alertmanager / vmalert routing config: ```bash yq e '.config.route // .alertmanagerConfig' \ helm-overrides///custom-values.yaml ``` ### 3. Render the chart ```bash helm template helm-templates/ \ -f helm-overrides///custom-values.yaml > /tmp/rendered.yaml ``` ### 4. Validate PromQL For Prometheus / `kube-prometheus-stack`: ```bash yq e 'select(.kind == "PrometheusRule")' /tmp/rendered.yaml > /tmp/rules.yaml promtool check rules /tmp/rules.yaml ``` For vmalert: ```bash yq e 'select(.kind == "VMRule" or .kind == "ConfigMap")' /tmp/rendered.yaml > /tmp/vmrules.yaml vmalert -dryRun -rule=/tmp/vmrules.yaml ``` For Mimir ruler — `mimirtool rules check` against the rendered ruler config. If validation fails: fix the PromQL. **Do not** merge a rule that fails parse. ### 5. Sanity-check expressions in Grafana Explore In a non-production manner: open Grafana on the cluster, paste the `expr:` into Explore, run over the last 6h. Confirm: - The series returned makes sense for the rule. - The threshold isn't always-firing or never-firing in normal operation. - Empty result is allowed (`absent_over_time(...)` style). ### 6. Open the PR ```bash git checkout -b alerts/- git add helm-overrides///custom-values.yaml git commit -m "alerts(/): " git push origin alerts/- gh pr create --base main ``` ### PR description template ```markdown ## Summary alert rule(s) on `` × ``. ## Rules touched | Rule | Change | New expr | New `for:` | Severity | Team | |------|--------|----------|-----------|----------|------| | `` | added | `` | `5m` | `critical` | `supply-platform` | ## Validation - [ ] `promtool check rules` / `vmalert -dryRun` passed - [ ] PromQL sanity-checked in Grafana Explore over last 6h - [ ] Threshold appropriate (no flap risk, no always-firing) - [ ] Routing labels (`severity`, `team`) verified against Alertmanager/vmalert routes - [ ] Runbook URL added (`annotations.runbook_url`) ## Approvers - Observability owner: - On-call lead (severity=): ``` ### 7. After merge — Sync and observe Sync the Argo Application. Watch: ```bash # vmalert kubectl --context= -n monitoring logs deploy/vmalert | grep -i "\|error" # kube-prometheus-stack kubectl --context= -n monitoring logs prometheus-k8s-0 | grep -i "\|error" ``` Open the alerting engine's UI (vmalert UI / Prometheus `/alerts`) to confirm the rule loaded. ### Watch for false positives Stay on the alert for at least one full `for:` window after sync. If the rule pages immediately and was not expected to → revert. --- ## Anti-patterns 1. **Sub-1m `for:` windows** — flap-prone. 2. **Alerting on absolute counts** instead of rates/ratios. 3. **`severity` flips without on-call sign-off** — silent re-route. 4. **Missing `runbook_url`** — pages without remediation steps waste on-call cycles. 5. **PromQL that uses `{job=~"foo.*"}` regex without anchors** — slow query, sometimes false matches. 6. **Inlining the rule body in the chart `templates/`** to "make it sticky" — that's a chart fork; see NEVER-DO. 7. **Adding the rule on every cluster in one PR** — surgical only. --- ## Rollback - Revert the PR. Sync. - vmalert/Prometheus reload picks up the previous ruleset within seconds. - Any in-flight alerts resolve on next evaluation. --- ## Related - Coding guideline: [../../global/coding-guidelines/observability.md](../../global/coding-guidelines/observability.md). - Procedure: [modify-observability-config.md](modify-observability-config.md) — broader observability edits. - Runbook: [../runbooks/metrics-gap.md](../runbooks/metrics-gap.md) — when an alert is silent because the metric is gone. - Schema: [../schemas/custom-values-schema.md](../schemas/custom-values-schema.md). - Escalation: [../../global/escalation-matrix.md](../../global/escalation-matrix.md).