10 KiB
Per AI Blitz Plan §platform.runbooks. Layer: 1. Repo: devops-infra-helm-charts.
Runbook — Metrics gap on a cluster / namespace
Type: Decision tree. Entry symptom: "Grafana panels are blank for
<cluster>/<namespace>/<service>," or an alert that should be firing isn't, or VictoriaMetrics showsno data. Layer: mostly Layer 2 (advisory). Layer 1 only when the fix is a values diff in this repo.
The observability path on Meesho's GKE fleet is:
workload Pod (exposes /metrics)
└→ victoria-metrics-agent (vmagent) scrapes
└→ remote_write to victoria-metrics-cluster (vmstorage)
└→ vmselect ← Grafana / vmalert query
A metrics gap can be at any hop. Walk this tree top-down — the most common root cause is hop 1 (scrape config).
Entry — gather context
CLUSTER=<cluster>
CTX=<kubectl-context>
NS=<workload-namespace-where-metric-is-missing>
METRIC=<metric-name-or-job-label>
# Confirm the cluster runs VM agent + cluster
ls helm-overrides/$CLUSTER | grep -E '^victoria-metrics-(agent|cluster)'
# Confirm pods are healthy
kubectl --context=$CTX -n monitoring get pods -l app.kubernetes.io/name=victoria-metrics-agent -o wide
kubectl --context=$CTX -n monitoring get pods -l app.kubernetes.io/name=vmstorage -o wide
If pods are missing/CrashLooping → that's the gap. Skip to §5.
Decision tree
START
│
├── Is the metric known to be emitted by the workload?
│ ├── No → §0 — Workload not emitting; out of scope (app team)
│ └── Yes →
│
├── §1 — Is vmagent scraping the workload?
│ ├── No → fix scrape config (Layer 1)
│ └── Yes →
│
├── §2 — Are scrape targets healthy (status=up)?
│ ├── Down → fix endpoint reachability (Layer 2 / Layer 1)
│ └── Up →
│
├── §3 — Are relabel rules dropping the metric?
│ ├── Yes → adjust relabel_configs (Layer 1)
│ └── No →
│
├── §4 — Is remote_write succeeding?
│ ├── No → vmagent → vmstorage path broken (Layer 2 / Layer 1)
│ └── Yes →
│
├── §5 — Is vmstorage healthy and ingesting?
│ ├── No → vmstorage outage (Layer 2)
│ └── Yes →
│
└── §6 — Is the Grafana datasource / tenant correct?
└── Misrouted query → fix datasource URL or tenant header (Layer 1)
§0 — Workload not emitting
Out of scope for this repo. Confirm with:
# Port-forward and curl /metrics directly
kubectl --context=$CTX -n $NS port-forward <pod> 9090:<metrics-port> &
curl -s localhost:9090/metrics | grep -i "$METRIC"
If /metrics is empty or doesn't contain $METRIC → app team. Stop.
§1 — Is vmagent scraping the workload?
# vmagent UI exposes /api/v1/targets
kubectl --context=$CTX -n monitoring port-forward svc/victoria-metrics-agent 8429:8429 &
curl -s localhost:8429/api/v1/targets | jq '.data.activeTargets[] | select(.labels.namespace=="'$NS'")'
If no targets for $NS:
- Look at
helm-overrides/$CLUSTER/victoria-metrics-agent/custom-values.yaml. - Check the
additionalScrapeConfigs(orconfig.scrape_configs) for a job that matches the workload's labels / namespace selector. - Common cause: a
kubernetes_sd_configsnamespaces.namesfilter excludes$NS. - Common cause: a missing
Pod/Service/PodMonitorannotationprometheus.io/scrape: "true"on the workload.
Layer 1 fix (if the gap is in the scrape config): edit helm-overrides/$CLUSTER/victoria-metrics-agent/custom-values.yaml per ../procedures/modify-observability-config.md.
Layer 2 fix (if the gap is on the workload — missing annotation): hand off to app team.
§2 — Targets healthy?
# In the same vmagent /targets output
curl -s localhost:8429/api/v1/targets | jq '.data.activeTargets[] | select(.health!="up") | {labels, lastError}'
If targets are down:
lastError |
Cause | Fix |
|---|---|---|
connection refused |
Workload not listening on declared port | App team — Layer 2. |
i/o timeout |
NetworkPolicy / firewall blocking vmagent → workload | Check NetworkPolicy in $NS. Often a NetworkPolicy allowing only intra-namespace traffic and not vmagent's namespace. Layer 2 — recommend the workload team allow vmagent. |
x509: certificate signed by unknown authority |
mTLS misconfigured | App team — Layer 2. |
404 Not Found |
Wrong path (default /metrics vs custom) |
Add metrics_path: in scrape config. Layer 1. |
§3 — Relabel rules dropping the metric?
yq e '.config.scrape_configs[].metric_relabel_configs, .config.scrape_configs[].relabel_configs' \
helm-overrides/$CLUSTER/victoria-metrics-agent/custom-values.yaml
Look for:
action: droprules with regexes matching$METRIC.action: keeprules whose regex excludes$METRIC.action: labeldropremoving a label the query uses.
Test in vmagent's UI under the Targets tab — it shows the labels post-relabel.
Layer 1 fix: loosen the relabel rule. Re-PR per ../procedures/modify-observability-config.md.
§4 — Remote_write succeeding?
kubectl --context=$CTX -n monitoring logs deploy/victoria-metrics-agent | grep -i 'remote_write\|error\|failed' | tail -20
| Symptom | Cause | Fix |
|---|---|---|
429 Too Many Requests from vmstorage |
vmstorage ingest saturated | Scale vmstorage / vminsert (Layer 1 — see observability.md). |
connection refused to vmstorage URL |
vmstorage Service down or wrong URL | Verify remoteWrite.url in vmagent values matches the live vmstorage Service DNS. Layer 1. |
out of bounds timestamp |
Clock skew on vmagent's node | Layer 2 — node time sync. |
series limit exceeded |
Cardinality bomb on vmstorage tenant | Layer 1 — drop the offending label. See observability.md §Cardinality. |
§5 — vmstorage healthy?
kubectl --context=$CTX -n monitoring get pods -l app.kubernetes.io/name=vmstorage -o wide
kubectl --context=$CTX -n monitoring describe pod vmstorage-0 | tail -30
kubectl --context=$CTX -n monitoring exec vmstorage-0 -- df -h /storage
Failure modes:
| Symptom | Layer | Fix |
|---|---|---|
Pod CrashLooping with out of disk |
Layer 1 | Bump persistence.size. Note: PVC growth requires the StorageClass to support allowVolumeExpansion: true. See ../schemas/storageclass-priorityclass-schema.md. |
| Pod CrashLooping with retention/index errors | Layer 2 | Escalate — may need data-side intervention. |
| Pod Pending | Layer 1 | Scheduling issue. See pod-pending-scheduling.md and ../../../wiki/analyses/ADR-A3-per-cluster-scheduling.md. |
| Pod Running but vmselect can't reach it | Layer 1 / 2 | Verify the headless Service and StatefulSet pod-DNS records. |
§6 — Grafana datasource / tenant correct?
yq e '.datasources.datasources.yaml.datasources[] | select(.name == "*VictoriaMetrics*" or .type == "prometheus")' \
helm-overrides/$CLUSTER/grafana/custom-values.yaml
| Sub-check | Action |
|---|---|
url: points at the correct in-cluster vmselect Service DNS |
If wrong, Layer 1 — fix the values per ../procedures/modify-observability-config.md. |
httpHeaderName1: X-Scope-OrgID (multi-tenant clusters only) |
If the cluster runs multi-tenant VM, the tenant header must be set. Layer 1 fix. |
Datasource url: points at an external *.meeshogcp.in host |
Forbidden — see ../../global/SANCTITY_RULES.md R3. Repoint at in-cluster Service DNS. |
Remediation summary
| Hop | Likely fix | Layer | Procedure |
|---|---|---|---|
| §1 scrape | Add scrape config / fix selector | 1 | ../procedures/modify-observability-config.md |
| §2 endpoint | Workload-side / NetworkPolicy | 2 | Hand off to app team |
| §3 relabel | Loosen drop rule | 1 | ../procedures/modify-observability-config.md |
| §4 remote_write | Scale vmstorage / fix URL | 1/2 | ../procedures/modify-observability-config.md |
| §5 vmstorage | PVC grow / scheduling fix | 1/2 | pod-pending-scheduling.md |
| §6 datasource | Fix Grafana datasource | 1 | ../procedures/modify-observability-config.md |
Escalation triggers
- §5 vmstorage outage with no obvious values-side fix → platform team (observability owner).
- Multi-cluster simultaneous metrics gap → platform team — cluster-level / control-plane issue.
- Cardinality explosion impacting vmstorage stability → platform team + workload owner — joint fix.
Done conditions
- The query that was returning
no datareturns the expected series in Grafana Explore. - No remote_write errors in vmagent logs for at least 5 minutes post-fix.
- Alerts that depend on the metric have transitioned from
pending/ silent back to expected state.
Related
- Procedure: ../procedures/modify-observability-config.md.
- Procedure: ../procedures/modify-alert-rules.md — if the gap is "alert silent" not "metric missing."
- Coding guideline: ../../global/coding-guidelines/observability.md.
- Runbook: pod-pending-scheduling.md — for §5 vmstorage scheduling failures.
- Runbook: argocd-sync-failure.md — if a sync didn't take.