Provisioned rather than built by hand in Grafana's UI — same reasoning
as the datasource: survives a pod restart, and a git diff shows what
changed. Four rows: ingress (Envoy total RPS + connections + response
class breakdown), service level (CPU/memory by namespace, filterable
via a $namespace template variable, plus a current-usage table), cluster
utilization (used vs actual node capacity, not an assumed limit), and
total resources (cores/memory/pods/disk).
Two scrape gaps found and fixed to make this possible, both in vmagent:
- Contour's own ingress Envoy (projectcontour namespace — the actual
data plane for everything routed through this homelab, hostPort
80/443) was not being scraped at all. Confirmed live: Cilium's
separate embedded Envoy (kube-system, its own L7 policy proxy) was
already flowing, via the annotation-based kubernetes-pods job — which
is what first showed envoy_* metrics existed in this cluster at all —
but Contour's Envoy carries no such annotation. Added an explicit job
targeting the projectcontour namespace by container port (8002, the
official chart's fixed Envoy metrics port) rather than guessing at
pod labels this cluster's auto-detected object names may not match.
- node-exporter, deployed two commits ago, was never actually being
scraped either: confirmed live that kubernetes-service-endpoints
(role: endpointslice, keyed on the Service's scrape annotation — where
that chart puts it) finds nothing in this cluster at all, not merely
down. Rather than chase why, added the same fix as Envoy: target the
pod directly by its declared container port (9100).
Verified against the live deployment (queried through vmui) before
writing a single panel: envoy_http_downstream_rq_total,
envoy_http_downstream_rq_xx, container_cpu_usage_seconds_total,
container_memory_working_set_bytes, machine_cpu_cores and
machine_memory_bytes all confirmed present with real data. The one
exception is the "Disk free" panel, which depends on the node-exporter
scrape fix landing in this same change — noted in the values file's own
comment as unverified until it actually deploys.
Also verified with `helm template`: the dashboard JSON round-trips
through the YAML values file and the chart's own ConfigMap templating
intact (19 panels both times), and vmagent's scrape_configs list still
carries all 8 chart defaults plus both new jobs — nothing lost by using
extraScrapeConfigs instead of overriding the full list by hand.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wajog7nELA3i8JWTjxYGHF
Lower RAM/disk for the same metric volume on an 8GB single node — VM's
own compression is the entire reason it exists as a project. Speaks
Prometheus's own query API (/api/v1/query), so nothing downstream
needed to change beyond the URL it points at: same PromQL, same
kubernetes-nodes-cadvisor-sourced container_cpu_usage_seconds_total /
container_memory_working_set_bytes toolshed's metrics work
(docs/PRODUCT-ARCHITECTURE.md step 5) already targets.
Three releases, matching the one-release-per-component convention
already used everywhere in this repo, since victoria-metrics-single has
no bundled scraper/exporter subcharts the way the Prometheus chart did:
- victoria-metrics-single: the TSDB + query engine. 3Gi/local-path,
7-day retention, resources capped at 512Mi — same trim reasoning as
the Prometheus server it replaces.
- vmagent: the scraper. Its default scrape_configs already includes
kubernetes-nodes-cadvisor and kubernetes-service-endpoints (the
chart's own comment says "COPY from Prometheus helm chart") — nothing
to override there, only remoteWrite pointed at victoria-metrics-single
and trimmed resources.
- node-exporter: pulled out of the removed Prometheus chart's bundled
subchart into its own standalone release, since victoria-metrics-single
has no equivalent. Same trim as before (32Mi/64Mi), same
prometheus.io/scrape annotation vmagent's default scrape config
already looks for.
Found and fixed while vendoring: helm-templates/victoria-metrics-single
already existed in this repo — a leftover GKE-targeted vendored copy
from the original Meesho monorepo import (commit b8575bb), with its own
templates/values.yaml for a different chart entirely. `mkdir -p` on an
already-existing directory silently did nothing, so the first vendoring
attempt left that old tree sitting alongside the new one and rendered
two ServiceAccounts/Services/StatefulSets with colliding names. Removed
outright rather than adapted, same reasoning as the ~35 charts
devops-helm-charts removed in its own prune — see this chart's own
Chart.yaml for the full note.
Verified with `helm template` against the real charts and these values
for all three releases individually, and again for the whole
generic-argo-apps-chart appSpec list (11 Applications render, including
the three new ones and nothing orphaned from the old prometheus entry).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wajog7nELA3i8JWTjxYGHF