Files
devops-infra-helm-charts-gcp/helm-overrides/k8s-admin-prd-ase1
Mukul SharmaandClaude Opus 5 0c68312765 Homelab Overview dashboard: envoy RPS, per-namespace CPU/mem, cluster utilization, totals
Provisioned rather than built by hand in Grafana's UI — same reasoning
as the datasource: survives a pod restart, and a git diff shows what
changed. Four rows: ingress (Envoy total RPS + connections + response
class breakdown), service level (CPU/memory by namespace, filterable
via a $namespace template variable, plus a current-usage table), cluster
utilization (used vs actual node capacity, not an assumed limit), and
total resources (cores/memory/pods/disk).

Two scrape gaps found and fixed to make this possible, both in vmagent:

- Contour's own ingress Envoy (projectcontour namespace — the actual
  data plane for everything routed through this homelab, hostPort
  80/443) was not being scraped at all. Confirmed live: Cilium's
  separate embedded Envoy (kube-system, its own L7 policy proxy) was
  already flowing, via the annotation-based kubernetes-pods job — which
  is what first showed envoy_* metrics existed in this cluster at all —
  but Contour's Envoy carries no such annotation. Added an explicit job
  targeting the projectcontour namespace by container port (8002, the
  official chart's fixed Envoy metrics port) rather than guessing at
  pod labels this cluster's auto-detected object names may not match.

- node-exporter, deployed two commits ago, was never actually being
  scraped either: confirmed live that kubernetes-service-endpoints
  (role: endpointslice, keyed on the Service's scrape annotation — where
  that chart puts it) finds nothing in this cluster at all, not merely
  down. Rather than chase why, added the same fix as Envoy: target the
  pod directly by its declared container port (9100).

Verified against the live deployment (queried through vmui) before
writing a single panel: envoy_http_downstream_rq_total,
envoy_http_downstream_rq_xx, container_cpu_usage_seconds_total,
container_memory_working_set_bytes, machine_cpu_cores and
machine_memory_bytes all confirmed present with real data. The one
exception is the "Disk free" panel, which depends on the node-exporter
scrape fix landing in this same change — noted in the values file's own
comment as unverified until it actually deploys.

Also verified with `helm template`: the dashboard JSON round-trips
through the YAML values file and the chart's own ConfigMap templating
intact (19 panels both times), and vmagent's scrape_configs list still
carries all 8 chart defaults plus both new jobs — nothing lost by using
extraScrapeConfigs instead of overriding the full list by hand.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wajog7nELA3i8JWTjxYGHF
2026-09-06 09:45:46 +05:30
..
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-31 08:05:13 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-31 07:55:46 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-31 13:17:36 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-31 10:46:04 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30
2026-08-26 03:39:42 +05:30

Cluster-based Custom Values

This folder contains the custom values.yaml files organized based on specific cluster names. Each subdirectory corresponds to a particular cluster and holds the configurations for the applications and tools deployed within that cluster.