Vendored prometheus-community/prometheus 29.27.1 the same way Contour/ArgoCD/Vault/Gitea/Harbor/Jenkins already are here — a thin Chart.yaml dependency plus a committed .tgz — rather than hand-written like postgresql, which has no official chart to vendor. Server only: alertmanager, kube-state-metrics, prometheus-node-exporter and prometheus-pushgateway are all enabled by default in this chart and all disabled here. None are needed for what actually consumes this — toolshed's per-app CPU/memory metrics read straight from the chart's built-in kubernetes-nodes-cadvisor scrape job (kubelet's own cAdvisor endpoint) — and each is its own pod on a node that was already at its 8GB ceiling before this. Trimmed for the same ceiling: 3Gi PVC on local-path (not the chart's 8Gi default), 7-day retention (not 15), resources capped at 512Mi. nameOverride pinned to exactly "prometheus" matters more here than for any other component pinning it: the chart's server Service renders as "<release-name>-server", so this is what makes it "prometheus-server" — the exact hostname toolshed's PROMETHEUS_URL was already seeded to point at, before this existed, so the connection is already correct the day this deploys. Verified with `helm template` against the real chart and these values: server-only object set (ClusterRole, ClusterRoleBinding, ConfigMap, Deployment, PVC, Service, ServiceAccount — nothing from the four disabled subcharts), and the rendered PVC/retention/resources/Service name all match what's written above. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Wajog7nELA3i8JWTjxYGHF
50 lines
2.0 KiB
YAML
50 lines
2.0 KiB
YAML
prometheus:
|
|
# Server only. alertmanager/kube-state-metrics/node-exporter/pushgateway
|
|
# are all enabled by default in this chart — none were asked for, and
|
|
# none are needed for what actually consumes this: toolshed's per-app
|
|
# CPU/memory come straight from kubelet's own cAdvisor endpoint (the
|
|
# kubernetes-nodes-cadvisor scrape job below, built into the server
|
|
# itself), not from any of these four. Each is its own pod on an 8GB
|
|
# node that was already at its ceiling before this — see claude.md's
|
|
# resource budget table.
|
|
alertmanager:
|
|
enabled: false
|
|
kube-state-metrics:
|
|
enabled: false
|
|
prometheus-node-exporter:
|
|
enabled: false
|
|
prometheus-pushgateway:
|
|
enabled: false
|
|
|
|
server:
|
|
persistentVolume:
|
|
# local-path-provisioner, this cluster's default StorageClass —
|
|
# installed right after Cilium precisely because kubeadm ships no
|
|
# default (unlike k3s). 3Gi rather than the chart's 8Gi default:
|
|
# `retention: 7d` below on one small cluster's worth of series
|
|
# comfortably fits, and the node has no room to spare. Not
|
|
# resizable in place with this provisioner, so sized deliberately
|
|
# rather than grown later.
|
|
size: 3Gi
|
|
storageClass: local-path
|
|
|
|
# 7 days, not the chart's 15-day default — a homelab whose entire
|
|
# purpose is proving a CI/CD pipeline has no use for a month of
|
|
# historical series, and every extra day is disk this node does not
|
|
# have spare.
|
|
retention: 7d
|
|
|
|
resources:
|
|
requests:
|
|
cpu: 50m
|
|
memory: 128Mi
|
|
limits:
|
|
memory: 512Mi
|
|
|
|
# The default kubernetes-nodes-cadvisor job (scheme https, bearer token
|
|
# from the pod's own ServiceAccount, metrics_path /metrics/cadvisor) is
|
|
# left exactly as the chart ships it — this is what
|
|
# container_cpu_usage_seconds_total and container_memory_working_set_bytes
|
|
# come from, and toolshed's metrics connection (internal/metrics)
|
|
# queries exactly those two, summed by namespace.
|