12 KiB
Schema — custom-values.yaml (Helm override)
Field-by-field annotation of the values file Argo CD's
valueFilesreferences. Authoritative forhelm-overrides/<cluster>/<app>/custom-values.yaml.Cross-reference: coding-guidelines/helm-values.md, ../../architecture.md.
What this file is — and isn't
A custom-values.yaml is not a Helm chart. It is the override layer Argo CD merges over the chart's own values.yaml (or, for thin-wrapper charts, the upstream subchart's defaults).
- Schema is owned by the chart, not by us. Every key here must be a key the chart understands. Adding a key the chart doesn't read does nothing.
- No multi-doc YAML. One document per file.
- No
apiVersion/kind. This file is values, not a Kubernetes manifest.
Top-level shape (typical)
# image registry / repository / tag
image:
registry: asia-southeast1-docker.pkg.dev
repository: meesho-devops-admin-0622/admin/sre/<image>
tag: <semver-or-sha>
pullPolicy: IfNotPresent
# release-name pinning (rarely changed once set)
fullnameOverride: <release>-<bu>-prd # e.g. kube-state-metrics-dbc-dsci-prd
# scaling
replicaCount: 3 # static, OR
autoscaling:
enabled: true
minReplicas: 3
maxReplicas: 10
targetCPUUtilizationPercentage: 70
# scheduling — bespoke per cluster
nodeSelector:
dedicated: <pool-key> # standard GKE
# cloud.google.com/compute-class: <cc> # GKE Autopilot
tolerations:
- key: dedicated
value: <pool-key>
effect: NoSchedule
# resources
resources:
requests:
cpu: 250m
memory: 512Mi
limits:
cpu: 1
memory: 2Gi
# persistence (stateful charts only)
persistence:
enabled: true
storageClass: sc-pd-ssd # MUST exist in manifests/storageclass/
size: 100Gi
accessModes: [ReadWriteOnce]
# secrets — by reference, never inline
existingSecret: <name-managed-by-external-secrets>
# probes
livenessProbe:
initialDelaySeconds: 30
periodSeconds: 10
readinessProbe:
initialDelaySeconds: 5
periodSeconds: 5
# ingress / Contour HTTPProxy (if applicable)
ingress:
enabled: false # most internal services
# chart-specific top-level keys
# (varies — read the chart's values.yaml for the full schema)
§image
| Field | Required | Convention | Notes |
|---|---|---|---|
image.registry |
yes (production) | asia-southeast1-docker.pkg.dev |
Meesho's Artifact Registry mirror. (SANCTITY_RULES R11) |
image.repository |
yes | meesho-devops-admin-0622/admin/sre/<image> |
Mirror path. |
image.tag |
yes (production) | semver, datestamp, or git SHA | Never latest. Never unpinned. |
image.pullPolicy |
no | IfNotPresent |
Always only for development; pulls slow rollout. |
image.pullSecrets |
no | omit | Mirror is public to fleet; pull secrets are a lock-in trap. |
Note: chart authors disagree on the shape — some use image: {repository, tag} (no registry), some use a single image: <fully-qualified>, some split into image.repository: <registry>/<repo>. Read the chart's own values.yaml before structuring this block.
§release identity
| Field | Required | Convention | Notes |
|---|---|---|---|
fullnameOverride |
varies | Set once at release creation; stable forever | Service DNS, PVC binding, ConfigMap refs depend on it. (SANCTITY_RULES R9) |
nameOverride |
rare | Omit unless explicitly needed | |
commonLabels / commonAnnotations |
rare | Omit unless the chart documents the pattern |
For dataplane (db-*) clusters: fullnameOverride: <kind>-dbc-<bu>-prd is the convention (e.g. kube-state-metrics-dbc-dsci-prd). For BU clusters, omit unless the chart's default name collides.
§scaling
Charts diverge sharply here. Common shapes:
# Pattern A — static replicas
replicaCount: 3
# Pattern B — chart-level autoscaling block
autoscaling:
enabled: true
minReplicas: 3
maxReplicas: 10
targetCPUUtilizationPercentage: 70
# Pattern C — separate HPA resource (KEDA / VictoriaMetrics)
hpa:
enabled: true
min: 3
max: 20
metrics:
- type: Resource
resource: {name: cpu, target: {type: Utilization, averageUtilization: 70}}
| Rule | Why |
|---|---|
Don't set replicaCount when autoscaling.enabled: true. |
HPA and the chart's static deployment fight; replica count flaps. |
Set both minReplicas and maxReplicas when autoscaling. |
Without min, the HPA can scale to zero on a quiet hour. |
minReplicas ≥ 2 for any production-traffic-path workload. |
Single-replica releases die on node drain. |
§scheduling
The single biggest source of silent mis-deploys. Per SANCTITY_RULES R5:
| Cluster | Key style |
|---|---|
GKE Autopilot (k8s-central-prd-ase1, k8s-dsgpu-prd-ase1, k8s-shared-int-ase1) |
cloud.google.com/compute-class: |
| Standard GKE | dedicated: |
# Standard GKE — Contour internal-0 on most clusters
nodeSelector:
dedicated: contour-internal-0
tolerations:
- key: dedicated
value: contour-internal-0
effect: NoSchedule
# often a second toleration to allow scheduling onto shared nodes
- key: dedicated
value: contour-shared
effect: NoSchedule
# GKE Autopilot — same Contour internal-0 on k8s-central-prd-ase1
nodeSelector:
cloud.google.com/compute-class: contour-internal-0-cc
tolerations:
- key: cloud.google.com/compute-class
value: contour-internal-0-cc
effect: NoSchedule
- key: cloud.google.com/compute-class
value: contour-shared-cc
effect: NoSchedule
The full per-cluster Contour matrix lives in contour-nodeselector-tolerations-summary.md. For non-Contour apps, copy from a sibling app on the same cluster.
| Field | Notes |
|---|---|
nodeSelector |
Hard requirement — no scheduling onto non-matching nodes. Use exactly one key (the cluster's pool key). |
tolerations |
Allow scheduling onto tainted nodes. May list multiple to allow shared nodes alongside dedicated. |
affinity.nodeAffinity |
Use sparingly; nodeSelector is sufficient for most cases here. |
topologySpreadConstraints |
Use for multi-zone resilience on chatty workloads (Contour, ingress-nginx). |
§resources
resources:
requests:
cpu: 250m
memory: 512Mi
limits:
cpu: 1
memory: 2Gi
| Rule | Why |
|---|---|
Always set requests in production. |
Without requests, pods are best-effort; first to evict on memory pressure. |
Set limits unless the chart explicitly recommends omitting them. |
Some sidecars (cert-manager, external-dns) deliberately skip limits. |
| Don't copy resources from another cluster. | Workload sizing is per-traffic-tier. Supply prd is not demand prd. |
| CPU limits cause throttling, not OOMKill. Memory limits cause OOMKill. Tune accordingly. |
§persistence
persistence:
enabled: true
storageClass: sc-pd-ssd
size: 100Gi
accessModes: [ReadWriteOnce]
Allowed storageClass values (must match a file under manifests/storageclass/):
| Class | Backend | Use case |
|---|---|---|
pd-standard-retain-dr |
GCP Persistent Disk standard, retention-on-delete | DR-critical state |
sc-pd-ssd |
GCP PD SSD | Default for write-heavy workloads (etcd, ClickHouse) |
sc-pd-standard |
GCP PD standard | Default for read-heavy / archival |
sc-filestore-standard |
GCP Filestore | Shared volumes (Jenkins build cache, JFrog filestore) |
| Rule | Why |
|---|---|
Never reference a storageClass that doesn't exist in manifests/storageclass/. |
PVC stays Pending forever. |
size must be set explicitly. |
Chart defaults are often wrong (8Gi for everything). |
accessModes for ReadWriteMany requires Filestore. PD-backed classes are RWO. |
§secrets
# Chart-specific — read the chart's values.yaml
existingSecret: <name>
# OR
auth:
existingSecret: <name>
# OR
envFrom:
- secretRef: {name: <name>}
| Rule | Why |
|---|---|
Never inline password:, apiKey:, etc. |
Pre-commit TruffleHog will catch many; it won't catch all. (SANCTITY_RULES R4) |
The Secret resource is materialised by the cluster's external-secrets app. If the secret name doesn't appear in helm-overrides/<cluster>/external-secrets/, it doesn't exist. |
|
Reference, don't author. Adding the secret material to a Helm-rendered Secret template defeats the External Secrets pattern. |
§probes
livenessProbe:
httpGet: {path: /healthz, port: http}
initialDelaySeconds: 30
periodSeconds: 10
timeoutSeconds: 3
failureThreshold: 3
readinessProbe:
httpGet: {path: /ready, port: http}
initialDelaySeconds: 5
periodSeconds: 5
# For slow-warm workloads (Jenkins, JFrog, ClickHouse)
startupProbe:
httpGet: {path: /healthz, port: http}
initialDelaySeconds: 60
periodSeconds: 10
failureThreshold: 60 # 10 minutes total runway
| Rule | Why |
|---|---|
Always set liveness + readiness for any long-running container. |
|
Use startupProbe for slow-warm workloads instead of inflating livenessProbe.initialDelaySeconds to 600 s. |
|
Don't set the same probe to both liveness and readiness — they have different purposes (kill vs gate). |
§ingress (Contour HTTPProxy / Ingress)
For external-facing apps, ingress is via Contour HTTPProxy. Specifics depend on the cluster's Contour topology (multiple Contour instances per cluster — see contour-nodeselector-tolerations-summary.md).
- The
external-dnsannotation is the DNS-binding step — match the cluster's DNS pattern. - TLS terminates at Contour with
cert-manager-issued certs; reference the issuer by name. - Internal services use
contour-internal-0/contour-internal-1; external services usecontour-external/contour-external-1. The class is set via theServiceannotation, not the values file directly.
End-to-end example — a typical observability-side override
helm-overrides/k8s-supply-prd-ase1/kube-state-metrics/custom-values.yaml:
image:
registry: asia-southeast1-docker.pkg.dev
repository: meesho-devops-admin-0622/admin/sre/kube-state-metrics
tag: v2.10.1
replicaCount: 2
resources:
requests:
cpu: 100m
memory: 256Mi
limits:
cpu: 500m
memory: 1Gi
nodeSelector:
dedicated: monitoring
tolerations:
- key: dedicated
value: monitoring
effect: NoSchedule
prometheus:
monitor:
enabled: true
honorLabels: true
End-to-end example — a dataplane override
helm-overrides/db-2516183257845181-c-1204-195038-428/victoria-metrics-agent/custom-values.yaml:
fullnameOverride: vmagent-dbc-supply-prd
image:
registry: asia-southeast1-docker.pkg.dev
repository: meesho-devops-admin-0622/admin/sre/victoria-metrics-agent
tag: v1.96.0
replicaCount: 1
resources:
requests:
cpu: 100m
memory: 256Mi
(Dataplane clusters intentionally have minimal overrides.)
Validation
# Render the chart locally with this values file
helm template <release> helm-templates/<chart> \
-f helm-overrides/<cluster>/<app>/custom-values.yaml | head -80
# Dry-run diff against the live cluster (requires kubectl context + helm-diff plugin)
helm diff upgrade <release> helm-templates/<chart> \
-f helm-overrides/<cluster>/<app>/custom-values.yaml
# Lint
yamllint helm-overrides/<cluster>/<app>/custom-values.yaml
helm lint helm-templates/<chart>
# Sanity: storageClass referenced exists
yq e '.persistence.storageClass' helm-overrides/<cluster>/<app>/custom-values.yaml \
| xargs -I{} ls manifests/storageclass/{}.yaml 2>/dev/null \
|| echo "WARN: storageClass not in manifests/storageclass/"