Files
devops-infra-helm-charts-gcp/helm-overrides/gke-toolshed-prd-usc1/argocd-admin-prd/custom-values.yaml
T
Mukul SharmaandClaude Opus 5 61bc4af1a0 Serve every tool on deployshed.com instead of nip.io
Harbor, Gitea, Argo CD, Jenkins, Vault, Grafana and vmui now answer on
their deployshed.com names alone. Each was already serving both while the
move was proved out; this removes the nip.io half.

The dual-hostname workarounds go with it. Jenkins' secondaryingress existed
only because its chart's primary ingress takes one hostName and a
certificate could not span both names — the real domain moves onto the
primary with jenkins-tls, which it already holds. Argo CD gets extraTls
rather than ingress.tls, because the boolean hardcodes secretName
argocd-server-tls and would request a second certificate for a name that
already has a valid one in argocd-deployshed-tls.

Harbor also changes in two ways beyond the hostname:

  - externalURL moves to https://harbor.infra.deployshed.com. Harbor hands
    this to docker clients in its own API responses and builds the push
    commands shown in its UI from it, so a stale value is what makes a
    correctly-configured registry still advertise the old address.

  - updateStrategy is now Recreate. Its jobservice and registry volumes are
    standard-rwo (ReadWriteOnce), and a RollingUpdate starts the new pod
    before the old one releases the disk, so the replacement hangs forever
    on Multi-Attach. The cluster was sitting in exactly that state, old pods
    serving while new ones stayed in ContainerCreating. The chart's own
    comment on this value recommends Recreate when RWM is unavailable. The
    cost is a brief outage during upgrades, which beats a rollout that
    cannot complete.

The private registry CA is not removed yet. Apps deployed before this move
recorded nip.io image references that only change when each is rebuilt, so
the old hostname stays served by a standalone Ingress until then.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LEsTefWWifp4ikvhHF5s6N
2026-09-17 09:31:26 +05:30

162 lines
6.5 KiB
YAML

argo-cd:
# GKE counterpart of helm-overrides/k8s-admin-prd-ase1/argocd-admin-prd.
#
# Installed once by hand with `helm install argocd-admin-prd` (namespace
# "argocd"), then manages itself through the argocd Application in
# devops-infra-argo-config-gcp, whose nameOverride matches that release.
# Deliberately no image tag pin, unlike the homelab: the chart's own
# appVersion (v3.5.2) governs, so the image cannot drift from the chart.
# A pin that outlives its chart is close to the failure this upgrade
# fixes — software older than the cluster it manages.
#
# Upgraded from chart 7.7.23 / Argo CD v2.13.8. Three v3 behaviour changes
# apply to this deployment, none of which needs a values change today:
# - logs RBAC is now enforced, so an account that reads pod logs needs
# an explicit `logs, get` policy. jenkins-ci below only syncs.
# - update/delete no longer inherit to an application's sub-resources.
# - resource tracking moves from labels to annotations, so the first
# sync after the upgrade re-stamps every managed resource.
# SSO still deferred, same as the homelab.
dex:
enabled: false
controller:
replicas: 1
resources:
requests:
cpu: 200m
memory: 400Mi
limits:
cpu: 500m
memory: 768Mi
redis-ha:
enabled: false
redis:
resources:
requests:
cpu: 50m
memory: 64Mi
limits:
memory: 128Mi
# repo-server does the manifest rendering (helm template per Application),
# so it is the component that actually saturates when many apps sync at
# once. Stateless, safe to run several behind its Service. With
# autoscaling on, the chart omits `replicas` from the Deployment, so the
# HPA and ArgoCD's own self-management do not fight over the count.
#
# CPU only. The chart's default also scales on memory, but a Go process
# does not hand memory back promptly after a spike, so a memory target
# scales up and then never scales down. Setting it to null removes it.
repoServer:
autoscaling:
enabled: true
minReplicas: 1
maxReplicas: 3
targetCPUUtilizationPercentage: 70
targetMemoryUtilizationPercentage: null
resources:
requests:
cpu: 100m
memory: 256Mi
limits:
cpu: 300m
memory: 512Mi
# API/UI. Stateless, sessions live in Redis, so replicas are
# interchangeable. Same CPU-only reasoning as repoServer above.
server:
autoscaling:
enabled: true
minReplicas: 1
maxReplicas: 3
targetCPUUtilizationPercentage: 70
targetMemoryUtilizationPercentage: null
# --insecure is NOT set here as an extra arg: configs.params below
# carries server.insecure, which is the supported way to express it and
# is what the chart renders into argocd-cmd-params-cm. Setting both
# works but leaves two places to disagree.
ingress:
enabled: true
ingressClassName: contour
hostname: "argocd.infra.deployshed.com"
# TLS via extraTls rather than the `tls: true` boolean, deliberately.
#
# The boolean hardcodes `secretName: argocd-server-tls` (see the
# chart's argocd-server/ingress.yaml). This deployment already holds a
# valid, issued certificate for this exact hostname in
# argocd-deployshed-tls, created by the standalone Ingress that served
# the real domain while nip.io was still on `hostname`. Flipping the
# boolean would ignore that and request a second certificate for the
# same name — a needless issuance and a gap while it is obtained.
#
# extraTls takes an explicit secretName, so the existing certificate is
# adopted as-is and the standalone Ingress can simply be deleted.
extraTls:
- hosts:
- argocd.infra.deployshed.com
secretName: argocd-deployshed-tls
resources:
requests:
cpu: 50m
memory: 128Mi
limits:
cpu: 200m
memory: 256Mi
# applicationSet.enabled no longer exists in this chart, and there is no
# replacement: unlike dex and notifications below, the ApplicationSet
# controller's Deployment has no conditional at all. replicas: 0 is the
# only lever — the Deployment exists but runs nothing. Carrying the old
# `enabled: false` forward would have quietly started the controller,
# since Helm ignores unknown keys.
#
# Nothing here uses the ApplicationSet CRD; Applications are rendered by
# generic-argo-apps-chart instead.
applicationSet:
replicas: 0
notifications:
enabled: false
configs:
# Contour terminates TLS in front of Argo CD; leaving Argo CD's own TLS
# on as well produces a redirect loop. This renders into
# argocd-cmd-params-cm, which the server actually reads.
#
# This file previously expressed it as server.extraArgs: [--insecure],
# inherited from the homelab. Both work, but only one should exist, and
# the rendered ConfigMap is the thing to check when it looks wrong.
params:
server.insecure: true
cm:
# https now that the only hostname served carries a real certificate.
# This is what ArgoCD builds its own links from, so leaving it http
# would hand out plain-HTTP URLs for a TLS-only deployment.
url: "https://argocd.infra.deployshed.com"
timeout.reconciliation: 3m
timeout.reconciliation.jitter: 60s
# No Ingress health override, unlike the homelab: there Contour sat
# behind hostPort, so nothing ever wrote an Ingress's load balancer
# status. On GKE Envoy gets a real LoadBalancer Service and Contour
# writes that status, so ArgoCD's built-in check works as intended.
#
# Scoped account for Jenkins' syncArgoApp step — same as the homelab.
accounts.jenkins-ci: apiKey
accounts.jenkins-ci.enabled: "true"
rbac:
policy.csv: |
p, jenkins-ci, applications, sync, webapp/*, allow
p, jenkins-ci, applications, get, webapp/*, allow
# Reached over cluster DNS, never through Contour — which is what lets
# Contour itself be ArgoCD-managed. The repos are private on this
# public-facing Gitea, so ArgoCD reads them with a repo-creds Secret
# (created at bootstrap, covering everything under gitadmin/), not
# anonymously as in the homelab.
repositories:
devops-infra-helm-charts-gcp:
url: http://gitea-http.gitea.svc.cluster.local:3000/gitadmin/devops-infra-helm-charts-gcp.git
devops-infra-argo-config-gcp:
url: http://gitea-http.gitea.svc.cluster.local:3000/gitadmin/devops-infra-argo-config-gcp.git