Mukul SharmaandClaude Opus 5 adb514c3be Issue public certificates for the real domain
The deployment has a domain of its own now, which is the first time a
public certificate has been possible here at all. The registry issuer
beside this one explains why: nip.io is not on the public suffix list and
every *.nip.io certificate shares one rate limit, so Let's Encrypt could
never serve the addresses this cluster has been using.

DNS-01 rather than HTTP-01, because every deployed app lives at
<app>.apps.<domain> and only a DNS-01 challenge can issue the wildcard
that covers all of them. HTTP-01 would mean a certificate per app,
requested the moment each one is created.

Two issuers, staging and production. Production allows five duplicate
certificates a week and a failing solver spends that allowance without
issuing anything, which can lock a domain out of certificates for days.
Staging has no meaningful limit and issues from an untrusted root, so a
browser warning is the signal that the plumbing works. Move to the
production issuer once a staging certificate appears.

The token reaches cert-manager the same way every other credential here
does: Vault, through External Secrets. It wants Zone -> DNS -> Edit on the
one zone and nothing else — enough to write the TXT record a challenge
needs, and no more. Until `vault kv put secret/cloudflare/dns-token` has
run, the ExternalSecret stays unfulfilled and the issuers cannot register.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LEsTefWWifp4ikvhHF5s6N
2026-09-17 01:24:54 +05:30
2026-08-26 04:03:34 +05:30

devops-infra-argo-config

GitOps control plane for infrastructure tooling across Meesho's Kubernetes fleet.

This repo manages ArgoCD Application resources for every infrastructure tool (Contour, VictoriaMetrics, Grafana, Kyverno, KEDA, external-secrets, Vault, etc.) deployed across ~19 clusters. It uses an App-of-Applications pattern: one parent Application per cluster renders child Applications from a appSpec[] list via a generic Helm chart.

Each environment tracks a dedicated branch — merging to that branch triggers immediate ArgoCD auto-sync with no staging gate:

Environment Branch
Production (prd) main
Staging (stg) develop
Integration (int) pre-prod

How it works

incubator/<env>/<cluster>.yaml          ← Parent Application (one per cluster)
    └── points at generic-argo-apps-chart/ + values/<env>/<cluster>-values.yaml
            └── renders one child Application per appSpec[] entry
                    └── sources charts + overrides from devops-infra-helm-charts

Directory structure

Directory Purpose
incubator/<env>/ Parent ArgoCD Application YAML, one per cluster
values/<env>/ Values files defining which tools deploy per cluster
generic-argo-apps-chart/ Helm chart that renders child Applications from appSpec[]
projects/ ArgoCD AppProject definitions (sre, sec)
external-name-service-*/ Cross-cluster DNS routing (ExternalName / MCS topology)
docs/ Agent-facing operational documentation
skills/ Parameterized agent tasks for common operations
wiki/ Architecture decisions and entity pages

Getting started

Common operations

Task Procedure
Add a tool to a cluster docs/platform/procedures/add-tool-to-cluster.md
Upgrade a chart version docs/platform/procedures/upgrade-chart-version.md
Onboard a new cluster docs/platform/procedures/add-new-cluster.md
Roll out a tool fleet-wide docs/platform/procedures/fleet-wide-tool-rollout.md
Debug sync failure docs/platform/runbooks/argocd-sync-failure.md
Debug Helm render error docs/platform/runbooks/render-failure.md
Find values inconsistencies across clusters docs/platform/runbooks/values-drift.md
Debug stuck deployment docs/platform/runbooks/deployment-stuck.md

Sister repos

  • devops-infra-helm-charts — Helm charts and custom-values.yaml overrides. Every appSpec[].chartDir and valuesDir must exist here.
  • devops-argo-config — Same pattern for service/application workloads (not infra tooling).
S
Description
No description provided
Readme
223 KiB