Files
devops-infra-helm-charts-gcp/docs/platform/runbooks/argocd-sync-failure.md
T
2026-08-26 03:39:42 +05:30

10 KiB

Runbook — Argo CD Sync Failure (infra release)

Type: Decision tree. Entry symptom: an Argo CD Application for an infra release is OutOfSync, errored, or stuck Progressing. Layer: mostly Layer 1 (read state, propose YAML diff). Some branches are Layer 2 (advisory).

This runbook handles sync failures for infra Applications routed by github.com/Meesho/devops-infra-argo-config and rendering against this repo's helm-overrides/<cluster>/<app>/.


Entry — gather context

APP=<release>                                # e.g. argocd, contour-internal-0
CLUSTER=<cluster-name>                        # e.g. k8s-supply-prd-ase1
NS=$(argocd app get $APP -o json | jq -r '.spec.destination.namespace')
PROJECT=$(argocd app get $APP -o json | jq -r '.spec.project')

argocd app get $APP                           # headline
argocd app get $APP -o json | jq -r '.status.conditions[]?'
argocd app get $APP -o json | jq -r '.status.operationState.message // empty'

Note which Argo CD instance you're hitting — most infra Applications live in a per-cluster Argo CD install.


Decision tree

START
  │
  └── Is `argocd app get $APP` known to this Argo at all?
        │
        ├── NO → §1 — Application not found
        │
        └── YES → What's the symptom?
              │
              ├── Sync failed with an error message       → §2 — Errored sync
              ├── Sync stuck `Progressing` for >5 min     → §3 — Stuck progressing
              ├── App is `OutOfSync` but Sync hasn't run  → §4 — OutOfSync only
              └── `Synced`+`Healthy` but workload bad     → §5 — Wrong workload (leave runbook)

§1 — Application not found

Sub-check Action
Are you on the right Argo CD instance? Most v2 infra apps live in per-cluster Argo CDs.
Was the app deboarded recently? git -C <argo-config> log --diff-filter=D -- 'apps/<cluster>/<app>*' and git log --diff-filter=D -- 'helm-overrides/<cluster>/<app>/'.
Did the values directory land on main? git log --all -- 'helm-overrides/<cluster>/<app>/'.

If the file should exist on main but the Application resource isn't created → Layer 2 — escalate to the platform team. The cluster's Argo CD bootstrap (ApplicationSet) may not be picking up the path.


§2 — Errored sync (read the error message)

§2a — repository not accessible / authentication required

Action
Check spec.source.repoURL is github.com/Meesho/devops-infra-helm-charts.git.
If yes, the credentials in Argo CD's repo-list need refreshing. Layer 2 — recommend platform team rotates credentials.

§2b — path 'X' does not exist in repo Y

path 'helm-overrides/k8s-supply-prd-ase1/argocd' does not exist
Action
ls helm-overrides/<cluster>/<app>/ on main.
If absent → either the values-side PR wasn't merged, or the path was typo'd in the sister-repo Application. Layer 1 — open a fix PR (sister repo).
If a blue-green migration just landed: the Application may be pointing at the old chart path that was retired. Layer 1 — repoint the Application to the new sibling path.

§2c — Helm template error / values file not found

open helm-overrides/<cluster>/<app>/custom-values.yaml: no such file or directory
Action
Verify the file exists on main: git ls-tree origin/main -- helm-overrides/<cluster>/<app>/custom-values.yaml.
If absent → onboarding is incomplete. Layer 1 — open the missing values-side PR.

§2d — unable to render manifests / template error

Helm template error inside the chart (missing required value, type mismatch).

Action
Reproduce locally: helm template <app> helm-templates/<chart> -f helm-overrides/<cluster>/<app>/custom-values.yaml.
Determine whether the fix is in this repo (rare — usually values shape changed) or helm-templates/<chart> (more common after a chart bump). Layer 1.
If the error is Cannot use existing release: ... — see §2h.

§2e — cluster not found / dial tcp ... no route to host

Action
Layer 2 — escalate to platform team. Cluster API server unreachable, or cluster-secret stale in Argo CD.
Do not edit spec.destination.{server,name} to redirect; that masks the underlying cluster issue.

§2f — forbidden: ... / admission webhook deny

admission webhook "validate.kyverno.svc-fail" denied the request
Action
Look at the rule that denied (Kyverno? PSP? OPA? GKE Autopilot policy?).
Often the chart's manifest violates a cluster policy (e.g. runAsUser: 0, missing securityContext).
Layer 1 — fix in chart values; pair with the policy team if the policy is wrong.
GKE Autopilot specifically denies many privileged settings — read the deny message carefully.

§2g — webhook errored: ... cert-manager / external-secrets / kyverno

A webhook that should validate the new resource is itself unhealthy.

Action
kubectl get pods -n cert-manager (or the relevant operator's namespace) — is it running?
Layer 2 — recommend recovering the webhook before re-syncing this app.

§2h — cannot patch ... immutable field

Most often: Deployment.spec.selector or StatefulSet.volumeClaimTemplates. A chart bump that changes labels.

Action
Read the upstream changelog to confirm.
Layer 2 — recommend deleting the old Deployment / StatefulSet (with the workload owner) so the chart can recreate it. Do not delete blindly — for StatefulSet, the PVCs survive but the rollout is disruptive.
For systemic immutable-field changes across a chart bump, this is a sign the bump should have been a blue-green migration. Roll back, plan the migration.

§2i — dependent CRD ... not installed

The chart needs a CRD that doesn't exist yet on the cluster.

Action
Check whether the chart includes the CRD in templates/crds/ (most upstream charts ship CRDs).
If yes: the chart's helm template may not include CRDs by default — Argo CD has IncludeCRDs semantics; check the Application's helm.skipCrds setting.
If the CRD is supposed to come from a different chart (cert-manager, kube-prometheus-stack): Layer 2 — sync that chart first.

§3 — Stuck Progressing for > 5 minutes

The sync started but resources aren't reconciling.

Sub-check Action
argocd app get $APP --refresh shows resource-level status. Look for Progressing resources.
Is a Deployment failing to roll out? kubectl rollout status deploy/<name> -n $NS. If yes → see pod-pending-scheduling.md or ingress-down.md.
Is a Job hung? kubectl describe job/<name> -n $NS. Old Jobs sometimes block syncs (Helm pre-/post-install hooks).
Is a PreSync/PostSync hook hanging? kubectl get pods -n $NS -l argocd.argoproj.io/hook=PostSync.

If the workload itself is the problem, leave this runbook.


§4 — OutOfSync only (no error, sync hasn't run)

Argo CD sees a diff between git and the cluster. Most infra apps are intentionally manual-sync (ADR-A5).

Sub-check Action
Is this expected? (e.g. you just merged a PR.) Click Sync.
Diff suspicious? (e.g. someone kubectl edit-ed.) argocd app diff $APP. If out-of-band edit happened, the GitOps contract was violated; recommend reverting the manual change or capturing it in a PR.
Diff has sat for > 1 day? Notify the app owner — manual-sync apps rot if no one clicks.

§5 — Synced and Healthy but workload misbehaving

Argo thinks all is fine; the workload is broken. Not a sync failure. Leave this runbook.

Symptom Where to go
Pods crashlooping pod-pending-scheduling.md §3
Ingress 5xx ingress-down.md
Specific feature broken App-team playbook

§6 — Special: blue-green migration in flight

If this app is in a <chart><chart>-<variant> migration:

  • Confirm which variant the Application points at (check spec.source.path).
  • The chart name may have changed in the new variant; release-name pinning via fullnameOverride may be required to keep the same Service DNS during cutover.
  • A failed sync mid-migration is the trigger to roll back (spec.source.path ← old) and Sync, not to push forward.
  • Read blue-green-chart-migration.md before deciding.

Escalation matrix

Symptom Action Escalate to
§1 + bootstrap looks healthy Investigate further App owner
§2a (repo auth) Confirm allowed repoURL; rotate creds Platform team
§2b (path missing) Fix in this repo or sister repo App owner
§2c, §2d (helm render) Reproduce; fix values or chart App owner / platform team
§2e (cluster unreachable) Don't edit destination Platform team
§2f (admission webhook) Fix in chart values App + policy team
§2g (webhook unhealthy) Recover the webhook first Platform team
§2h (immutable field) Probably needs blue-green Platform team
§3 (stuck > 30 min) Check pod events; consider workload rollback App owner

Done conditions

  • argocd app get $APP shows Synced + Healthy.
  • The PR or manual fix that resolved it is on main.
  • If the failure was caused by a regression, a postmortem / RCA is scheduled.