added repo

This commit is contained in:
Your Name
2026-08-26 04:03:34 +05:30
parent 2389ec1fd6
commit 1055e1394f
150 changed files with 12395 additions and 0 deletions
+138
View File
@@ -0,0 +1,138 @@
# Runbook: Deployment Stuck After Sync
> Symptom → diagnosis → remediation for tools that ArgoCD reports as `Synced` but whose Kubernetes resources are not healthy.
>
> **Layer:** 2 (Advisory). Agents diagnose and suggest; humans execute kubectl/ArgoCD commands.
>
> Per AI Blitz Plan §5.2. See also: [argocd-sync-failure.md](argocd-sync-failure.md) (for sync errors), [render-failure.md](render-failure.md) (for Helm render errors).
---
## Symptoms
- ArgoCD Application shows `Synced` but `Degraded` or `Unknown` health
- Pods are in `Pending`, `CrashLoopBackOff`, `ImagePullBackOff`, or `OOMKilled`
- Tool was working before a chart upgrade or values change
- Application is `Synced + Healthy` but the tool isn't functioning as expected
---
## Decision tree
```text
Application shows Synced but Degraded?
├── Check pod status in the tool's namespace:
│ kubectl get pods -n <namespace> --context=<cluster>
├── Pod status is Pending?
│ └── → [A] Resource constraints or node issues
├── Pod status is CrashLoopBackOff?
│ └── → [B] Application crash — check logs
├── Pod status is ImagePullBackOff or ErrImagePull?
│ └── → [C] Image registry issue
├── Pod status is OOMKilled?
│ └── → [D] Memory limit too low
├── Pod runs but tool is unhealthy?
│ └── → [E] Config or connectivity issue
└── Pods are fine but ArgoCD shows Degraded?
└── → [F] Health check misconfiguration
```
---
## [A] Pod Pending
**Diagnosis:**
```bash
kubectl describe pod <pod-name> -n <namespace> --context=<cluster> | grep -A 20 Events
```
Common causes and fixes:
| Event message | Cause | Fix |
| ------------- | ----- | --- |
| `Insufficient cpu/memory` | Node resource exhaustion | Check node capacity; may need to adjust resource requests in `custom-values.yaml` (change in `devops-infra-helm-charts`) |
| `did not match node affinity` | Node selector or affinity mismatch | Review `nodeSelector`/`affinity` in `custom-values.yaml` |
| `PersistentVolumeClaim not bound` | PVC not provisioned | Check storage class and PVC events |
| `Unschedulable` | No nodes available | Check if cluster has sufficient nodes or if Karpenter/cluster-autoscaler is stuck |
---
## [B] CrashLoopBackOff
**Diagnosis:**
```bash
kubectl logs <pod-name> -n <namespace> --context=<cluster> --previous
```
Look for startup errors, missing config, failed connections. The fix usually requires changing `custom-values.yaml` in `devops-infra-helm-charts` — that's a separate PR in the sister repo.
---
## [C] ImagePullBackOff
**Diagnosis:**
```bash
kubectl describe pod <pod-name> -n <namespace> --context=<cluster> | grep image
```
Common causes:
- Image tag doesn't exist (bad chart version pin)
- Image registry credentials expired (check `external-secrets` or `imagePullSecrets`)
- Private registry unreachable from the cluster
Fix depends on root cause. Usually requires a chart fix in `devops-infra-helm-charts` or a credential rotation.
---
## [D] OOMKilled
**Diagnosis:**
```bash
kubectl top pod -n <namespace> --context=<cluster>
kubectl describe pod <pod-name> -n <namespace> --context=<cluster> | grep -A 5 "Last State"
```
Fix: Increase `resources.limits.memory` in `custom-values.yaml` in `devops-infra-helm-charts`. Open PR there.
---
## [E] Config or Connectivity Issue
Tool is running but not working correctly.
**Diagnosis:** Check tool-specific logs. Most infra tools write structured logs.
Common causes:
- Wrong endpoint URL in `custom-values.yaml`
- Service dependency not available (e.g., Vault unreachable for external-secrets)
- Wrong namespace for a referenced service
Fix: Update `custom-values.yaml` in `devops-infra-helm-charts`.
---
## [F] ArgoCD Health Check Misconfiguration
ArgoCD uses health checks to determine if an Application is healthy. Some custom resources have incorrect or missing health checks.
**Diagnosis:**
```bash
argocd app get <appName> --show-operation | grep -A 10 "Health Status"
```
If resources are actually healthy but ArgoCD says Degraded, this may be a health check definition issue in the ArgoCD config itself (not this repo).
Escalate to platform team via [escalation-matrix.md](../../global/escalation-matrix.md).
---
## Escalation
If diagnosis points to a cluster-level issue (node pressure, networking, GKE problem) rather than a configuration issue, escalate via [escalation-matrix.md](../../global/escalation-matrix.md).