Build failed: "Cloning into '.'..." with no further output, then
git config later failing "fatal: not in a git directory" despite the
cloned file being present and editable. Likely cause: this stage just
started running in container('docker-cli') (root — docker:27-cli's
base has no non-root USER) right when it broke, while deleteDir()
just before it runs via the Jenkins agent's own JNLP process (a
different, non-root UID) — git refuses to trust a repo directory
owned by a different UID than the current process, and that refusal
can surface as an unrelated-looking error on a later command rather
than a clear ownership error on the clone itself. Adds
`git config --global --add safe.directory '*'` (safe here — this
workspace is a throwaway, container-local checkout for one build) and
set -e so a genuinely failed clone stops the script immediately
instead of running later commands against partial state.
Replaces per-build on-demand installs (apk add bash/python3/curl,
curl-downloading yq to /tmp) in runHooks.groovy, syncArgoApp.groovy,
and updateHelmTag.groovy with a single custom image
(build-tools.Dockerfile) that has all of it baked in once, at
image-build time — not repeated on every single pipeline run.
updateHelmTag.groovy also now runs inside container('docker-cli')
(previously unwrapped, defaulting to the auto-injected jnlp agent
container, which is why it needed the curl-downloaded yq fallback in
the first place — that container has git but not yq).
dind-pod.yaml's docker-cli container now points at
harbor.192.168.1.7.nip.io/homelab/build-tools:1 instead of the stock
docker:27-cli — this image needs building and pushing once before any
build using this pod template will work; see build-tools.Dockerfile's
header comment.
demo-go-app's Deployment failed to pull: "dial tcp: lookup
harbor-core.harbor.svc.cluster.local on 127.0.0.53:53: server
misbehaving". docker push worked from the Jenkins build pod because
it has pod-network DNS (CoreDNS); pulling for a real Deployment
happens via containerd on the node itself, using the node's host-level
resolver, which has no route to *.svc.cluster.local at all. Switches
buildDocker.groovy's push target and dind-pod.yaml's
--insecure-registry flag to the Contour ingress hostname instead,
which resolves via normal public DNS (nip.io) from both pods and the
host.
Build #14/#15 kept failing "ARGOCD_TOKEN is empty" even after
confirming the actual Kubernetes Secret has a real token value. Root
cause: the check used Groovy's env.ARGOCD_TOKEN, which is Jenkins'
own pipeline-level environment map — populated from build parameters,
environment{} blocks, withEnv, etc. — not the container's actual OS
environment. A container-scoped env: entry in a podTemplate YAML
(dind-pod.yaml's secretKeyRef) never populates that Groovy map; it
only sets the real process environment inside that container, which
sh steps correctly inherit. So this check was always going to see
null regardless of how correctly Vault/ESO/the Secret were wired —
every fix to that chain was chasing the wrong problem. Moves the
emptiness check into the shell script itself, where $ARGOCD_TOKEN
genuinely resolves.
Build #13: "apk: not found". This sh step isn't wrapped in
container('docker-cli') — unlike runHooks/buildDocker, it never
specifies a container, so it runs in the auto-injected jnlp agent
container (Debian-based jenkins/inbound-agent), not the Alpine
docker-cli one. git clone worked fine in the same step (Debian image
bundles git), just no apk. Fetches the static mikefarah/yq binary via
curl instead of relying on any particular package manager being
present, to /tmp rather than /usr/local/bin since the agent likely
runs as a non-root user.
Build #11 failed: "yq: not found". docker:27-cli's Alpine base
doesn't ship yq by default, same gap as the bash/python3 on-demand
installs already in runHooks.groovy. Alpine renamed the mikefarah/yq
package from `yq` to `yq-go` at v3.20 (and an unrelated Python-based
tool is also sometimes packaged as plain `yq` on other distros, with
incompatible syntax) — tries yq-go first, falls back to yq, rather
than assume which Alpine version docker:27-cli currently ships.
Build #10 failed before its sh step even ran:
"IllegalArgumentException: named capturing group is missing trailing
'}'". Root cause: replaceFirst('http://', "http://\${GIT_USER}:...")
— replaceFirst's *replacement* argument is parsed with Java
regex-replacement syntax, where ${name} means "substitute named
capture group", not literal text. The pattern 'http://' has no named
groups, so Java's regex engine choked trying to resolve
${GIT_USER}/${GIT_PASS} as group references. Replaced with a plain
string split + concatenation, which has no regex-replacement
semantics to collide with, while still keeping \${GIT_USER}/
\${GIT_PASS} literal in the Groovy string so the shell (not Groovy)
expands them from the credential-bound env vars at sh-step time.
Confirmed via a direct token request to Harbor's own token endpoint
(same username/password from harbor-robot-dockerconfig) that the
robot account genuinely has push access to homelab/demo-go-app —
Harbor returned a valid token with actions:[pull,push]. So the actual
`docker push` failure ("unauthorized... action: push") wasn't a
permissions problem at all.
Root cause: kubernetes.io/dockerconfigjson secrets are required to
store their data under the fixed key `.dockerconfigjson`. Mounting
the secret without remapping that key meant the file that actually
landed at /root/.docker was named `.dockerconfigjson`, not
`config.json` — the only filename docker's CLI reads for stored
credentials. Docker found nothing there and pushed unauthenticated,
which Harbor correctly rejected. Adds an items: remap so the mounted
file is named config.json.
Build #7 failed with "Cannot connect to the Docker daemon at
tcp://localhost:2375" right at the first sh step. Verified the dind
entrypoint script directly (docker-library/docker's
dockerd-entrypoint.sh) — with DOCKER_TLS_CERTDIR="" and a
dash-prefixed arg it correctly builds
`dockerd --host=tcp://0.0.0.0:2375 --insecure-registry=...`, so the
--insecure-registry flag added last commit isn't logically wrong.
This is a container-start race instead: Kubernetes doesn't guarantee
ordering between containers in the same pod, so the docker-cli
container's first sh can fire before dockerd in the sibling container
has finished its startup checks (iptables detection etc. run on every
start). Polls `docker info` for up to 60s before the actual build
instead of assuming instant availability.
docker push was hanging: "Client.Timeout exceeded while awaiting
headers" doing a TLS handshake against harbor-core.harbor.svc.cluster.local,
which only ever speaks plain HTTP (TLS disabled cluster-wide by
design). Docker defaults to HTTPS for any bare registry hostname
regardless of network path — that default has nothing to do with
whether traffic routes through Contour/Ingress, contrary to what an
earlier pending-items note assumed. Adds --insecure-registry to the
dind container's dockerd startup args.
Build #5 got through checkout, loadConfig, and the pre_build hook,
then hit two real bugs at the actual docker build step:
1. `docker build` failed with "mkdir /root/.docker/buildx: read-only
file system" — dind-pod.yaml mounts the Harbor push-auth secret
read-only at /root/.docker, but modern docker defaults to
BuildKit/buildx, which wants to write its own state there. Forces
the classic builder via DOCKER_BUILDKIT=0 instead.
2. The subsequent error-handling itself then threw a
NullPointerException — every stage file (including untouched
legacy ones from the real devops-lib) reads env.FAILURE when
setting currentBuild.result, but nothing in this repo ever defined
it, so it was null. Defined once in homelabPipeline.groovy rather
than touching 40+ individual occurrences across every stage file.
Drops AGENTS.md, BUGS_AND_IMPROVEMENTS_REPORT.md, CLAUDE.md,
ai-blitz/*, and docs/{SECURITY.md,acronyms.md,adr/*} — leftover
documentation and task scaffolding from the original org-wide devops-
lib that don't describe this homelab's simplified single-service
pipeline. These were already missing from the working tree from
earlier cleanup; this commit just records that state.
Renames src/com/meesho -> src/com/homelab, resources/com/meesho ->
resources/com/homelab, resources/org/meesho -> resources/org/homelab
(via git mv, preserving history), and sweeps every remaining
occurrence of "meesho" (any casing) out of package declarations,
imports, libraryResource() paths, and comments across the whole repo.
Also drops the per-user allowlist in vars/eksCICD.groovy, which
hardcoded real former-colleagues' emails and doesn't apply to a
single-person homelab — that branch is now permanently skipped rather
than deleted outright, to avoid hand-editing the escape-sequence-heavy
echo blocks it guards (eksCICD.groovy itself is unused legacy code,
not called by homelabPipeline.groovy).
Does not touch the ~114 files that were already missing from the
working tree but still tracked in the prior commit — that's unrelated
pre-existing state, left as-is.