infra

Author	SHA1	Message	Date
Viktor Barzin	b858561bd0	Merge remote-tracking branch 'origin/master' Some checks failed ci/woodpecker/push/default Pipeline failed Details	2026-06-24 20:59:39 +00:00
Viktor Barzin	a7704f46a6	deploy goldmane-edge-aggregator: durable who-talks-to-whom edge trail (#58 , ADR-0014) Infra side of ADR-0014: an mTLS gRPC consumer of Calico Goldmane's Flows API that records the namespace-pair edge-set in CNPG and posts a daily new-edge digest to #security. Adds the goldmane-edge-aggregator stack, the pg-goldmane-edges Vault rotation role (Tier-0 vault state updated here), and the namespace in the ghcr-credentials allowlist. Cert: REUSES the operator-minted, Tigera-CA-signed whisker-backend client cert (Goldmane verifies only the CA chain, not identity) instead of minting from the Tigera CA private key. This avoids putting the CA key in TF state AND the hashicorp/tls provider, which is incompatible with this repo's global generate-providers/lockfile pattern (it broke every stack's lockfile). Verified live: aggregator streaming flows, 174 edges in Postgres across 50x54 namespaces, db+slack ExternalSecrets synced, digest dry-run formats correctly, private image pulls via the Kyverno-synced ghcr-credentials. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-24 20:59:39 +00:00
Viktor Barzin	aa510e3600	instagram-poster: force_conflicts on ESO manifests (fix apply) All checks were successful ci/woodpecker/push/default Pipeline was successful Details The ESO v1 migration (2026-06-22) made the external-secrets controller own .spec.refreshInterval via server-side apply, so terraform apply of the two ExternalSecret manifests fails with a field-manager conflict (Woodpecker #348), which blocked the replicas=0 scale-down from landing. Add force_conflicts=true to both, matching the grafana/woodpecker/traefik fix applied to other stacks the same day. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-24 20:49:53 +00:00
Viktor Barzin	53834deb24	instagram-poster: scale to 0 (unused, dead ExternalSecret) Some checks failed ci/woodpecker/push/default Pipeline failed Details Viktor confirmed the Instagram Graph poster isn't used. Its ExternalSecret has been dead on missing Vault keys (ig_graph_long_lived_token, ig_business_account_id), so the deployment sat at 0/1 firing DeploymentReplicasMismatch. Setting replicas=0 stops the alert and makes the scale-down durable (a bare kubectl scale reverts on the next stack apply). Re-set to 1 after minting a Meta long-lived token + populating the Vault keys. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-24 20:45:30 +00:00
Viktor Barzin	8dd9a3978d	Merge remote-tracking branch 'forgejo/master' into wizard/homelab-vault All checks were successful ci/woodpecker/push/default Pipeline was successful Details	2026-06-24 12:25:52 +00:00
Viktor Barzin	65b2df1222	fix(monitoring): force_conflicts on grafana_db_creds ExternalSecret The external-secrets controller owns .spec.refreshInterval via SSA, so a plain terraform apply of the monitoring stack conflicts. Latent until 2026-06-24 (the homelab-vault loki-rules change was the first monitoring apply in a while and surfaced it). force_conflicts lets TF win — same pattern as woodpecker/traefik/ k8s-version-upgrade stacks.	2026-06-24 12:25:36 +00:00
Viktor Barzin	1d0388da12	Merge remote-tracking branch 'origin/master' All checks were successful ci/woodpecker/push/default Pipeline was successful Details	2026-06-24 12:22:58 +00:00
Viktor Barzin	92361f36db	calico: enable Goldmane + Whisker (Calico 3.30 OSS flow observability) Turns on Calico 3.30's native east-west flow observability so we can see which Service talks to which (ADR-0014, issue #57). Enabled via the operator CRs directly (kubectl_manifest Goldmane + Whisker, name=default) rather than the Helm goldmane/whisker flags, because the goldmanes/whiskers CRDs already exist and this sidesteps the helm-upgrade CR-before-CRD ordering issue. Whisker notifications=Disabled so the UI doesn't call the external Tigera endpoint. Applied supervised: creating the Goldmane CR re-rendered calico-node with the FELIX_FLOWLOGSGOLDMANESERVER env (operator auto-wires Felix — no manual FelixConfiguration); calico-node rolled cleanly 7/7, tigerastatus healthy, goldmane is receiving flows from all nodes, Whisker UI serves. Durable Loki persistence is NOT included here: the Goldmane emitter is Calico Cloud/Enterprise-gated with no OSS knob to aim it at Loki (the CR can override only name+resources, not env), so a durable trail needs a small custom gRPC consumer of goldmane:7443 — tracked in issue #58. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-24 12:22:48 +00:00
Viktor Barzin	e711b2f971	feat(monitoring): homelab vault traceability alerts (TOTP-fetch + volume) Some checks failed Build infra CLI / build (push) Waiting to run Details ci/woodpecker/push/default Pipeline failed Details Adds a Loki ruler group (lane=security -> #security) for the homelab vault op-log: VaultwardenTOTPFetched (every 2nd-factor fetch is visible) and VaultwardenFetchVolumeHigh (>100 fetches/10m backstop). The audit spine (Vault audit device, reads of secret/data/workstation/claude-users/*) is already captured. True CLI-bypass detection needs cross-stream correlation (follow-up).	2026-06-24 10:31:32 +00:00
Viktor Barzin	0293b5c634	android-emulator: fix idle-sleeper dying with SIGPIPE before it could sleep All checks were successful ci/woodpecker/push/default Pipeline was successful Details Caught live-testing the previous commit: every sleeper run exited 141 (SIGPIPE) in ~1s with no output, never reaching the scale-down. Cause: `set -o pipefail` + `dumpsys power \| awk '...; exit'` — awk closes the pipe after the first match while `kubectl exec` is still streaming dumpsys, so the exec gets SIGPIPE, pipefail makes the pipeline 141, and set -e kills the script before any echo. (My earlier dry-run missed it because it didn't run under `set -euo pipefail`.) Fix: drop pipefail; capture each exec to a var (`\|\| true`) then parse with awk reading to END (no early `exit`), so nothing can SIGPIPE mid-stream and a failed/booting exec falls through to the fail-safe "do not sleep" branch. Also fetch the pod name via jsonpath instead of `-o name \| head -1` (no pipe to SIGPIPE, no `pod/` prefix to strip), and exec `adb` directly without the `sh -c` wrapper. Verified live: ran the corrected script as the gate ServiceAccount against the stuck emulator (idle ~120h) — it logged "idle >= 6h ... scaling to zero" and patched the deployment to replicas=0. The 6+ day pod is now asleep. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-24 08:57:36 +00:00
Viktor Barzin	839fdb33c2	android-emulator: sleep after 6h idle (activity-based), fix never-sleeping All checks were successful ci/woodpecker/push/default Pipeline was successful Details The emulator was meant to scale to zero when idle but had been up 6+ days straight despite ~5 days with no real use. Two bugs: 1. The idle check counted ESTABLISHED TCP connections to the adb/noVNC ports. A forgotten `adb connect` (no disconnect) holds that transport open forever, so every 15-min run saw "active" and reset the counter -- it never reached the sleep branch. (Right now: 4 such stale transports from pods on k8s-node3/node4.) 2. Even when it did reach the sleep branch, `kubectl scale --replicas=0` failed Forbidden -- the gate ServiceAccount can patch `deployments` but not `deployments/scale`. Switch the sleeper to measure actual use: time since last user activity (taps/keys/app-launches, incl. noVNC clicks) from `dumpsys power` vs guest uptime. No interaction for 6h -> sleep. This ignores idle/forgotten connections entirely. Scale down with a direct replicas patch on the named deployment (same path the wake gate scales up), so it needs only the existing `deployments` patch grant -- no `deployments/scale`. Now stateless (drops the idle-counter annotation; gate.py no longer sets it) and lighter on etcd. Fail-safe: any read error (e.g. mid-boot) does not sleep. Requested by Viktor: turn the dev-only emulator off when it hasn't been used for 6h. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-24 08:49:23 +00:00
Viktor Barzin	566447a698	k8s-upgrade: preflight kubeadm-plan gate must pass explicit target (minor-upgrade fix) All checks were successful ci/woodpecker/push/default Pipeline was successful Details Last night's 1.34.9->1.35.6 run passed the ESO/kyverno compat gate (the migration worked!) but ABORTED at the kubeadm-plan-target gate: it ran `kubeadm upgrade plan` with NO version, so master's old 1.34.9 kubeadm auto-proposed only the current minor (Loki: "falling back to stable-1.34") and plan_target != 1.35.6 -> abort. That gate worked for patch upgrades but never for minors. Fix: pass the explicit `v$TARGET_VERSION` (verified on master: `kubeadm upgrade plan v1.35.6` emits "kubeadm upgrade apply v1.35.6"). Works for patches too. Applied live to the ConfigMap before tonight's run; deleted the failed preflight-1-35-6 job. Also: ESO 2.x took SSA ownership of .spec.refreshInterval, so terraform's apply of the k8s-upgrade-creds ExternalSecret hit a field-manager conflict. Added field_manager.force_conflicts=true (benign — interval is semantically identical). This pattern affects all 104 migrated ESs fleet-wide (follow-up). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-24 06:06:14 +00:00
Viktor Barzin	98d2b89614	calico: bump tigera-operator mem limit 256Mi -> 512Mi (OOM crashloop fix) All checks were successful ci/woodpecker/push/default Pipeline was successful Details The operator OOM-crashlooped on 2026-06-23: it idles at ~246Mi with a ~266Mi startup spike (re-listing resources to build informer caches), both at/over the 256Mi limit, so the first time the pod restarted it could never finish startup (exit 137 OOMKilled, leader-elect, OOM, repeat). A latent landmine — the limit was always too tight; it only bit once the pod restarted. Data plane was never affected (calico-node 7/7, tigerastatus green throughout). 512Mi gives headroom (now ~246Mi steady, verified stable 0 restarts). NOT caused by the ESO migration (which never touched calico); cluster churn was at most the trigger that exposed the tight limit. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-23 12:46:28 +00:00
Viktor Barzin	68c240b8de	Merge remote-tracking branch 'origin/master' Some checks failed ci/woodpecker/push/default Pipeline failed Details	2026-06-23 09:56:25 +00:00
Viktor Barzin	7d297dc6b1	eso: complete migration — chart 2.6.0, all CRs on v1, 1.35 gate cleared Phase 3 of the ESO 0.12->2.6 migration (the last k8s-1.35 compat-gate blocker). Climbed external-secrets 0.16.2 -> 0.17.0 -> ... -> 2.6.0 one minor at a time, each hop applied + verified (ES sync held at 109 Ready every hop; atomic=true rollback safety net). Crossed the 0.17 cutoff (v1beta1 serving removed) only after Phase 2 put all 104 ExternalSecrets + 2 ClusterSecretStores on external-secrets.io/v1. Result: compat-gate now returns "OK: cluster is safe to upgrade to 1.35.6" (EXIT 0) — the autonomous version-check chain will take k8s 1.34 -> 1.35 on its next nightly run. Also fixes the repo-wide stale-lock issue that broke CI pipeline 332: the terragrunt-generated providers.tf declares gavinbunney/kubectl + telmate/proxmox, but ~28-39 stacks' committed .terraform.lock.hcl predated that ("Inconsistent dependency lock file: no version selected"). Reconciled via `tg init -upgrade` and committed so `terragrunt apply`/CI work cleanly again. Docs: .claude/CLAUDE.md ESO line corrected (104 ESs, v1, chart 2.6.0); plan doc marked COMPLETE. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-23 09:55:51 +00:00
Viktor Barzin	59f2beda21	chrome-service: run real Google Chrome (H.264/AAC codecs) for the browser All checks were successful ci/woodpecker/push/default Pipeline was successful Details Point the chrome-service container at the new chrome-service-browser image and launch /opt/google/chrome/chrome instead of the bundled Chromium. Fixes MEDIA_ERR_SRC_NOT_SUPPORTED on H.264/AAC video (Instagram Reels etc.) in the noVNC view — bundled Chromium has those codecs compiled out; only real Chrome carries them. connect_over_cdp callers (tripit fare scrape, homelab browser, snapshot-harvester) attach over raw CDP (version-tolerant) — validated after rollout. Image is built off-infra on GHA (prior commit) → public ghcr. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-22 21:15:36 +00:00
Viktor Barzin	df1ec1879d	chrome-service: build a real-Chrome browser image (H.264/AAC codecs) Some checks failed ci/woodpecker/push/default Pipeline was successful Details Build chrome-service-browser / build (push) Has been cancelled Details Add an infra-owned image (Playwright base + google-chrome-stable) + its GHA build workflow. The bundled Chromium ships proprietary codecs compiled out, so H.264/AAC video (Instagram Reels, X, most .mp4) fails in the noVNC view with MEDIA_ERR_SRC_NOT_SUPPORTED; only real Google Chrome carries those codecs (libffmpeg swap + Chrome-for-Testing both ruled out). This commit only builds the image (→ ghcr.io/viktorbarzin/chrome-service-browser); a follow-up flips main.tf's launch to it once the image exists + is public. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-22 21:01:17 +00:00
Viktor Barzin	c670cb7118	eso: Phase 2 — migrate all 104 ExternalSecrets + 2 ClusterSecretStores to v1 Some checks failed ci/woodpecker/push/default Pipeline failed Details The API rewrite half of the ESO 0.12->2.6 migration (last k8s-1.35 compat-gate blocker). Done on chart 0.16.2, which serves BOTH external-secrets.io/v1beta1 and v1, so this is the safe window — MUST land before 0.17 removes v1beta1 (there is no conversion webhook). Pure apiVersion bump, schema is byte-identical: 106 occurrences (104 ExternalSecrets + 2 ClusterSecretStores vault-kv/vault-database) across 73 .tf files, v1beta1 -> v1, no other field changes. Validated live first on tandoor (single, non-coupled, synced ES): the kubernetes_manifest apiVersion bump forces a REPLACE; the target Secret is cascade-GC'd for ONE ~0.3s poll then ESO recreates it (identical value re-synced from Vault, new UID) and the ES returns SecretSynced=True on v1. Running pods keep their mounted copy through the sub-second blip. All 110 target Secrets were snapshotted to /tmp first as a backstop. CI applies the changed stacks serially (staged rollout); watching aggregate ES sync back to 108 synced (2 pre-existing dead: instagram-poster, payslip-ingest). Next: Phase 3 climb 0.16.2 -> 2.6.0. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-22 19:13:04 +00:00
Viktor Barzin	98cd535b97	authentik: lock chrome.viktorbarzin.me noVNC to Viktor only All checks were successful ci/woodpecker/push/default Pipeline was successful Details The chrome-service noVNC exposes Viktor's live logged-in browser sessions (Instagram etc. — he'll sign in there for homelab browser to reuse). It was auth="required" = any authenticated user, and "Home Server Admins" includes emo (emil.barzin@gmail.com), so the admin group is not a sufficient gate. Add a host-specific case to the domain-wide forward-auth restriction allowing only Viktor's accounts (vbarzin@gmail.com + akadmin break-glass); everyone else, incl. emo, is denied at the noVNC. emo's AGENT already can't reach the browser (read-only RBAC blocks port-forward); this closes the human noVNC path. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-22 18:09:27 +00:00
Viktor Barzin	a3cdc0d6d0	chrome-service: size headed Chrome window to fill Xvfb (noVNC cut-off) All checks were successful ci/woodpecker/push/default Pipeline was successful Details The noVNC view showed the browser in the top-left with the rest of the framebuffer black. Cause: Chrome launched with no --window-size, and there's no window manager, so it opened at its profile-persisted (smaller) size inside the 1280x720 Xvfb. Add --window-size=1280,720 --window-position=0,0 so the window fills the screen on every launch (fresh pods/profiles too). Live windows were already resized via CDP as a stopgap. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-22 18:00:20 +00:00
Viktor Barzin	c7ead032ec	chrome-service: fix noVNC stuck-"Connecting" (x11vnc fd-sweep under nofile=2^31) Some checks failed ci/woodpecker/push/default Pipeline was successful Details Build chrome-service-novnc / build (push) Has been cancelled Details The noVNC view hung on "Connecting" forever then timed out. Root cause: x11vnc sweeps the entire fd table (fcntl per fd) on every client connection, and containerd grants pods RLIMIT_NOFILE=2^31, so the RFB handshake never completes (websockify accepts the WS and dials localhost:5900, but x11vnc never sends its banner — verified: handshake timed out at 8s, x11vnc had burned 1h41m CPU spinning). Same bug + fix the android-emulator stack already carries. Cap nofile before x11vnc starts, in two places: - files/novnc/entrypoint.sh: `ulimit -n 65536` (root fix, makes the image correct) - main.tf novnc container: `command = ["bash","-c","ulimit -n 65536; exec /entrypoint.sh"]` so the cap applies deterministically on rollout even though the image is :latest/IfNotPresent (a rebuilt entrypoint isn't guaranteed to be re-pulled). Also documents the gotcha + diagnosis in docs/architecture/chrome-service.md and notes the black-when-idle behaviour + the autoconnect URL. (A live x11vnc relaunch with the cap already unblocked the running pod; this makes it survive restarts.) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-22 17:34:03 +00:00
Viktor Barzin	20ca5ee624	tripit: REEL_PROVIDER=anonymous — actually fetch reels (was fake canned caption) All checks were successful ci/woodpecker/push/default Pipeline was successful Details REEL_PROVIDER was unset, so the reel pipeline used FakeReelExtractor, which returns a CANNED caption — every pasted (tripit #120) or forwarded reel produced a DUMMY Saved Place instead of reading the real reel. Set REEL_PROVIDER=anonymous in app_env (covers the web Deployment + the ingest CronJob) so AnonymousReelExtractor does the real anonymous read. Verified live from the cluster: yt-dlp fetched a real IG /p/ caption (no IG_GRAPHQL_DOC_ID needed — the internal-API path is an optional optimisation; yt-dlp fallback works). LLM extraction + Nominatim POI geocoding were already real (prior commits); this was the last fake link in the chain. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-22 17:30:47 +00:00
Viktor Barzin	f46b69f372	tripit: enable real LLM + Nominatim on the web Deployment (in-app reel paste #120 ) All checks were successful ci/woodpecker/push/default Pipeline was successful Details The web Deployment ran LLM_MODE=fake with no reel geocoder — only the ingest-plans CronJob had real providers. The in-app reel-URL paste feature (tripit #120) runs ingest_reel IN the web pod (BackgroundTask), so the Deployment now needs real extraction: LLM_MODE=llamacpp (qwen3vl-8b; qwen3-8b segfaults on the current llama-swap image) with the ADR-0033 claude-agent-service fallback, plus REEL_GEOCODER_PROVIDER=nominatim for venue->city/country POI geocoding. Set in app_env (feeds the Deployment; the CronJobs already had these via extra_env). Bonus: this also un-fakes the in-app booking share import, which used the same fake LLM. MAIL_INGEST_ENABLED stays false on the Deployment (only the CronJob polls mail). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-22 16:50:04 +00:00
Viktor Barzin	59f2070e56	tripit: switch mail-ingest LLM_MODEL qwen3-8b -> qwen3vl-8b (qwen3-8b segfaults) All checks were successful ci/woodpecker/push/default Pipeline was successful Details The qwen3-8b GGUF segfaults on load on the current llama-swap :cuda image ("common_init_from_params: failed to create context"; llama-swap returns 502), which broke ALL tripit mail ingest text extraction — booking emails AND forwarded reels (status=failed, "no place could be read"). The GGUF isn't corrupt (valid header, full size, worked for weeks) — it's a llama.cpp/image regression. Rather than pin the SHARED llama-swap image (cross-user blast radius), repoint the ingest-plans CronJob at qwen3vl-8b, an already-provisioned 8B model that loads fine and extracts flight numbers + places reliably. Restores the auto-path (reels resolve via the Nominatim geocoder; bookings parse again). The broken qwen3-8b GGUF is a separate, non-urgent llama-cpp cleanup. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-22 15:52:09 +00:00
Viktor Barzin	f96cde35bd	tripit: enable Nominatim POI geocoding for reel→Wishlist ingest All checks were successful ci/woodpecker/push/default Pipeline was successful Details Forwarded reels (tripit ADR-0031) geocode their venue to map a Saved Place to a country + city, but the reel route was wired to the global geocoder, which here is GEOCODER_PROVIDER=openmeteo (city-level, name-based). OpenMeteo returns nothing for a venue query like "Time Out Market, Lisbon" so reels never resolved and no Saved Place was created. The app fix (tripit 3c62d596) gave the reel route its own geocoder behind REEL_GEOCODER_PROVIDER; set it to nominatim on the ingest-plans CronJob (the only one running the reel route) so forwarded reels resolve to real venue coords + city + country. Isolated from the global geocoder, which stays openmeteo for weather/tours. Verified Nominatim resolves the venue from the cluster. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-22 14:59:37 +00:00
Viktor Barzin	aeed461591	Revert "feat(monitoring): Tempo + OTel Collector for tripit tracing, hardened (ADR-0032 Phase 2)" All checks were successful ci/woodpecker/push/default Pipeline was successful Details This reverts commit `1595bddfc2`.	2026-06-22 08:31:17 +00:00
Viktor Barzin	1595bddfc2	feat(monitoring): Tempo + OTel Collector for tripit tracing, hardened (ADR-0032 Phase 2) Some checks failed ci/woodpecker/push/default Pipeline failed Details Re-land Phase 2 after the first attempt's two failure modes, both fixed: - tempo.resources set under the correct single-binary chart key (was OOMKilled on the namespace LimitRange default when mis-placed at top level). - atomic=true + cleanup_on_fail=true on BOTH helm releases — a failed install auto-rolls-back instead of leaving a stuck/orphaned release (memory #6479). Tempo (single-binary, proxmox-lvm 20Gi, 30d) + OTel Collector (contrib; otlp -> redaction -> batch -> tempo) + Tempo datasource + additive trace_id->Tempo derivedField on Loki + tripit LOG_FORMAT=json/OTEL_EXPORTER_OTLP_ENDPOINT. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-22 08:17:59 +00:00
Viktor Barzin	464e0bfb97	Revert "feat(monitoring): Tempo + OTel Collector for tripit tracing (ADR-0032 Phase 2)" All checks were successful ci/woodpecker/push/default Pipeline was successful Details This reverts commit `7513468a2d`.	2026-06-22 06:46:56 +00:00
Viktor Barzin	72dcb125d5	Revert "fix(monitoring): tempo OOMKilled — move resources under tempo.resources" This reverts commit `a02782d11f`.	2026-06-22 06:46:56 +00:00
Viktor Barzin	a02782d11f	fix(monitoring): tempo OOMKilled — move resources under tempo.resources Some checks failed ci/woodpecker/push/default Pipeline failed Details Pipeline #315 failed: tempo-0 CrashLoopBackOff / OOMKilled (exit 137). The single-binary grafana/tempo chart (v1.24.4) takes container resources at tempo.resources, not a top-level resources: — so my block was ignored and the pod fell to the namespace LimitRange default and OOMed. Set tempo.resources explicitly (req 256Mi / limit 2Gi). tripit + existing monitoring were unaffected throughout. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-22 06:44:31 +00:00
Viktor Barzin	7513468a2d	feat(monitoring): Tempo + OTel Collector for tripit tracing (ADR-0032 Phase 2) Some checks failed ci/woodpecker/push/default Pipeline failed Details Stand up the cluster's first trace store + OTLP ingress so tripit's OpenTelemetry spans (Phase 1, already live in prod) export and correlate with logs: - Grafana Tempo (single-binary, filesystem on proxmox-lvm 20Gi, 30d) - OTel Collector (contrib; otlp -> redaction deny-list backstop -> batch -> tempo) - Grafana: a Tempo datasource + an ADDITIVE trace_id->Tempo derivedField on the Loki datasource (no uid change, so existing dashboards are unaffected) - tripit deployment: LOG_FORMAT=json + OTEL_EXPORTER_OTLP_ENDPOINT -> the Collector Additive (new helm releases; Loki/Prometheus/Grafana untouched). Offline 'terraform validate' clean; full plan+apply runs in CI (locked git-crypt blocks a local plan as non-admin). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-22 06:31:11 +00:00
Viktor Barzin	ac27e41fde	Merge remote-tracking branch 'origin/master' All checks were successful ci/woodpecker/push/default Pipeline was successful Details	2026-06-21 20:41:35 +00:00
Viktor Barzin	296deda3b4	eso: Phase 1 — climb chart 0.12.1 -> 0.16.2 (transition version) + atomic First half of the ESO 0.12->2.6 migration (docs/plans/2026-06-21-eso-0.12-to-2.x-migration-design.md), clearing the LAST k8s-1.35 compat-gate blocker. Stepped one minor at a time on k8s 1.34 (no k8s interleave — cluster already on 1.34, ESO bands are conservative tested ranges not hard limits): 0.12.1 -> 0.13.0 -> 0.14.4 -> 0.15.1 -> 0.16.2. Each hop applied + verified: controller healthy, all 108 live ExternalSecrets stayed SecretSynced (2 pre-existing dead — instagram-poster, payslip-ingest — missing Vault data, untouched). Added atomic=true + timeout=600 (ESO had no rollback safety net). 0.16.2 serves BOTH v1beta1 AND v1 (storedVersions now ["v1beta1","v1"]) — the safe window to rewrite all 104 CRs to v1 (Phase 2) before 0.17 removes v1beta1. State auto-committed per hop by scripts/tg (Tier-0 SOPS). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-21 20:41:30 +00:00
Viktor Barzin	de2250f667	immich-frame: set photo date format to dd/MM/yyyy All checks were successful ci/woodpecker/push/default Pipeline was successful Details The photo date overlay was showing US-style MM/dd/yyyy — ImmichFrame's built-in default when PhotoDateFormat is unset. Viktor wants UK day/month/year ordering instead. Pin PhotoDateFormat to the date-fns pattern "dd/MM/yyyy" (uppercase MM = month; lowercase mm would render minutes). The config map carries reloader.stakater.com/match, so Reloader restarts the immich-frame pod automatically on apply. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-21 20:36:43 +00:00
Viktor Barzin	0bae025b9b	wealth dashboard: spend-down figures in today's money (inflation-adjusted) All checks were successful ci/woodpecker/push/default Pipeline was successful Details Viktor asked whether the spend-down numbers were inflation-adjusted — they were not (all nominal). He chose to switch the card to today's money, so every row now shows constant purchasing power for life. Each row is a die-with-zero annuity at the REAL rate (1+g)/1.03−1 (3% inflation), spending a constant inflation-adjusted amount (the actual pounds withdrawn rise with inflation) until net worth hits £0 at age 100: • No growth (0%) → £12/day, £370/mo, £4,446/yr (negative real: loses to inflation) • Inflation (3%) → £43/day, £1,315/mo, £15,776/yr (0% real: holds value) • Market (7%) → £130/day, £3,942/mo, £47,300/yr (~3.9% real) Title now flags "(today's £)". Same panel/layout; only the SQL, title, and tooltip changed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-21 20:13:59 +00:00
Viktor Barzin	3fb6284e2b	immich-frame: use 24-hour clock (ClockFormat HH:mm) All checks were successful ci/woodpecker/push/default Pipeline was successful Details Viktor asked to switch the Immich photo-frame shown on the Portal kitchen appliance to a 24-hour clock. immichFrame defaults ClockFormat to 'hh:mm' (12-hour) and we never overrode it, so the frame was showing 12-hour time. Set ClockFormat: "HH:mm" (date-fns 24h token) in the frame Settings.yml ConfigMap; Reloader restarts the pod on apply. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-21 20:10:51 +00:00
Viktor Barzin	e89de86af0	wealth dashboard: spend-down table → three growth scenarios All checks were successful ci/woodpecker/push/default Pipeline was successful Details Viktor wanted the spend-down card to compare three portfolio-growth scenarios rather than the previous floor-vs-4%-real pair. The table now has three rows, each a die-with-zero annuity (drain net worth to £0 by age 100) spending a constant number of ACTUAL (nominal) pounds, differing only by the assumed nominal growth rate: • No growth (0%) → £43/day, £1,315/mo, £15,776/yr (= NW ÷ years) • Inflation (3%) → £106/day, £3,233/mo, £38,792/yr (NEW) • Avg market (7%) → £220/day, £6,703/mo, £80,435/yr This keeps the £43 no-growth floor he anchored on. The old third row was "4% real" (£133) expressed in today's money; it's replaced by the 7%-nominal market row (£220, actual pounds) so all three rows share one basis (nominal pounds) and are directly comparable. 3%/7% are hardcoded (one-line SQL edit). Table height 4→5 for the extra row; panels below shifted down 1. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-21 20:06:29 +00:00
Viktor Barzin	85d42f2c13	wealth dashboard: merge spend-down tiles into one compact table All checks were successful ci/woodpecker/push/default Pipeline was successful Details Viktor wanted the six separate spend-down stat tiles consolidated into a single, more compact card with the figures laid out as rows. Replaces stat panels 9220-9225 with one table panel (id 9220) in the Overview row: 2 rows (Floor / 4% real) × 3 columns (per day / month / year). Same underlying math and live values (£43/£1,315/£15,776 floor; £133/£4,039/£48,463 at 4% real). w=9 instead of the full-width tile row, so it takes ~a third of the width. Note: this intentionally overrides the "table panels live at the bottom" layout convention — Viktor chose to keep this headline KPI glanceable at the top of the dashboard rather than scroll for it. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-21 19:55:57 +00:00
Viktor Barzin	63add2a126	feat(tripit): finalize ADR-0028 auth env — AUTH_MODE=normal, trips@ sender, trust XFF All checks were successful ci/woodpecker/push/default Pipeline was successful Details Now that the native-auth rollout is complete: (1) AUTH_MODE hybrid->normal — the legacy Authentik OIDC-bearer + forward-auth arms were removed in #96, and 'hybrid' already resolved to 'normal' via backward-compat parsing; this makes it explicit and corrects the now-false comment. (2) SMTP_FROM plans@->trips@ — the dedicated native-auth sender; the trips@->spam@ send-as alias is live + verified (RCPT 250). (3) TRUST_FORWARDED_FOR=true — so #95's per-IP signup rate-limit keys on the real client behind Traefik, not the shared ingress pod IP. Env-only; the Deployment image is KEEL_IGNORE_IMAGE (lifecycle-ignored), so this does NOT touch the running image. Reloader restarts the pod to pick up the new env. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-21 19:50:20 +00:00
Viktor Barzin	166a2bcab4	wealth dashboard: add "spend-down to £0 at 100" stat tiles All checks were successful ci/woodpecker/push/default Pipeline was successful Details Viktor wanted a glanceable number on the Wealth dashboard for how much he can spend for the rest of his life — spending the whole net worth down to zero by age 100. Adds a third line of six stat tiles to the Overview section, two equations × three cadences (per day / month / year): • FLOOR — net worth ÷ time remaining to age 100. Treats the money as cash (no growth, no inflation): a conservative lower bound. ≈ £43/day, £1.3k/mo, £15.8k/yr. • 4% REAL — die-with-zero annuity: the constant, inflation-adjusted spend that drains the balance to £0 at 100 while it keeps earning 4% real. PMT = NW·r/(1−(1+r)^−n). ≈ £133/day, £4.0k/mo, £48.5k/yr. Horizon is today → his 100th birthday (DOB 1998-10-04 → 2098-10-04), computed live so the figures tick as net worth and the horizon move. Net worth reuses the existing latest-per-account dav_corrected math, so the tiles always agree with the "Net worth (current)" stat (pension included; target £0). The 4% real rate is hard-coded per his "keep it simple, just a number" steer — a one-line SQL edit to change later. Layout: tiles inserted at y=9; all sections below shifted down 4 rows. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-21 19:48:30 +00:00
Viktor Barzin	51838a4ec7	kyverno: 3.6.1 -> 3.8.1 (app 1.16 -> 1.18.1) — clears the k8s-1.35 compat-gate block All checks were successful ci/woodpecker/push/default Pipeline was successful Details kyverno v1.16 supports k8s <=1.34, so it was one of the two addons blocking the autonomous 1.35 upgrade (compat gate, nightly). v1.18 supports 1.35. Stepped one minor at a time per the kyverno upgrade guide (per-minor CRD notes): 3.6.1 (1.16) -> 3.7.2 (1.17.2) -> 3.8.1 (1.18.1), each hop applied + verified supervised. atomic=true (auto-rollback on a failed rollout) + forceFailurePolicyIgnore (admissions stay open mid-roll) kept it safe. Values schema confirmed compatible across 3.6->3.8 (forceFailurePolicyIgnore still under features:). Verified after each hop: all 17 ClusterPolicies stayed Ready, admission controller 2/2, no destroys/replaces in plan. Final 1.18.1: images v1.18.1, mutating webhook live (server-side dry-run injects ndots:2 in a non-excluded ns). compat-gate vs 1.35.6 now lists ONLY external-secrets (kyverno cleared). ESO 0.12->2.x (v1beta1->v1, 73 files) is the last remaining 1.35 blocker — to be planned. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-21 17:21:38 +00:00
Viktor Barzin	ead876ec65	k8s-upgrade: nightly Slack report monitor + scope chain-failed alert to phases All checks were successful ci/woodpecker/push/default Pipeline was successful Details Adds a daily visibility layer so every night's autonomous-upgrade outcome is reviewable at a glance during the upgrade-cleanup window (Viktor: "track every night's upgrade for the next 7 days; clean up all bugs and blockers"). Last night (2026-06-20) confirmed BOTH prior fixes work in production: the detector resolved target 1.35.6 (k8s_upgrade_available) and the compat gate correctly REFUSED it (k8s_upgrade_blocked=1 -> K8sUpgradeBlocked) because ESO v0.12 (<=1.31) and kyverno v1.16 (<=1.34) don't support 1.35. What's here: - CronJob k8s-upgrade-nightly-report (06:07 UTC) -> one Slack summary/morning: running version, detector freshness, detected target, outcome (no-op / blocked+live reasons / upgraded / in-progress / detector-stale), recent jobs. Read-only: reads Pushgateway gauges + live nodes/jobs, re-runs compat-gate.py for fresh blockers; reuses the chain SA + slack_webhook + scripts ConfigMap. Pure helpers unit-tested (test_nightly_report.py, 8 cases incl. a real v-prefix bug TDD caught). Verified end-to-end in-cluster (posted to Slack). - K8sUpgradeChainJobFailed regex scoped from `k8s-upgrade-.` to `k8s-upgrade-(preflight\|master\|worker\|postflight)-.` so the new report job (or any future helper) can't false-trip the chain-wedged alarm. Manual state repair (no git artifact): imported the orphaned `alert-digest` CronJob into the monitoring stack state (`tg import module.monitoring.kubernetes_cron_job_v1.alert_digest monitoring/alert-digest`). Root cause: when alert_digest was added (2026-06-12) the apply recorded its ConfigMap + Secret but not the CronJob, so every full monitoring apply since has failed with `cronjobs.batch "alert-digest" already exists` (Woodpecker pipeline 298 today) — surviving only via targeted prometheus applies. Now in state, so monitoring CI applies cleanly again. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-21 16:57:44 +00:00
Viktor Barzin	7270e2be3b	monitoring: K8sUpgradeChainJobFailed must not double-fire on a compat-gate block Some checks failed ci/woodpecker/push/default Pipeline failed Details Last night (2026-06-20) the detector + compat-gate fixes worked: the chain resolved target 1.35.6 and the gate correctly REFUSED it (ESO 0.12 + kyverno 1.16 don't support 1.35), pushing k8s_upgrade_blocked=1 -> K8sUpgradeBlocked fired as designed. But the refusal also made the preflight Job exit 1 (block() exits 1 on purpose so the Failed Job re-spawns nightly), which tripped K8sUpgradeChainJobFailed too — a duplicate, misleading "pipeline wedged" alarm for what is the intended halt-and-alert outcome. Fix: gate the alert with `unless on() k8s_upgrade_blocked == 1`. A deliberate block sets that gauge (and it stays 1 until the next preflight resets it), so the chain-job-failed alert is suppressed for the blocked period; a genuine wedge / crash / halt-on-alert exits 1 WITHOUT setting it, so it still fires (preserving the alert's original purpose — catching the pre-in_flight preflight failure that hid the 5-day 1.34.9 wedge). Runbook + automated-upgrades docs updated to match. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-21 16:35:35 +00:00
Viktor Barzin	cc4bb8ffe8	wealth dashboard: show price freshness for all 3 holdings, not just worst Some checks failed ci/woodpecker/push/default Pipeline failed Details Viktor wanted the freshness tile to cover all three main holdings (META, VUAG, VUSA), not only the single stalest one. Dropped LIMIT 1 so the stat renders one value per held position (worst-first), switched the tile to horizontal orientation so the three values sit side-by-side, and updated the description. Each value is coloured by its own age threshold (META red ~2mo, the Vanguard ETFs green ~2d). No threshold or datasource change. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-21 14:49:33 +00:00
Viktor Barzin	c23b03864e	traefik/crowdsec: delete dead Yaegi plugin + middleware CRD + captcha (PR2/2) Zero live ingresses reference traefik-crowdsec@kubernetescrd (PR1 + a cluster-wide targeted ingress re-apply confirmed 0), so the crowdsec Middleware CRD and the broken Yaegi bouncer plugin can be removed without orphaning any router. Removes: the `crowdsec` Middleware, the crowdsec-bouncer plugin (static config + initContainer download + state.json entry), the captcha template ConfigMap + volume + captcha.html, the Turnstile widget + data.cloudflare_accounts, and the 3 now-unused module vars. Also drops the `crowdsec` middleware from the catch-all error-pages IngressRoute chain (the one remaining CRD-level reference, which an Ingress-annotation grep does not surface) so that router is not orphaned when the Middleware is deleted; it keeps rate-limit. Enforcement is fully handled out-of-band now: cs-firewall-bouncer (in-kernel nftables, direct hosts) + Cloudflare IP-List/WAF (proxied hosts). The api-token-middleware plugin is deliberately preserved (still used by paperless-mcp). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-21 13:35:13 +00:00
Viktor Barzin	68d9058f85	cleanup: fully remove orphaned council-complaints app The council-complaints app (Islington civic-reporting pilot) has been abandoned. It was already dead in the cluster (deployments scaled 0/0, image only on the decommissioned registry.viktorbarzin.me which 404s), and it was never in Terraform — only docs + a kyverno comment referenced it. Its live cluster resources (namespace, both NFS-backed PVs, ingresses) were torn down out-of-band via kubectl (nothing in TF to drift from); the DB-dump PVC was backed up to NFS first. This removes the remaining repo references to the live app: - service-catalog.md: drop the council-complaints row - ci-cd.md + .claude/CLAUDE.md: drop it from the GHA->ghcr app list - kyverno require-trusted-registries: the registry.viktorbarzin.me/* allowlist comment claimed council-complaints as the last referencer; rewrite it (no live workload pulls from that registry now; only stale completed Job records still carry the ref). The allowlist line itself is kept (registry-scoped, not app-specific). Historical point-in-time plan docs (docs/plans/2026-05-16-auto-upgrade- apps-{design,plan}.md) still mention it inside a frozen "10 GHA-migrated repos (memory id=388)" snapshot; left as-is so the dated record stays accurate. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-21 13:32:10 +00:00
Viktor Barzin	6dc3ce139f	wealth dashboard: expand all rows by default + inline the freshness stat Some checks failed ci/woodpecker/push/default Pipeline failed Details Two follow-ups Viktor asked for on the Price freshness panel: - Expand every section by default. Grafana's collapsed rows hide their child panels; just flipping collapsed=false leaves a non-canonical shape (confirmed via the Grafana API that it keeps the panels nested rather than hoisting them), so each row is now collapsed=false + panels=[] with its children hoisted to top-level -- the exact form Grafana writes when you expand-and-save. Row headers revert to their original y (the child y-coords were already expanded-layout coordinates). - Stop the freshness stat from taking its own line. It's now the 6th tile in the existing returns row (1d/7d/30d/90d/12mo + freshness), all width 4 at y=5; the collapsed-row y-shift from the previous commit is undone. No query or threshold changes. The large diff is mechanical: 12 child panels re-indent from nested to top-level. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-21 13:29:25 +00:00
Viktor Barzin	ddbdbca7e9	wealth dashboard: add "Price freshness" stat for stalest held quote Some checks failed ci/woodpecker/push/default Pipeline failed Details Viktor was worried about stale prices silently distorting net worth. Confirmed it's real: META's quote has been frozen at 2026-04-17 (65 days old) while the dashboard keeps valuing the ~55-share position at that stale close; the Vanguard ETFs are current. Nothing flagged it. Adds one compact stat to the Overview row showing the most out-of-date HELD position's quote age (symbol + humanised age), colour-coded: green <=4d (weekend/bank-holiday tolerant), amber 5-9d, red >=10d. Pure read of the quote_latest mirror via the wealth-pg datasource, held positions only, LEFT JOIN so a held symbol with no quote at all sorts as max-stale. The six collapsed rows below shift down 4 grid units to make room; no other panel touched. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-21 12:23:45 +00:00
Viktor Barzin	b1bbe42821	homelab ha token: dedicated openclaw/ha-tokens secret + least-priv RBAC for emo Some checks are pending Build infra CLI / build (push) Waiting to run Details ci/woodpecker/push/default Pipeline was successful Details `ha token` originally read openclaw/openclaw-secrets -> skill_secrets, which only cluster admins can read — so it hung/failed for the non-admin operator it was built for (emo = emil.barzin@gmail.com, OIDC group "Home Server Admins", whose identity is deliberately barred from secrets in the openclaw namespace). Split the HA tokens into a dedicated secret openclaw/ha-tokens (keys sofia/london) with a Role + RoleBinding granting `get` on JUST that secret to the Home Server Admins group (k8s RBAC can't scope to a JSON sub-key, hence a separate object). emo now resolves the HA token with their own identity, WITHOUT gaining the rest of skill_secrets (slack_webhook, uptime_kuma_password). openclaw's own deployment keeps reading openclaw-secrets — purely additive. - stacks/openclaw/ha_tokens.tf: new secret + least-privilege Role/RoleBinding - cli/cmd_ha.go: read openclaw/ha-tokens (raw base64 per-instance key); drop JSON parse - README + ADR-0012 updated; VERSION -> v0.7.1 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-21 10:45:32 +00:00
Viktor Barzin	71d0af084e	traefik/crowdsec: remove 6 hard-coded middleware refs the variable sweep missed (PR1/2) The first PR1 commit only dropped the ingress_factory reference + the 8 exclude_crowdsec call sites. But the crowdsec middleware is ALSO hard-coded (not via the variable) in 6 more ingresses that build their middleware chain by hand: owntracks, the monitoring Helm values (grafana + prometheus + alertmanager), and the reverse-proxy module + its own separate ingress factory. Remove all 6 so that after the full-cluster apply NO live ingress references traefik-crowdsec@kubernetescrd — the precondition for PR2 deleting the CRD. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-21 00:17:40 +00:00

1 2 3 4 5 ...

1551 commits