infra

Viktor Barzin 60b2b1cdfc cluster-health: emergency-stop Keel + roll back image downgrades + quota raises Keel was rewriting tag strings (not just digests) despite the keel.sh/match-tag=true annotation injected by the Kyverno inject-keel-annotations ClusterPolicy. That annotation was supposed to constrain Keel to digest-only watches under the deployment's CURRENT tag. It didn't. Casualties confirmed today (live image rewritten to a lower version): uptime-kuma (:2 → :1, 4h CrashLoopBackOff because v1 boots into SQLite mode and can't read the v2 db-config.json → MariaDB store); n8n (:1.80.5 → :0.1.2, silent — EEXIST mkdir /root/.n8n loop); beads-server/dolt-workbench (:0.3.73 → :0.1.0, GraphQL schema mismatch on addDatabaseConnection); wealthfolio (:3.2.1 → :2.0 → :3.2 string truncate); plus historical ones previously fixed (claude-memory :71b32438 → :17, forgejo 11.0.14 → 1.18, onlyoffice 9.3.1 → 4.0.0.9, shlink 5.0.2 → 1.16.1). Changes: * stacks/keel: replicaCount = 0 in the helm values. Pod went from 1/1 to 0/0. Keep off until either match-tag is root-caused or every enrolled workload migrates to a content-addressed (SHA) pin. * stacks/uptime-kuma: pin image to louislam/uptime-kuma:2.3.2 (was :2, bumped to :1 by Keel). Full opt-out: keel.sh/policy=never on BOTH the deployment label (matches Kyverno's exclude rule so the inject-keel- annotations ClusterPolicy stops mutating) AND the annotation (so Keel itself respects). Removed keel.sh/policy from lifecycle.ignore_changes so TF owns it as `never` and can't drift back to `force`. * stacks/beads-server: pin dolt-workbench to dolthub/dolt-workbench:0.3.73 on both seed-config and workbench containers (was :latest, Keel rolled to :0.1.0). * stacks/wealthfolio: pin to afadil/wealthfolio:3.2.1 (was :3.2 truncated by Keel from the prior live :3.2.1). * stacks/monitoring: monitoring-quota requests.memory 16Gi → 20Gi. Cluster grew from 5 to 7 workers (k8s-node5/6 added 2026-05-26) and alloy's per-pod request jumped 50Mi → 562Mi earlier today; combined with new-node DS pods (loki-canary, node-exporter, sysctl-inotify) the quota tipped to 100% and blocked every new pod create with FailedCreate. Raising the cap unblocked the four affected DaemonSets in one shot. * stacks/immich: tier-quota requests.memory 20Gi → 24Gi, limits.memory 32Gi → 40Gi. Was at 88% with VPA still creeping up on immich-server's face-detection burst behaviour. * stacks/{excalidraw,immich,n8n}: providers.tf + .terraform.lock.hcl updated by `tg init -upgrade` to record telmate/proxmox 3.0.2-rc07 (matches the 21 other stacks that already declare it). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>		2026-05-26 18:48:50 +00:00
..
dashboards	fire-planner: COL refresh CronJob + Grafana Cost-of-Living dashboard	2026-05-22 14:15:38 +00:00
server-power-cycle	Add broker-sync Terraform stack (#7 )	2026-04-17 21:17:45 +01:00
alloy.yaml	alloy: move resources to alloy.* (chart key bug); 1Gi limit fixes IO storm	2026-05-26 02:08:35 +00:00
Dockerfile	extract monitoring, nvidia, mailserver, cloudflared, kyverno from platform [ci skip]	2026-03-17 21:34:11 +00:00
goflow2.tf	[infra] Sweep dns_config ignore_changes across all pod-owning resources [ci skip]	2026-04-18 21:19:48 +00:00
grafana.tf	fire-planner: COL refresh CronJob + Grafana Cost-of-Living dashboard	2026-05-22 14:15:38 +00:00
grafana_chart_values.yaml	monitoring: protect grafana ingress with authentik + disable anonymous	2026-05-10 17:01:50 +00:00
idrac.tf	infra: document auth = "app\|none" tier on every legacy ingress	2026-05-11 19:25:48 +00:00
k8s-monitoring-values.yaml	cleanup: remove calibre and audiobookshelf stacks after ebooks migration [ci skip]	2026-03-25 23:56:07 +02:00
loki.tf	alloy: move resources to alloy.* (chart key bug); 1Gi limit fixes IO storm	2026-05-26 02:08:35 +00:00
loki.yaml	monitoring/loki: bump memory request 2Gi → 3Gi (close gap to 4Gi limit)	2026-05-24 01:10:55 +00:00
main.tf	cluster-health: emergency-stop Keel + roll back image downgrades + quota raises	2026-05-26 18:48:50 +00:00
prometheus.tf	fix: HA Sofia REST sensors + PVC drift safety	2026-05-10 21:48:29 +00:00
prometheus_chart_values.tpl	monitoring: alerts for proxmox-csi LUN saturation per node	2026-05-26 02:45:13 +00:00
prometheus_snmp_chart_values.yaml	extract monitoring, nvidia, mailserver, cloudflared, kyverno from platform [ci skip]	2026-03-17 21:34:11 +00:00
pve_exporter.tf	[infra] Sweep dns_config ignore_changes across all pod-owning resources [ci skip]	2026-04-18 21:19:48 +00:00
snmp_exporter.tf	infra: document auth = "app\|none" tier on every legacy ingress	2026-05-11 19:25:48 +00:00
ups_snmp_values.yaml	extract monitoring, nvidia, mailserver, cloudflared, kyverno from platform [ci skip]	2026-03-17 21:34:11 +00:00