infra

Author	SHA1	Message	Date
Viktor Barzin	2ab5b94748	t3-safe-restart: extract shared safe-restart library from t3-autoupdate Pull the per-unit backup->restart->verify->recover routine (and the small helpers it needs) out of t3-autoupdate.sh into a sourced library, so a second job (the upcoming idle migrator) can reuse the exact same audited recovery path instead of forking safety-critical code. safe_restart_unit returns non-zero on failure (after recovery+freeze) rather than exiting, so callers control flow. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-21 12:28:53 +00:00
Viktor Barzin	0cebeeb0ee	t3-idle-migrate: implementation plan Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-21 12:26:05 +00:00
Viktor Barzin	ddbdbca7e9	wealth dashboard: add "Price freshness" stat for stalest held quote Some checks failed ci/woodpecker/push/default Pipeline failed Details Viktor was worried about stale prices silently distorting net worth. Confirmed it's real: META's quote has been frozen at 2026-04-17 (65 days old) while the dashboard keeps valuing the ~55-share position at that stale close; the Vanguard ETFs are current. Nothing flagged it. Adds one compact stat to the Overview row showing the most out-of-date HELD position's quote age (symbol + humanised age), colour-coded: green <=4d (weekend/bank-holiday tolerant), amber 5-9d, red >=10d. Pure read of the quote_latest mirror via the wealth-pg datasource, held positions only, LEFT JOIN so a held symbol with no quote at all sorts as max-stale. The six collapsed rows below shift down 4 grid units to make room; no other panel touched. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-21 12:23:45 +00:00
Viktor Barzin	9503bed589	t3-idle-migrate: design for graceful overnight restart of deferred t3-serve instances Viktor hit the t3 'Client and server versions differ' warning. Root cause: the daily gated autoupdate defers a user's t3-serve restart whenever that user has an active agent at the 04:00 window, so anyone busy every night (long-lived/AFK sessions) never migrates and the client/server version skew persists for days. This design adds a small idle-gated overnight job that drains those deferrals -- restarting a deferred instance onto the current binary only when no turn is in flight (state.sqlite active_turn_id) and it's been quiet for a buffer, so the migration lands in a real quiet gap instead of killing in-flight agent turns. Reuses the autoupdate's proven backup->restart->verify->recover path via a shared helper (approach C from the brainstorm). Design doc only; no behavior change yet. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-21 12:04:22 +00:00
Viktor Barzin	b1bbe42821	homelab ha token: dedicated openclaw/ha-tokens secret + least-priv RBAC for emo Some checks are pending Build infra CLI / build (push) Waiting to run Details ci/woodpecker/push/default Pipeline was successful Details `ha token` originally read openclaw/openclaw-secrets -> skill_secrets, which only cluster admins can read — so it hung/failed for the non-admin operator it was built for (emo = emil.barzin@gmail.com, OIDC group "Home Server Admins", whose identity is deliberately barred from secrets in the openclaw namespace). Split the HA tokens into a dedicated secret openclaw/ha-tokens (keys sofia/london) with a Role + RoleBinding granting `get` on JUST that secret to the Home Server Admins group (k8s RBAC can't scope to a JSON sub-key, hence a separate object). emo now resolves the HA token with their own identity, WITHOUT gaining the rest of skill_secrets (slack_webhook, uptime_kuma_password). openclaw's own deployment keeps reading openclaw-secrets — purely additive. - stacks/openclaw/ha_tokens.tf: new secret + least-privilege Role/RoleBinding - cli/cmd_ha.go: read openclaw/ha-tokens (raw base64 per-instance key); drop JSON parse - README + ADR-0012 updated; VERSION -> v0.7.1 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-21 10:45:32 +00:00
viktor	a091689603	Merge pull request 'traefik/crowdsec: remove dead plugin middleware reference (PR1/2)' (#8 ) from wizard/cs-deplugin-refs into master Some checks failed ci/woodpecker/push/default Pipeline failed Details	2026-06-21 00:17:51 +00:00
Viktor Barzin	71d0af084e	traefik/crowdsec: remove 6 hard-coded middleware refs the variable sweep missed (PR1/2) The first PR1 commit only dropped the ingress_factory reference + the 8 exclude_crowdsec call sites. But the crowdsec middleware is ALSO hard-coded (not via the variable) in 6 more ingresses that build their middleware chain by hand: owntracks, the monitoring Helm values (grafana + prometheus + alertmanager), and the reverse-proxy module + its own separate ingress factory. Remove all 6 so that after the full-cluster apply NO live ingress references traefik-crowdsec@kubernetescrd — the precondition for PR2 deleting the CRD. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-21 00:17:40 +00:00
Viktor Barzin	7bd4612edf	ci: scripts/tg waits out a contended state lock (-lock-timeout) All checks were successful ci/woodpecker/push/default Pipeline was successful Details The infra CI pipeline was failing often — ~38% of the last 50 runs didn't succeed. The single biggest cause (8 of 19 non-successes) was Tier-1 stack applies dying instantly with "Error acquiring the state lock". Tier-0 stacks already degrade gracefully (Vault advisory lock → the pipeline skips a locked stack). Tier-1 stacks have no such fallback: they rely on terraform's pg-backend pg_advisory_lock, and scripts/tg ran terragrunt with no -lock-timeout, so any concurrent lock holder was fatal — a Woodpecker-killed run whose PG lock wasn't reaped yet (PL266 killed → PL267 failed the same second), a human/agent applying locally, or the daily drift `plan`. Fix: scripts/tg now passes -lock-timeout (default 5m, override TG_LOCK_TIMEOUT) on every state-locking verb (plan/apply/destroy/refresh), so a contended lock WAITS for the holder to finish instead of failing. -auto-approve behaviour for non-interactive applies is unchanged. Central wrapper change → covers CI, plus local human/agent applies; no CI image rebuild (tg is read from the repo). Adds a hermetic pytest (stub terragrunt + preset PG_CONN_STR) pinning the arg-injection. Docs updated in AGENTS.md + .claude/CLAUDE.md. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-21 00:15:39 +00:00
Viktor Barzin	84a18a5529	traefik/crowdsec: remove dead Yaegi-plugin middleware reference (PR1/2) The Traefik CrowdSec (Yaegi) bouncer plugin enforces nothing on Traefik 3.7.5 (handler never invoked) and is fully superseded by the cs-firewall-bouncer (in-kernel nftables drop on direct hosts) + the Cloudflare IP-List/WAF rule (proxied hosts). Drop the `traefik-crowdsec@kubernetescrd` middleware from the ingress_factory chain and the 8 explicit `exclude_crowdsec = true` call sites, and delete the now-unused `exclude_crowdsec` variable. This is PR1 of a 2-phase removal: the reference is removed FIRST (a shared-module change → full-cluster apply re-renders every ingress without the middleware) so that PR2 can delete the `crowdsec` Middleware CRD + the plugin itself WITHOUT leaving any ingress pointing at a missing middleware (which would error those routers). PR2 MUST NOT land until this has fully applied and zero live ingresses reference traefik-crowdsec@kubernetescrd. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-21 00:15:12 +00:00
viktor	9774ae3d19	Merge pull request 'crowdsec: firewall-bouncer cluster-wide (remove node2 pin)' (#7 ) from wizard/cs-fw-allnodes into master All checks were successful ci/woodpecker/push/default Pipeline was successful Details	2026-06-21 00:08:15 +00:00
Viktor Barzin	c92590ae85	crowdsec: roll firewall-bouncer cluster-wide (remove node2 validation pin) One-node validation on k8s-node2 passed: kernel nftables sets created in both input and forward chains (policy accept), ~31k decisions loaded, a known banned scanner confirmed in the drop set, pod stable 4h+ with no collateral. Remove the nodeSelector so the DaemonSet runs on every node — direct-host enforcement now survives a MetalLB VIP failover to any worker. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-21 00:07:45 +00:00
viktor	4f1c998468	Merge pull request 'rybbit sync: exclude CAPI + per_page=500 fix' (#6 ) from wizard/crowdsec-syncfix into master All checks were successful ci/woodpecker/push/default Pipeline was successful Details	2026-06-21 00:05:50 +00:00
Viktor Barzin	f55bb6c422	rybbit: sync excludes CAPI blocklist + fix CF items per_page (500) The edge CF IP List can't hold the ~31k CAPI community blocklist (already enforced in-kernel by the firewall-bouncer), so the sync now skips origin=CAPI and carries only high-signal local/curated decisions (+ a 9000 safety cap). Also fixes the list-items GET: per_page=1000 returned a misleading CF 400 'invalid or expired cursor' (10027); the endpoint max is 500. Verified live: crowdsec_ban populates (4 IPs) and the sync exits 0. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-21 00:05:05 +00:00
Viktor Barzin	6d5d3726d6	Merge remote-tracking branch 'origin/master' into wizard/ha-cli-verbs Some checks are pending Build infra CLI / build (push) Waiting to run Details ci/woodpecker/push/default Pipeline was successful Details	2026-06-20 23:46:29 +00:00
Viktor Barzin	48225f2dea	homelab CLI v0.7: add `ha token` + `ha ssh` for Home Assistant Mined another devvm user's Claude sessions for repeated, hand-rolled command patterns worth absorbing into the shared CLI. The dominant signal was Home Assistant "Sofia" work: a `kubectl \| base64 \| jq` token-extraction pipeline re-derived ~420x, and a bespoke non-interactive `ssh -o …` invocation reinvented ~30x — every session. The existing `home-assistant-sofia.py` already covers the API but goes unused from an arbitrary cwd (needs an env var set + a cwd-relative path), so agents bypassed it and hand-rolled everything. Add two verbs covering exactly the gaps the `ha` MCP can't (entity state/control stays with the MCP): - `ha token [--instance sofia\|london]` (read): resolves the long-lived API token live from k8s secret openclaw/openclaw-secrets via the ambient kubeconfig — no pre-set env var. Composes as `curl -H "Authorization: Bearer $(homelab ha token)"`. - `ha ssh [--instance sofia\|london] -- <cmd>` (write): deterministic non-interactive ssh to the HA host using the invoking user's key. Also fix the root cause: `home-assistant-sofia.py` now falls back to `homelab ha token` when its env var is unset (works from any directory), and the home-assistant skill points agents at these verbs + `homelab metrics query` instead of hand-rolled curls. README + ADR-0012 + AGENTS.md updated per the per-verb-group convention. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-20 23:46:09 +00:00
Viktor Barzin	46166c63b2	fix(authentik): long-lived social-login sessions + shield auth from CrowdSec lockout All checks were successful ci/woodpecker/push/default Pipeline was successful Details Viktor's passkeys all vanished and he was suddenly being asked to log in multiple times a day instead of ~monthly. Root cause: on 2026-06-18 an ad-hoc tripit passkey E2E test (run from the devvm as akadmin via python-httpx) cleaned up "the demo user's" passkeys with GET /core/users/?search={demo} then DELETE each device of users[0] — but the fuzzy search returned the REAL account, so it wiped all 6 real passkeys. Losing passkeys forced fallback to Google login, and the social-login stage (default-source-authentication-login) had the provider default session_duration=seconds=0, which falls back to UNAUTHENTICATED_AGE=2h — hence the constant re-logins. (Password + passkey logins were already weeks=4.) Changes: - authentik: adopt default-source-authentication-login into Terraform (import) and pin session_duration=weeks=4, so Google/GitHub/Facebook logins last as long as password/passkey. Immediate relief without re-enrolling. - authentik: document the provider-schema gotcha — authentik_stage_identification exposes no webauthn_stage / enable_remember_me attribute, so they must NOT be in ignore_changes (commit `4e882989` removed them for this reason; re-adding breaks every apply). The passkey break was purely the missing device records, not drift. - edge (rybbit): shield auth so a CrowdSec hit can never wall a user out of login — carve authentik.viktorbarzin.me + public-auth out of the zone WAF block rule, make the LAPI->edge sync ban-only (stop downgrading captcha to a hard block), and set exclude_crowdsec on the Authentik UI ingress (auth keeps rate-limiting). - docs: record the session-duration change, the edge enforcement + auth carve-out (previously undocumented), and the pre-existing broken crowdsec-cf-sync CronJob (CF cursor pagination 400 + ~31k IPs vs list capacity -> edge list inert). Passkey re-enrollment is a manual user action (devices are gone from the DB); nothing auto-re-deletes them. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-20 23:40:22 +00:00
Viktor Barzin	600f1f933c	Create Claude auth state directories All checks were successful ci/woodpecker/push/default Pipeline was successful Details The first live renewal run showed systemd could not create state beneath a read-only home sandbox. Provision each user's writable state directory before enabling the timer so automatic renewal can run.	2026-06-20 20:25:55 +00:00
Viktor Barzin	7f1788a106	Merge remote-tracking branch 'origin/master' into wizard/claude-auth-renew All checks were successful ci/woodpecker/push/default Pipeline was successful Details	2026-06-20 20:22:20 +00:00
Viktor Barzin	ff67e9d422	Fix workstation package manifest parsing The approved Claude token renewal deployment could not run because setup-devvm passed inline package comments to apt as package names. Strip inline comments so the persisted all-user setup remains reproducible.	2026-06-20 20:22:05 +00:00
Viktor Barzin	524b874036	state(vault): update encrypted state Some checks failed ci/woodpecker/push/default Pipeline was canceled Details	2026-06-20 20:14:53 +00:00
Viktor Barzin	7050b0441e	Merge remote-tracking branch 'origin/master' into wizard/claude-auth-renew Some checks failed ci/woodpecker/push/default Pipeline failed Details	2026-06-20 20:11:09 +00:00
Viktor Barzin	bc2fbc712c	Merge remote-tracking branch 'origin/master' into wizard/claude-auth-renew	2026-06-20 20:10:48 +00:00
Viktor Barzin	02d14796cc	feat(mailserver): add trips@ send-as alias for TripIt native auth email (ADR-0028) Some checks failed ci/woodpecker/push/default Pipeline was canceled Details TripIt's native signup-verification + account-recovery mail (ADR-0028) sends From: trips@viktorbarzin.me while authenticating SMTP as spam@. With SPOOF_PROTECTION on, Postfix smtpd_sender_login_maps requires an EXPLICIT alias (the @domain catch-all doesn't satisfy it) — mirrors the existing plans@->spam@ grant. Must be applied + verified before TripIt flips SMTP_FROM to trips@, else every verification/recovery send is rejected 550. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-20 20:10:47 +00:00
Viktor Barzin	5549fc3672	Add per-user Claude auth renewal Each workstation user needs a continuously valid Claude token under their own Enterprise identity. Store only that user's OAuth state in an isolated Vault path, renew and verify it automatically, recover from Vault when possible, and alert when interactive SSO is required.	2026-06-20 20:10:40 +00:00
Viktor Barzin	3278588325	chore(authentik): tear down obsolete tripit-enrollment (ADR-0020 superseded by ADR-0028) All checks were successful ci/woodpecker/push/default Pipeline was successful Details TripIt external users are now LOCAL TripIt accounts (ADR-0028 native passkey + Authentik OIDC), so the Authentik-side self-enrollment machinery is dead. Removes the tripit-enrollment + tripit-recovery flows and all their stages/prompts/policies/bindings, the tripit-email-stages blueprint (+yaml), and the 'TripIt External' group; reverts the admin-services-restriction fence branch that contained those users (its sole member, the leftover tripit-demo@ test account, was deleted first, so the revert affects zero live principals). Real external collaborators (type=external) are untouched. tg plan: 0 add, 1 change (the policy expression), 20 destroy (all tripit_*). Closes tripit#97; moots the B2 per-app OIDC fences. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-20 20:04:24 +00:00
viktor	834c5e6a2a	Merge pull request 'CrowdSec proxied: single CF list (block-only) + firewall-bouncer re-apply' (#5 ) from wizard/crowdsec-1list into master Some checks failed ci/woodpecker/push/default Pipeline failed Details	2026-06-20 19:31:01 +00:00
Viktor Barzin	7cf93a0587	crowdsec+rybbit: proxied edge to single CF list (block-only) + retrigger firewall-bouncer apply CF account hard-limits to 1 Rules List, so proxied enforcement uses one crowdsec_ban list + one WAF block rule; the sync writes both ban and captcha decisions into it (captcha downgraded to block at the edge). Drops the second list + managed_challenge rule. Trivial touch to firewall_bouncer.tf to make CI re-apply crowdsec and recreate the DaemonSet (tar fix already in master; stale orphan was cleared). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-20 19:29:43 +00:00
viktor	1406d8a391	Merge pull request 'Fix CF ruleset import id + depends_on' (#4 ) from wizard/crowdsec-fix2 into master Some checks failed ci/woodpecker/push/default Pipeline failed Details	2026-06-20 19:13:03 +00:00
Viktor Barzin	f2b089e267	rybbit: fix cloudflare_ruleset import id (zone/ 3-part form) + depends_on lists v4.52.7 import id must be zone/<zone_id>/<ruleset_id>; add depends_on so the crowdsec_ban/captcha lists exist before the WAF rules reference them. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-20 19:12:29 +00:00
viktor	58fc6d5061	Merge pull request 'Fix CrowdSec firewall-bouncer tar + CF WAF ruleset import' (#3 ) from wizard/crowdsec-fixes into master Some checks failed ci/woodpecker/push/default Pipeline failed Details	2026-06-20 19:06:15 +00:00
Viktor Barzin	a351a66843	crowdsec+rybbit: fix firewall-bouncer tar extraction (busybox) + import existing CF WAF ruleset - initContainer used GNU tar --wildcards which fails on the busybox curl image (pod Init:Error); switch to extract-all + cp via shell glob. - cloudflare_ruleset hit the per-zone singleton conflict; import the existing 'default' http_request_firewall_custom ruleset and manage all rules — CrowdSec ban/captcha first, the pre-existing disabled skip rule preserved verbatim. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-20 19:04:30 +00:00
viktor	70e8ce1021	Merge pull request 'CrowdSec real enforcement: edge WAF (proxied) + firewall-bouncer (direct)' (#2 ) from wizard/crowdsec-enforcement into master Some checks failed ci/woodpecker/push/default Pipeline failed Details	2026-06-20 09:42:41 +00:00
Viktor Barzin	ca8d617e72	rybbit: use 'Account Rule Lists' permission group for the CF sync token (v4) tg plan verified the agent's guess 'Account Filter Lists Edit/Read' is not a key in the v4.52.7 permission-group map; the live CF API lists the correct account-scoped groups as 'Account Rule Lists Read'/'Write'. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-20 09:41:41 +00:00
Viktor Barzin	0c56290af0	chore(forgejo): re-trigger apply of git.timeout/gc.auto (changed-stack skip) All checks were successful ci/woodpecker/push/default Pipeline was successful Details `910d5892` landed the [git.timeout] + [git.config] env in master, but the CI apply skipped stacks/forgejo (the changed-stack-diff race after a sync-merge), so the Forgejo deployment never picked it up. A trivial comment touch to force a clean apply of the stack so the durable push-mirror fix actually takes effect. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-20 09:19:53 +00:00
Viktor Barzin	cc4bfb593b	rybbit: proxied CrowdSec enforcement via Cloudflare IP Lists + WAF rule Replaces the Worker+KV approach (which only covered the ~27 routed hosts) with a zone-wide mechanism that covers ALL proxied hosts: two CF account IP Lists (crowdsec_ban, crowdsec_captcha) + one zone WAF custom rule that blocks `(ip.src in $crowdsec_ban)` and managed-challenges `(ip.src in $crowdsec_captcha)`. No per-request Worker, no cookie machinery — the rybbit Worker stays analytics-only. lapi_kv_sync.py now full-reconciles the two lists from LAPI (fail-safe: a LAPI blip skips the run and freezes the last-known-good block set; serializes CF bulk ops since CF allows one pending op per account). A least-privilege CF API token (Account Filter Lists Edit) is minted in TF. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-20 09:18:33 +00:00
Viktor Barzin	7e646e1c7c	crowdsec: add cs-firewall-bouncer DaemonSet (direct-host nftables enforcement) Drops banned source IPs in-kernel via nftables (hooks input+forward, so DNAT'd LoadBalancer traffic is caught before reaching Traefik) for DIRECT hosts — the direct-side replacement for the dead Traefik plugin, zero per-request hop. No published image exists, so an initContainer fetches the pinned official static binary (v0.0.34) onto a stock debian-slim base (nftables backend uses netlink directly, no nft CLI needed). hostNetwork + NET_ADMIN/NET_RAW (not privileged). Config (with api_key) in a Secret, Reloader-annotated. crowdsec ns is already in the Kyverno wave-1 exclude list, so the privileged/hostNetwork pod is admitted. Pinned to k8s-node2 (runs a Traefik pod) for one-node validation before the nodeSelector is removed to roll cluster-wide. Fail-open by element timeout if the bouncer stops. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-20 09:11:08 +00:00
Viktor Barzin	53117b193a	portal-realtime: deploy the v2 full-duplex voice agent (Pipecat) All checks were successful ci/woodpecker/push/default Pipeline was successful Details New stack for the realtime voice agent — v2 of the portal-assistant brain path. One persistent WebSocket per conversation: continuous mic audio -> Silero VAD turn-taking -> Whisper STT (portal-stt) -> streaming Claude brain (claude-agent-service) -> edge-tts (portal-tts) -> audio out, with barge-in. Reuses all three upstream cluster services; nothing new is spun up. Public Cloudflare ingress (proxied, WebSocket) at portal-realtime.viktorbarzin.me with the app's own DEVICE_TOKEN as the edge gate (auth="app" — Authentik would break the native Portal client). No buffering middleware: it would break the streaming WebSocket. Image ghcr.io/viktorbarzin/portal-assistant-realtime (private ghcr, pulled with ghcr_pull_token). Sibling to the v1 portal-assistant gateway, which stays live. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-20 08:23:17 +00:00
Viktor Barzin	44cac6f4e2	gitignore: ignore Python test artifacts (__pycache__, *.pyc, .pytest_cache) All checks were successful ci/woodpecker/push/default Pipeline was successful Details Introduced the first pytest file in the tree (stacks/k8s-version-upgrade/scripts/test_compat_gate.py); running it leaves an untracked __pycache__/ dir. Ignore the standard Python build artifacts so test runs don't show up as working-tree noise or get committed by accident. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-20 08:17:03 +00:00
Viktor Barzin	b58fe8cb1a	docs(k8s-upgrade): record detector Packages-probe -L fix + compat-gate patch scope All checks were successful ci/woodpecker/push/default Pipeline was successful Details Two corrections to the runbook matching today's code fixes: - The next-minor patch probe (GET .../Packages) also needs `-L`; it lacked it until 2026-06-20 and silently no-op'd the 2026-06-19 nightly run. Both probes now follow the 302. - The compat gate's addon check is scoped to minor jumps — patches within the running minor are never addon-blocked (target_minor <= running_minor returns early), so a conservative ceiling like ESO 0.12 -> 1.31 no longer false-blocks a 1.34.x patch. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-20 08:16:20 +00:00
Viktor Barzin	e5250f417e	k8s-version-upgrade: compat gate must not false-block patch upgrades All checks were successful ci/woodpecker/push/default Pipeline was successful Details The compat gate compared every addon's matrix ceiling against the target k8s minor unconditionally. That is correct for a minor JUMP, but it also blocked patch upgrades within the minor the cluster is ALREADY running: ESO v0.12's matrix ceiling is 1.31, the cluster runs 1.34.9, so a target of 1.34.10 (a patch) was refused with "external-secrets supports k8s <= 1.31; target 1.34 exceeds it" — even though the running cluster is itself proof ESO 0.12 works on 1.34. That silently defeats autonomous patching (it would have bitten the moment a 1.34.10 was published). Fix: a target at or below the running minor crosses into no new k8s minor, so every installed addon is already empirically proven on it — check_addons now returns no reasons when target_minor <= running_minor. Added running_minor() (oldest kubelet across nodes, mirroring the detector; RUNNING_K8S env override for tests) and pass it in. Minor jumps are unchanged: 1.34->1.35 still blocks on ESO 0.12 + kyverno 1.16. removed-API + containerd checks are naturally inert for patches (no API removal / containerd floor inside a minor) and keep running as defence. Added test_compat_gate.py (8 cases) covering both paths. Verified end-to-end against live Prometheus: target 1.34.10 -> EXIT 0 (safe), target 1.35.6 -> EXIT 2 (blocked on ESO+kyverno). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-20 08:14:50 +00:00
Viktor Barzin	38675b7922	crowdsec: register kvsync + firewall bouncer keys in LAPI Seeds two new bouncers at LAPI startup (BOUNCER_KEY_kvsync, BOUNCER_KEY_firewall) from Vault secret/platform, mirroring the existing BOUNCER_KEY_traefik wiring. These are the two halves of the real enforcement that replaces the dead Yaegi plugin: kvsync authenticates the LAPI->Cloudflare-KV sync (proxied edge Worker), firewall authenticates the cs-firewall-bouncer DaemonSet (direct-host nftables). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-20 08:12:38 +00:00
Viktor Barzin	a9384a4067	Merge remote-tracking branch 'origin/master' Some checks failed ci/woodpecker/push/default Pipeline failed Details	2026-06-20 08:09:16 +00:00
Viktor Barzin	44a98d408e	k8s-version-upgrade: detector next-minor probe must follow 302 (curl -sfL) The next-minor Packages query used `curl -sf` without -L. pkgs.k8s.io 302-redirects every request to a backing host, so without -L curl returned an empty body, NEXT_MINOR_PATCH came back empty, and the detector fell through to "No upgrade needed". That is exactly why last night's 23:00 chain no-op'd instead of resolving the 1.35 next-minor target (1.35.6) and handing it to the compat gate. `curl -sfL` follows the redirect and returns the Packages file (verified: -sf -> empty, -sfL -> 1.35.6). Mirrors the same -L fix already applied to the Release availability probe (-sILo) above. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-20 08:09:08 +00:00
Viktor Barzin	910d589205	fix(forgejo): raise git-op timeouts + lower gc.auto to stop push-mirror timeouts Some checks failed ci/woodpecker/push/default Pipeline was canceled Details The tripit Forgejo->GitHub push-mirror silently stalled: `git cat-file --batch-all-objects` over the NFS-backed repo exceeded the default git deadline once ~4500 loose objects accumulated (gc.auto's 6700 threshold hadn't fired), so pushes stopped reaching GitHub and prod deploys stalled. Raise [git.timeout] (DEFAULT/MIRROR/GC) so a slow object enumeration can't abort the mirror, and set [git.config] gc.auto=1000 so post-push autogc + the git_gc_repos cron keep repos packed (the real fix). A one-off forced gc already unblocked tripit; this prevents recurrence across all repos. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-20 08:08:50 +00:00
Viktor Barzin	45bed1c133	Merge remote-tracking branch 'origin/master' All checks were successful ci/woodpecker/push/default Pipeline was successful Details	2026-06-20 08:07:23 +00:00
Viktor Barzin	e1736d2e5c	calico: hop 3.28.5->3.30.7 (operator v1.38.13) — restores a SUPPORTED Calico/k8s-1.34 pairing. Disabled new-in-3.30 Goldmane/Whisker (their CRs render before crds/ install on helm upgrade; we use Prometheus/Loki). calico-node 7/7 on quay/v3.30.7, tigerastatus green. Applied manually + verified overnight.	2026-06-20 08:07:08 +00:00
Viktor Barzin	4d9fdbc7f7	rybbit: add CrowdSec LAPI -> Cloudflare KV sync script (proxied edge control plane) Pure-stdlib script (alert_digest pattern, runs on stock python:3.12-alpine) that projects CrowdSec Ip-scope ban/captcha decisions into the Workers KV namespace the edge Worker reads on each proxied request. Full-reconcile per run so an un-ban clears from the edge within one interval; fail-safe (a LAPI read error skips the run and leaves existing bans to expire by TTL = fail-open, never a stale all-block). TF wiring (KV namespace + CronJob + key registration) follows. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-20 08:05:11 +00:00
Viktor Barzin	0ac176da01	crowdsec: whitelist internal/LAN/tailnet CIDRs at the decision layer Preparing for real CrowdSec enforcement (edge Cloudflare Worker for proxied hosts + cs-firewall-bouncer for direct hosts). Both enforce by dropping the real source IP, so if an internal/RFC1918 address ever ended up in a ban decision it could blackhole legitimate internal traffic. Whitelisting the cluster/LAN/tailnet ranges (10/8, 172.16/12, 192.168/16, 100.64/10) at the CrowdSec parser layer makes that structurally impossible — a trusted source can never produce a decision in the first place. Public IP already whitelisted. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-20 08:03:46 +00:00
Viktor Barzin	3e3fdb34f0	homelab: v0.6.0 — usage telemetry (usage top), evidence-driven verb prioritization Some checks are pending Build infra CLI / build (push) Waiting to run Details ci/woodpecker/push/default Pipeline was successful Details Answers the question that drove the whole CLI — which verbs to add next — with data instead of one maintainer's habits, and resolves the cross-user-usage ask in-bounds (no reading anyone's home). - emit on dispatch: every verb fire-and-forgets one Loki line {job,user,verb} + "exit=N ver=X". ONLY the verb path + exit code — never args, paths, flags, or secrets (the emit never sees arguments). Best-effort: 800ms timeout, errors swallowed, never affects the command; opt-out HOMELAB_TELEMETRY=0. Discovery verbs (manifest/version/help) and usage itself don't self-record. - usage top [--since 30d] [--user U] [--json]: ranks verbs via sum by (verb)(count_over_time({job="homelab-usage"}[…])) against the shared Loki. Cross-user analytics WITHOUT touching ~/.claude — the privacy-preserving answer to "what does the team use". - Loki sink (zero new infra, dogfoods v0.5 logs path); push verified HTTP 204 no auth. ADR docs/adr/0011. Live-verified: ran 4 verbs, usage top ranked them correctly (metrics query=2). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>	2026-06-19 22:29:01 +00:00
Viktor Barzin	666fefd22b	calico: hop 3.26->3.28.5 (operator v1.34.13); calico-node 7/7 healthy, tigerastatus green, kube-controller-manager restarted (3.28 UID change). Applied manually + verified. Some checks failed ci/woodpecker/push/default Pipeline failed Details	2026-06-19 22:09:23 +00:00

1 2 3 4 5 ...

4497 commits