infra

Viktor Barzin f201e4573e immich: fix slow context search — prewarm clip_index + latency alert/healthcheck Context (smart) search latency was caused by the 665MB vchord clip_index decaying out of PG shared_buffers (~33% resident -> ~1.8s cold ANN reads vs ~4ms warm), NOT by yesterday's ML MODEL_TTL/clip-keepalive change (CLIP textual is warm ~15ms on GPU). The postStart prewarm runs once at pod start and pg_prewarm.autoprewarm only re-warms at startup, so the index decays under job buffer-pressure over days. - clip-index-prewarm CronJob (immich, /5): pg_prewarm('clip_index') keeps the whole index resident -> searches stay ~4ms. - immich-search-probe CronJob (immich, /5): times a random-vector ANN query + reads clip_index residency, pushes gauges to the Pushgateway. - Prometheus alerts ImmichSmartSearchSlow / ImmichClipIndexColdCache / ImmichSearchProbeStale (+ inhibition when the probe is stale). - cluster_healthcheck.sh check #46 check_immich_search (TOTAL_CHECKS 45->46). - Docs: infra CLAUDE.md immich note, monitoring.md, cluster-health skill. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>		2026-06-05 09:19:07 +00:00
..
agent-task-tracking.md	Add agent task tracking documentation	2026-04-15 17:11:26 +00:00
authentication.md	docs/auth: sync to current `auth` enum (required/app/public/none)	2026-05-11 19:28:42 +00:00
automated-upgrades.md	claude-agent-service: wire parallel execution (git-crypt mount, memory, MAX_CONCURRENCY)	2026-06-03 10:24:24 +00:00
backup-dr.md	backup: stop offsite-copying regenerable data; shrink nextcloud backup; pin nextcloud image	2026-06-01 15:15:26 +00:00
chrome-service.md	chrome-service: open NP for Traefik → noVNC sidecar (port 6080)	2026-05-07 18:40:11 +00:00
ci-cd.md	[forgejo] Phases 3+4+5: cutover, decommission, docs sweep	2026-05-07 18:30:02 +00:00
compute.md	immich: GPU-accelerate video transcoding (NVENC + NVDEC)	2026-05-29 18:05:34 +00:00
databases.md	redis: revert 3-node Sentinel HA to single standalone instance [ci skip]	2026-05-30 17:49:43 +00:00
dns.md	phpipam-pfsense-import: every 5min → hourly	2026-04-26 22:48:43 +00:00
homepage.md	add homepage auto-discovery documentation [ci skip]	2026-03-25 13:06:43 +02:00
incident-response.md	[claude-agent-service] Migrate all pipelines from DevVM SSH to K8s HTTP	2026-04-18 10:12:02 +00:00
llama-cpp.md	immich: set MACHINE_LEARNING_MODEL_TTL 0->600 to stop GPU VRAM hog	2026-06-02 20:16:11 +00:00
mailserver.md	monitoring: bring EmailRoundtripStale threshold docs in sync with for:20m	2026-04-21 22:39:46 +00:00
monitoring.md	immich: fix slow context search — prewarm clip_index + latency alert/healthcheck	2026-06-05 09:19:07 +00:00
multi-tenancy.md	add architecture documentation for all infrastructure subsystems [ci skip]	2026-03-24 00:55:25 +02:00
networking.md	technitium: CoreDNS rewrite forgejo.viktorbarzin.me -> Traefik ClusterIP	2026-06-04 07:34:30 +00:00
overview.md	gpu: schedule off NFD label, not k8s-node1 hostname	2026-04-22 13:43:07 +00:00
secrets.md	docs: comprehensive audit and update of all architecture docs and runbooks [ci skip]	2026-04-06 13:21:05 +03:00
security.md	security(wave1): W1.7 analysis snapshot — observation data → allowlist plan	2026-05-22 15:22:25 +00:00
storage.md	storage docs: document the per-VM SCSI-LUN cap (proxmox-csi)	2026-05-26 02:56:27 +00:00
vpn.md	docs: Technitium DNS IP — 10.0.20.101 → 10.0.20.201	2026-05-23 08:53:52 +00:00
wave1-egress-observation-2026-05-22.md	security(wave1): W1.7 analysis snapshot — observation data → allowlist plan	2026-05-22 15:22:25 +00:00