# Garage metrics + monitoring setup (2026-08-30 session) ## Working recipe: enable Prometheus /metrics on Garage v2.1 Root cause of the "Unsupported authorization method" mystery: Garage v2.x serves the admin API (metrics + admin endpoints) on a SEPARATE port configured via `[admin] api_bind_addr`. Without it, admin requests collide with the S3 parser on port 3900 and always return S3-style XML errors — even with a valid `Authorization: Bearer ` header. The earlier "no working HTTP admin call was found" section of SKILL.md was WRONG; the HTTP admin API works fine once the separate port exists. Config change per node (in /opt/garage/garage.toml, `[admin]` section): ```toml [admin] api_bind_addr = ":3903" # 10.8.0.1 (vps01), 10.8.0.2 (bigbox), 10.8.0.4 (vps02) metrics_token = "" # equals admin_token in this cluster admin_token = "" ``` - `api_bind_addr` in `[s3_api]` is a DIFFERENT key — anchor patches on the `[admin]` header when replacing, or the s3 entry gets found first. - Bind to the WG IP (not 0.0.0.0) so metrics don't leak to the internet. Verification: - `garage admin-token list` (via --entrypoint CLI, meta mounted) shows `metrics_token (from daemon configuration) ... Scope: Metrics`. - `curl -H "Authorization: Bearer " http://:3903/metrics` → Prometheus text (~90-110 KB per node), self-documented `# HELP` lines. - `curl http://:3903/health` → 200 if quorum, 503 otherwise. ## Rollout procedure (worked) 1. Patch config on all 3 nodes (vps02 needs `sudo`: file owned by root; vps01 too). 2. Restart ONE node at a time with ~8s pauses, order vps02 → vps01 → bigbox, `docker compose restart garage` (config re-read at container start). 3. After all restarts, `garage status` must show HEALTHY NODES (all 3 up, v2.1.0). ## Firewall gotchas per host - vps01: ufw active, policy DROP. New port needed `sudo ufw allow from 10.8.0.0/24 to any port 3903 proto tcp`. Symptom without it: WG ping OK but TCP connect hangs (timeout) — RPC/other ports were already allowed, 3903 wasn't. - vps02: nftables, INPUT policy ACCEPT — no rule needed. - bigbox: localhost/WG direct, no firewall issue observed. ## Monitoring stack targets (prometheus.yml) - `garage` job: 10.8.0.1:3903, 10.8.0.2:3903, 10.8.0.4:3903, path /metrics, bearer auth with the shared token. - `garage_health` job: same targets, path /health. - `node` job: 10.8.0.2:9100 (bigbox), 10.8.0.1:9100, 10.8.0.4:9100 (node-exporter, not yet deployed as of session end). - Grafana on :3001 (3000 occupied by gitea). Loki :3100, promtail via docker socket. - Alerts (alerts.yml): GarageNodeDown (up==0 for 2m), GarageNoQuorum (<2 of 3 up for 2m), resync error counter > 0, RPC node health == 0. ## Environment notes - bigbox cannot resolve/reach general web (DNS fails for dl.garagehq.de, github.com raw is reachable; Firecrawl web tools not configured → web_search returns an error about FIRECRAWL_API_KEY). Authoritative docs: git.deuxfleurs.fr (Deuxfleurs/garage, branch main-v1) — `doc/book/reference-manual/admin-api.md` and `doc/book/reference-manual/configuration.md`. - `docker exec` into garage container impossible (scratch image). Use --entrypoint CLI. - cli/ dir in /opt/garage is a 10-byte "Not Found" placeholder — ignore.