Files
garage-s3-cluster/references/metrics-and-monitoring.md
T
2026-09-06 13:50:54 +00:00

3.3 KiB

Garage metrics + monitoring setup (2026-08-30 session)

Working recipe: enable Prometheus /metrics on Garage v2.1

Root cause of the "Unsupported authorization method" mystery: Garage v2.x serves the admin API (metrics + admin endpoints) on a SEPARATE port configured via [admin] api_bind_addr. Without it, admin requests collide with the S3 parser on port 3900 and always return S3-style XML errors — even with a valid Authorization: Bearer <admin_token> header. The earlier "no working HTTP admin call was found" section of SKILL.md was WRONG; the HTTP admin API works fine once the separate port exists.

Config change per node (in /opt/garage/garage.toml, [admin] section):

[admin]
api_bind_addr = "<WG-IP>:3903"   # 10.8.0.1 (vps01), 10.8.0.2 (bigbox), 10.8.0.4 (vps02)
metrics_token = "<token>"        # equals admin_token in this cluster
admin_token = "<token>"
  • api_bind_addr in [s3_api] is a DIFFERENT key — anchor patches on the [admin] header when replacing, or the s3 entry gets found first.
  • Bind to the WG IP (not 0.0.0.0) so metrics don't leak to the internet.

Verification:

  • garage admin-token list (via --entrypoint CLI, meta mounted) shows metrics_token (from daemon configuration) ... Scope: Metrics.
  • curl -H "Authorization: Bearer <token>" http://<node>:3903/metrics → Prometheus text (~90-110 KB per node), self-documented # HELP lines.
  • curl http://<node>:3903/health → 200 if quorum, 503 otherwise.

Rollout procedure (worked)

  1. Patch config on all 3 nodes (vps02 needs sudo: file owned by root; vps01 too).
  2. Restart ONE node at a time with ~8s pauses, order vps02 → vps01 → bigbox, docker compose restart garage (config re-read at container start).
  3. After all restarts, garage status must show HEALTHY NODES (all 3 up, v2.1.0).

Firewall gotchas per host

  • vps01: ufw active, policy DROP. New port needed sudo ufw allow from 10.8.0.0/24 to any port 3903 proto tcp. Symptom without it: WG ping OK but TCP connect hangs (timeout) — RPC/other ports were already allowed, 3903 wasn't.
  • vps02: nftables, INPUT policy ACCEPT — no rule needed.
  • bigbox: localhost/WG direct, no firewall issue observed.

Monitoring stack targets (prometheus.yml)

  • garage job: 10.8.0.1:3903, 10.8.0.2:3903, 10.8.0.4:3903, path /metrics, bearer auth with the shared token.
  • garage_health job: same targets, path /health.
  • node job: 10.8.0.2:9100 (bigbox), 10.8.0.1:9100, 10.8.0.4:9100 (node-exporter, not yet deployed as of session end).
  • Grafana on :3001 (3000 occupied by gitea). Loki :3100, promtail via docker socket.
  • Alerts (alerts.yml): GarageNodeDown (up==0 for 2m), GarageNoQuorum (<2 of 3 up for 2m), resync error counter > 0, RPC node health == 0.

Environment notes

  • bigbox cannot resolve/reach general web (DNS fails for dl.garagehq.de, github.com raw is reachable; Firecrawl web tools not configured → web_search returns an error about FIRECRAWL_API_KEY). Authoritative docs: git.deuxfleurs.fr (Deuxfleurs/garage, branch main-v1) — doc/book/reference-manual/admin-api.md and doc/book/reference-manual/configuration.md.
  • docker exec into garage container impossible (scratch image). Use --entrypoint CLI.
  • cli/ dir in /opt/garage is a 10-byte "Not Found" placeholder — ignore.