Files
garage-s3-cluster/references/metrics-and-monitoring.md
T
2026-09-06 13:50:54 +00:00

62 lines
3.3 KiB
Markdown

# Garage metrics + monitoring setup (2026-08-30 session)
## Working recipe: enable Prometheus /metrics on Garage v2.1
Root cause of the "Unsupported authorization method" mystery:
Garage v2.x serves the admin API (metrics + admin endpoints) on a SEPARATE port
configured via `[admin] api_bind_addr`. Without it, admin requests collide with the
S3 parser on port 3900 and always return S3-style XML errors — even with a valid
`Authorization: Bearer <admin_token>` header. The earlier "no working HTTP admin call
was found" section of SKILL.md was WRONG; the HTTP admin API works fine once the
separate port exists.
Config change per node (in /opt/garage/garage.toml, `[admin]` section):
```toml
[admin]
api_bind_addr = "<WG-IP>:3903" # 10.8.0.1 (vps01), 10.8.0.2 (bigbox), 10.8.0.4 (vps02)
metrics_token = "<token>" # equals admin_token in this cluster
admin_token = "<token>"
```
- `api_bind_addr` in `[s3_api]` is a DIFFERENT key — anchor patches on the `[admin]`
header when replacing, or the s3 entry gets found first.
- Bind to the WG IP (not 0.0.0.0) so metrics don't leak to the internet.
Verification:
- `garage admin-token list` (via --entrypoint CLI, meta mounted) shows
`metrics_token (from daemon configuration) ... Scope: Metrics`.
- `curl -H "Authorization: Bearer <token>" http://<node>:3903/metrics` → Prometheus text
(~90-110 KB per node), self-documented `# HELP` lines.
- `curl http://<node>:3903/health` → 200 if quorum, 503 otherwise.
## Rollout procedure (worked)
1. Patch config on all 3 nodes (vps02 needs `sudo`: file owned by root; vps01 too).
2. Restart ONE node at a time with ~8s pauses, order vps02 → vps01 → bigbox,
`docker compose restart garage` (config re-read at container start).
3. After all restarts, `garage status` must show HEALTHY NODES (all 3 up, v2.1.0).
## Firewall gotchas per host
- vps01: ufw active, policy DROP. New port needed
`sudo ufw allow from 10.8.0.0/24 to any port 3903 proto tcp`.
Symptom without it: WG ping OK but TCP connect hangs (timeout) — RPC/other ports
were already allowed, 3903 wasn't.
- vps02: nftables, INPUT policy ACCEPT — no rule needed.
- bigbox: localhost/WG direct, no firewall issue observed.
## Monitoring stack targets (prometheus.yml)
- `garage` job: 10.8.0.1:3903, 10.8.0.2:3903, 10.8.0.4:3903, path /metrics,
bearer auth with the shared token.
- `garage_health` job: same targets, path /health.
- `node` job: 10.8.0.2:9100 (bigbox), 10.8.0.1:9100, 10.8.0.4:9100 (node-exporter,
not yet deployed as of session end).
- Grafana on :3001 (3000 occupied by gitea). Loki :3100, promtail via docker socket.
- Alerts (alerts.yml): GarageNodeDown (up==0 for 2m), GarageNoQuorum (<2 of 3 up
for 2m), resync error counter > 0, RPC node health == 0.
## Environment notes
- bigbox cannot resolve/reach general web (DNS fails for dl.garagehq.de, github.com
raw is reachable; Firecrawl web tools not configured → web_search returns an error
about FIRECRAWL_API_KEY). Authoritative docs: git.deuxfleurs.fr
(Deuxfleurs/garage, branch main-v1) — `doc/book/reference-manual/admin-api.md` and
`doc/book/reference-manual/configuration.md`.
- `docker exec` into garage container impossible (scratch image). Use --entrypoint CLI.
- cli/ dir in /opt/garage is a 10-byte "Not Found" placeholder — ignore.