mirror of
https://gitverse.ru/kpa39l/garage-s3-cluster.git
synced 2026-09-29 09:15:07 +00:00
Initial commit: Hermes skill garage-s3-cluster
This commit is contained in:
@@ -0,0 +1,62 @@
|
||||
# Garage metrics + monitoring setup (2026-08-30 session)
|
||||
|
||||
## Working recipe: enable Prometheus /metrics on Garage v2.1
|
||||
|
||||
Root cause of the "Unsupported authorization method" mystery:
|
||||
Garage v2.x serves the admin API (metrics + admin endpoints) on a SEPARATE port
|
||||
configured via `[admin] api_bind_addr`. Without it, admin requests collide with the
|
||||
S3 parser on port 3900 and always return S3-style XML errors — even with a valid
|
||||
`Authorization: Bearer <admin_token>` header. The earlier "no working HTTP admin call
|
||||
was found" section of SKILL.md was WRONG; the HTTP admin API works fine once the
|
||||
separate port exists.
|
||||
|
||||
Config change per node (in /opt/garage/garage.toml, `[admin]` section):
|
||||
```toml
|
||||
[admin]
|
||||
api_bind_addr = "<WG-IP>:3903" # 10.8.0.1 (vps01), 10.8.0.2 (bigbox), 10.8.0.4 (vps02)
|
||||
metrics_token = "<token>" # equals admin_token in this cluster
|
||||
admin_token = "<token>"
|
||||
```
|
||||
- `api_bind_addr` in `[s3_api]` is a DIFFERENT key — anchor patches on the `[admin]`
|
||||
header when replacing, or the s3 entry gets found first.
|
||||
- Bind to the WG IP (not 0.0.0.0) so metrics don't leak to the internet.
|
||||
|
||||
Verification:
|
||||
- `garage admin-token list` (via --entrypoint CLI, meta mounted) shows
|
||||
`metrics_token (from daemon configuration) ... Scope: Metrics`.
|
||||
- `curl -H "Authorization: Bearer <token>" http://<node>:3903/metrics` → Prometheus text
|
||||
(~90-110 KB per node), self-documented `# HELP` lines.
|
||||
- `curl http://<node>:3903/health` → 200 if quorum, 503 otherwise.
|
||||
|
||||
## Rollout procedure (worked)
|
||||
1. Patch config on all 3 nodes (vps02 needs `sudo`: file owned by root; vps01 too).
|
||||
2. Restart ONE node at a time with ~8s pauses, order vps02 → vps01 → bigbox,
|
||||
`docker compose restart garage` (config re-read at container start).
|
||||
3. After all restarts, `garage status` must show HEALTHY NODES (all 3 up, v2.1.0).
|
||||
|
||||
## Firewall gotchas per host
|
||||
- vps01: ufw active, policy DROP. New port needed
|
||||
`sudo ufw allow from 10.8.0.0/24 to any port 3903 proto tcp`.
|
||||
Symptom without it: WG ping OK but TCP connect hangs (timeout) — RPC/other ports
|
||||
were already allowed, 3903 wasn't.
|
||||
- vps02: nftables, INPUT policy ACCEPT — no rule needed.
|
||||
- bigbox: localhost/WG direct, no firewall issue observed.
|
||||
|
||||
## Monitoring stack targets (prometheus.yml)
|
||||
- `garage` job: 10.8.0.1:3903, 10.8.0.2:3903, 10.8.0.4:3903, path /metrics,
|
||||
bearer auth with the shared token.
|
||||
- `garage_health` job: same targets, path /health.
|
||||
- `node` job: 10.8.0.2:9100 (bigbox), 10.8.0.1:9100, 10.8.0.4:9100 (node-exporter,
|
||||
not yet deployed as of session end).
|
||||
- Grafana on :3001 (3000 occupied by gitea). Loki :3100, promtail via docker socket.
|
||||
- Alerts (alerts.yml): GarageNodeDown (up==0 for 2m), GarageNoQuorum (<2 of 3 up
|
||||
for 2m), resync error counter > 0, RPC node health == 0.
|
||||
|
||||
## Environment notes
|
||||
- bigbox cannot resolve/reach general web (DNS fails for dl.garagehq.de, github.com
|
||||
raw is reachable; Firecrawl web tools not configured → web_search returns an error
|
||||
about FIRECRAWL_API_KEY). Authoritative docs: git.deuxfleurs.fr
|
||||
(Deuxfleurs/garage, branch main-v1) — `doc/book/reference-manual/admin-api.md` and
|
||||
`doc/book/reference-manual/configuration.md`.
|
||||
- `docker exec` into garage container impossible (scratch image). Use --entrypoint CLI.
|
||||
- cli/ dir in /opt/garage is a 10-byte "Not Found" placeholder — ignore.
|
||||
Reference in New Issue
Block a user