Initial commit: Hermes skill garage-s3-cluster

This commit is contained in:
estorozhenko
2026-09-06 13:50:54 +00:00
commit 5b8c31d329
5 changed files with 253 additions and 0 deletions
+60
View File
@@ -0,0 +1,60 @@
---
name: garage-s3-cluster
description: Garage S3 cluster status checks on vps01/bigbox/vps02.
---
# Garage S3 cluster (vps01 / bigbox / vps02)
Garage = self-hosted S3-compatible storage cluster, replication_factor=3, read/write_quorum=2.
Runs in Docker on bigbox from /opt/garage. Bucket "obsidian" holds obsidian-vault backups.
## Topology (WireGuard 10.8.0.0/24)
- vps01: 10.8.0.1:3901 (hostname 5599453-kpa39l, zone vps01)
- bigbox: 10.8.0.2:3901 (zone home) — /opt/garage, wg0 = 10.8.0.2/24
- vps02: 10.8.0.4:3901 (zone vps02)
- S3 API (incl. admin API): 0.0.0.0:3900 on every node; RPC: WG IP:3901
- Config: /opt/garage/garage.toml (rpc_secret, admin_token, bootstrap_peers inside)
## Status check — the reliable way
Container `dxflrs/garage:v2.1.0` is SCRATCH: no shell, no CLI in PATH.
`docker exec garage garage ...` and `docker exec garage sh ...` both fail with
"executable file not found". The CLI binary lives at `/garage` INSIDE the image — run it
with `--entrypoint`, mounting BOTH config and meta (the node key lives in meta/; skipping
the meta mount gives "Unable to read node key. It will be generated..."):
```
docker run --rm \
-v /opt/garage/garage.toml:/etc/garage.toml:ro \
-v /opt/garage/meta:/var/lib/garage/meta \
--entrypoint /garage dxflrs/garage:v2.1.0 status
```
Subcommands that work at top level: `status` (health table — HEALTHY NODES),
`stats` (block manager + cluster-wide usage), `bucket list`.
Pitfall: `garage cluster status` does NOT exist in v2.1 — "Found argument 'cluster' which
wasn't expected". There is no `cluster` subcommand; status/stats are top-level.
Healthy signals in `garage stats`: "resync queue length: 0", "blocks with resync errors: 0",
MklTodo/GcTodo/InsQueue all 0.
## Admin HTTP API (port 3900) — do not rely on it
- `Authorization: Bearer <admin_token>` → "Unsupported authorization method" (S3-style XML error)
- `X-Garage-Admin-Token: <admin_token>` → AccessDenied "anonymous access"
- `/v2/status`, `/v1/status`, `/cluster/status` all same behavior.
- The CLI route above is the dependable path; no working HTTP admin call was found.
## Files in /opt/garage (bigbox)
- docker-compose.yml — service garage, network_mode: host, volumes: garage.toml, meta/, data/
- garage.toml — full config; admin_token duplicated here and in admin_token file
- admin_token — admin token (65 hex chars); secret — RPC secret (not for admin API)
- cli/ — BROKEN: garage-v2.1.0-linux-x86_64.tar.gz is a 10-byte "Not Found" placeholder. Ignore; use the in-image CLI above.
- meta/, data/ — LMDB storage (db_engine = "lmdb")
## Reachability probe (over WG)
```
for ip in 10.8.0.1 10.8.0.2 10.8.0.4; do
timeout 3 bash -c "echo > /dev/tcp/$ip/3901" 2>/dev/null && echo "$ip:3901 OK" || echo "$ip:3901 FAIL"
done
```
Script: `scripts/garage-status.sh` — runs the CLI status command with correct mounts.
+98
View File
@@ -0,0 +1,98 @@
# Grafana NO DATA debugging + monitoring hygiene (2026-09-02 session)
Symptom: dashboard panels show only "NO DATA" while Prometheus itself has data.
Three distinct root causes hit in one session — check in this order.
## 1. Datasource `uid` not pinned → Grafana generates a random one
If `uid:` is missing from the datasource in
`grafana/provisioning/datasources/datasources.yml`, Grafana assigns a RANDOM uid
at provisioning time (visible in its sqlite: `SELECT uid FROM data_source` →
e.g. `P8E80F9AEF21F6940`). Dashboards reference datasources by `"uid"` in each
panel (`"uid": "Prometheus"`), so panels silently find nothing → NO DATA.
Fix: pin the uid to match the dashboards:
```yaml
- name: Prometheus
type: prometheus
uid: Prometheus # must match panel "uid" refs
url: http://172.28.0.1:9090
isDefault: true
```
## 2. Prometheus on host network → `http://prometheus:9090` unresolvable from Grafana
Prometheus + blackbox run `network_mode: host` (needed to reach WireGuard
10.8.0.0/24), while Grafana is on the docker bridge network. The service name
`prometheus` does NOT resolve inside the bridge net
(`docker exec grafana getent hosts prometheus` → nothing). Loki, being on the
same bridge net, resolves fine — so one datasource works and the other doesn't.
Fix: point the datasource at the host via the bridge gateway:
```yaml
url: http://172.28.0.1:9090 # gateway IP = host as seen from Grafana's bridge net
```
Find the gateway: `docker inspect grafana --format '{{range .NetworkSettings.Networks}}{{.Gateway}} {{end}}'`
Verify from inside: `docker exec grafana wget -qO- http://172.28.0.1:9090/-/healthy`
Caveat: the bridge subnet is stable per docker daemon, but recreating the network
can change it.
## 3. Panels/alerts reference NON-EXISTENT metric names
Garage admin-API metrics (:3903) have NO `garage_` prefix. Only
`garage_build_info`, `garage_local_disk_avail/total`, `garage_replication_factor`
carry the prefix. Everything else is bare: `api_s3_request_counter`,
`block_resync_queue_length`, `block_resync_errored_blocks`, `cluster_*`
(connected_nodes, healthy, partitions_all_ok, layout_node_connected), `table_*`, `rpc_*`.
Phantom names that silently render NO DATA (never fire / never plot):
- `garage_block_count` → real: `block_resync_queue_length` + `block_resync_errored_blocks`
- `garage_rpc_node_health_is_up` → real: `cluster_layout_node_connected`
- `garage_api_s3_request_counter` → real: `api_s3_request_counter` (use `rate(api_s3_request_counter[5m])`)
- `garage_block_resync_error_count` → real: `block_resync_errored_blocks`
Check what actually exists before writing queries:
`curl 'http://127.0.0.1:9090/api/v1/label/__name__/values'` filtered by `job=garage`,
or `curl 'http://127.0.0.1:9090/api/v1/query?query={job="garage"}'` and list `__name__`.
## Provisioning reload semantics (Grafana)
- Dashboards: re-read automatically (~30s, updateIntervalSeconds).
- Datasources: provisioned ONLY at Grafana startup → after editing
datasources.yml you MUST `docker compose restart grafana`.
## Readable legends: relabel instance → hostname (not the raw scrape address)
By default `instance` = scrape address: tproxy shows `127.0.0.1:18081` (the LOCAL
end of the SSH tunnel — NOT "monitoring the local interface", just the transport),
garage shows `10.8.0.x:3903`. Fix at the SOURCE in prometheus.yml relabel_configs,
not per-panel legend, so panels/explore/alerts are consistent:
```yaml
relabel_configs:
- target_label: instance
replacement: vps03:8081 # tproxy
- source_labels: [__address__] # garage / node IP→hostname mapping
regex: 10.8.0.2:3903
target_label: instance
replacement: bigbox:3903
```
Result legends: `vps01:3903 / bigbox:3903 / vps02:3903` (garage),
`vps03:8081` (tproxy), `vps01:9100` etc (node). After changing relabel, restart
Prometheus and re-check `/api/v1/targets` — new labels apply immediately.
## Alert rules: same phantom-name trap
alerts.yml had `garage_block_resync_error_count` and `garage_rpc_node_health_is_up`
(never fire). Fixed to `block_resync_errored_blocks > 0` and
`cluster_layout_node_connected == 0`. Validate with
`docker exec prometheus promtool check config /etc/prometheus/prometheus.yml`
→ expect "6 rules found", all rules eval.
## Grafana sqlite inspection (no CLI auth needed)
Grafana's admin password may be unknown (changed in UI). Inspect state directly:
```
docker cp grafana:/var/lib/grafana/grafana.db /tmp/grafana.db
python3 -c "import sqlite3; print(sqlite3.connect('/tmp/grafana.db').execute('SELECT name,uid,url FROM data_source').fetchall())"
# dashboard JSON lives in table dashboard, column data
```
+62
View File
@@ -0,0 +1,62 @@
# Garage metrics + monitoring setup (2026-08-30 session)
## Working recipe: enable Prometheus /metrics on Garage v2.1
Root cause of the "Unsupported authorization method" mystery:
Garage v2.x serves the admin API (metrics + admin endpoints) on a SEPARATE port
configured via `[admin] api_bind_addr`. Without it, admin requests collide with the
S3 parser on port 3900 and always return S3-style XML errors — even with a valid
`Authorization: Bearer <admin_token>` header. The earlier "no working HTTP admin call
was found" section of SKILL.md was WRONG; the HTTP admin API works fine once the
separate port exists.
Config change per node (in /opt/garage/garage.toml, `[admin]` section):
```toml
[admin]
api_bind_addr = "<WG-IP>:3903" # 10.8.0.1 (vps01), 10.8.0.2 (bigbox), 10.8.0.4 (vps02)
metrics_token = "<token>" # equals admin_token in this cluster
admin_token = "<token>"
```
- `api_bind_addr` in `[s3_api]` is a DIFFERENT key — anchor patches on the `[admin]`
header when replacing, or the s3 entry gets found first.
- Bind to the WG IP (not 0.0.0.0) so metrics don't leak to the internet.
Verification:
- `garage admin-token list` (via --entrypoint CLI, meta mounted) shows
`metrics_token (from daemon configuration) ... Scope: Metrics`.
- `curl -H "Authorization: Bearer <token>" http://<node>:3903/metrics` → Prometheus text
(~90-110 KB per node), self-documented `# HELP` lines.
- `curl http://<node>:3903/health` → 200 if quorum, 503 otherwise.
## Rollout procedure (worked)
1. Patch config on all 3 nodes (vps02 needs `sudo`: file owned by root; vps01 too).
2. Restart ONE node at a time with ~8s pauses, order vps02 → vps01 → bigbox,
`docker compose restart garage` (config re-read at container start).
3. After all restarts, `garage status` must show HEALTHY NODES (all 3 up, v2.1.0).
## Firewall gotchas per host
- vps01: ufw active, policy DROP. New port needed
`sudo ufw allow from 10.8.0.0/24 to any port 3903 proto tcp`.
Symptom without it: WG ping OK but TCP connect hangs (timeout) — RPC/other ports
were already allowed, 3903 wasn't.
- vps02: nftables, INPUT policy ACCEPT — no rule needed.
- bigbox: localhost/WG direct, no firewall issue observed.
## Monitoring stack targets (prometheus.yml)
- `garage` job: 10.8.0.1:3903, 10.8.0.2:3903, 10.8.0.4:3903, path /metrics,
bearer auth with the shared token.
- `garage_health` job: same targets, path /health.
- `node` job: 10.8.0.2:9100 (bigbox), 10.8.0.1:9100, 10.8.0.4:9100 (node-exporter,
not yet deployed as of session end).
- Grafana on :3001 (3000 occupied by gitea). Loki :3100, promtail via docker socket.
- Alerts (alerts.yml): GarageNodeDown (up==0 for 2m), GarageNoQuorum (<2 of 3 up
for 2m), resync error counter > 0, RPC node health == 0.
## Environment notes
- bigbox cannot resolve/reach general web (DNS fails for dl.garagehq.de, github.com
raw is reachable; Firecrawl web tools not configured → web_search returns an error
about FIRECRAWL_API_KEY). Authoritative docs: git.deuxfleurs.fr
(Deuxfleurs/garage, branch main-v1) — `doc/book/reference-manual/admin-api.md` and
`doc/book/reference-manual/configuration.md`.
- `docker exec` into garage container impossible (scratch image). Use --entrypoint CLI.
- cli/ dir in /opt/garage is a 10-byte "Not Found" placeholder — ignore.
+10
View File
@@ -0,0 +1,10 @@
#!/usr/bin/env bash
# Garage cluster health: node table + block manager + usage.
# Works from bigbox; the dxflrs/garage:v2.1.0 image is scratch (no shell), so the
# CLI is invoked via --entrypoint /garage with config + meta mounted.
set -euo pipefail
GE=/opt/garage
exec docker run --rm \
-v "$GE/garage.toml:/etc/garage.toml:ro" \
-v "$GE/meta:/var/lib/garage/meta" \
--entrypoint /garage dxflrs/garage:v2.1.0 status
+23
View File
@@ -0,0 +1,23 @@
# Scrape config for the Garage cluster (3 nodes over WireGuard).
# Copy into prometheus.yml scrape_configs. Replace <metrics_token> with the
# metrics_token from [admin] in garage.toml (all nodes share one in this cluster).
- job_name: 'garage'
metrics_path: /metrics
scheme: http
authorization:
type: Bearer
credentials: '<metrics_token>'
static_configs:
- targets: ['10.8.0.1:3903', '10.8.0.2:3903', '10.8.0.4:3903']
labels:
cluster: garage
# Health probe (200 = quorum ok, 503 = lost quorum)
- job_name: 'garage_health'
metrics_path: /health
scheme: http
static_configs:
- targets: ['10.8.0.1:3903', '10.8.0.2:3903', '10.8.0.4:3903']
labels:
cluster: garage