commit 5b8c31d329fca20f74e98f5fe3b9b74b01d925ee Author: estorozhenko Date: Sun Sep 6 13:50:54 2026 +0000 Initial commit: Hermes skill garage-s3-cluster diff --git a/SKILL.md b/SKILL.md new file mode 100644 index 0000000..d69b957 --- /dev/null +++ b/SKILL.md @@ -0,0 +1,60 @@ +--- +name: garage-s3-cluster +description: Garage S3 cluster status checks on vps01/bigbox/vps02. +--- + +# Garage S3 cluster (vps01 / bigbox / vps02) + +Garage = self-hosted S3-compatible storage cluster, replication_factor=3, read/write_quorum=2. +Runs in Docker on bigbox from /opt/garage. Bucket "obsidian" holds obsidian-vault backups. + +## Topology (WireGuard 10.8.0.0/24) +- vps01: 10.8.0.1:3901 (hostname 5599453-kpa39l, zone vps01) +- bigbox: 10.8.0.2:3901 (zone home) — /opt/garage, wg0 = 10.8.0.2/24 +- vps02: 10.8.0.4:3901 (zone vps02) +- S3 API (incl. admin API): 0.0.0.0:3900 on every node; RPC: WG IP:3901 +- Config: /opt/garage/garage.toml (rpc_secret, admin_token, bootstrap_peers inside) + +## Status check — the reliable way +Container `dxflrs/garage:v2.1.0` is SCRATCH: no shell, no CLI in PATH. +`docker exec garage garage ...` and `docker exec garage sh ...` both fail with +"executable file not found". The CLI binary lives at `/garage` INSIDE the image — run it +with `--entrypoint`, mounting BOTH config and meta (the node key lives in meta/; skipping +the meta mount gives "Unable to read node key. It will be generated..."): + +``` +docker run --rm \ + -v /opt/garage/garage.toml:/etc/garage.toml:ro \ + -v /opt/garage/meta:/var/lib/garage/meta \ + --entrypoint /garage dxflrs/garage:v2.1.0 status +``` + +Subcommands that work at top level: `status` (health table — HEALTHY NODES), +`stats` (block manager + cluster-wide usage), `bucket list`. +Pitfall: `garage cluster status` does NOT exist in v2.1 — "Found argument 'cluster' which +wasn't expected". There is no `cluster` subcommand; status/stats are top-level. + +Healthy signals in `garage stats`: "resync queue length: 0", "blocks with resync errors: 0", +MklTodo/GcTodo/InsQueue all 0. + +## Admin HTTP API (port 3900) — do not rely on it +- `Authorization: Bearer ` → "Unsupported authorization method" (S3-style XML error) +- `X-Garage-Admin-Token: ` → AccessDenied "anonymous access" +- `/v2/status`, `/v1/status`, `/cluster/status` all same behavior. +- The CLI route above is the dependable path; no working HTTP admin call was found. + +## Files in /opt/garage (bigbox) +- docker-compose.yml — service garage, network_mode: host, volumes: garage.toml, meta/, data/ +- garage.toml — full config; admin_token duplicated here and in admin_token file +- admin_token — admin token (65 hex chars); secret — RPC secret (not for admin API) +- cli/ — BROKEN: garage-v2.1.0-linux-x86_64.tar.gz is a 10-byte "Not Found" placeholder. Ignore; use the in-image CLI above. +- meta/, data/ — LMDB storage (db_engine = "lmdb") + +## Reachability probe (over WG) +``` +for ip in 10.8.0.1 10.8.0.2 10.8.0.4; do + timeout 3 bash -c "echo > /dev/tcp/$ip/3901" 2>/dev/null && echo "$ip:3901 OK" || echo "$ip:3901 FAIL" +done +``` + +Script: `scripts/garage-status.sh` — runs the CLI status command with correct mounts. \ No newline at end of file diff --git a/references/grafana-nodata-debugging.md b/references/grafana-nodata-debugging.md new file mode 100644 index 0000000..7bad31d --- /dev/null +++ b/references/grafana-nodata-debugging.md @@ -0,0 +1,98 @@ +# Grafana NO DATA debugging + monitoring hygiene (2026-09-02 session) + +Symptom: dashboard panels show only "NO DATA" while Prometheus itself has data. +Three distinct root causes hit in one session — check in this order. + +## 1. Datasource `uid` not pinned → Grafana generates a random one + +If `uid:` is missing from the datasource in +`grafana/provisioning/datasources/datasources.yml`, Grafana assigns a RANDOM uid +at provisioning time (visible in its sqlite: `SELECT uid FROM data_source` → +e.g. `P8E80F9AEF21F6940`). Dashboards reference datasources by `"uid"` in each +panel (`"uid": "Prometheus"`), so panels silently find nothing → NO DATA. + +Fix: pin the uid to match the dashboards: +```yaml +- name: Prometheus + type: prometheus + uid: Prometheus # must match panel "uid" refs + url: http://172.28.0.1:9090 + isDefault: true +``` + +## 2. Prometheus on host network → `http://prometheus:9090` unresolvable from Grafana + +Prometheus + blackbox run `network_mode: host` (needed to reach WireGuard +10.8.0.0/24), while Grafana is on the docker bridge network. The service name +`prometheus` does NOT resolve inside the bridge net +(`docker exec grafana getent hosts prometheus` → nothing). Loki, being on the +same bridge net, resolves fine — so one datasource works and the other doesn't. + +Fix: point the datasource at the host via the bridge gateway: +```yaml +url: http://172.28.0.1:9090 # gateway IP = host as seen from Grafana's bridge net +``` +Find the gateway: `docker inspect grafana --format '{{range .NetworkSettings.Networks}}{{.Gateway}} {{end}}'` +Verify from inside: `docker exec grafana wget -qO- http://172.28.0.1:9090/-/healthy` +Caveat: the bridge subnet is stable per docker daemon, but recreating the network +can change it. + +## 3. Panels/alerts reference NON-EXISTENT metric names + +Garage admin-API metrics (:3903) have NO `garage_` prefix. Only +`garage_build_info`, `garage_local_disk_avail/total`, `garage_replication_factor` +carry the prefix. Everything else is bare: `api_s3_request_counter`, +`block_resync_queue_length`, `block_resync_errored_blocks`, `cluster_*` +(connected_nodes, healthy, partitions_all_ok, layout_node_connected), `table_*`, `rpc_*`. + +Phantom names that silently render NO DATA (never fire / never plot): +- `garage_block_count` → real: `block_resync_queue_length` + `block_resync_errored_blocks` +- `garage_rpc_node_health_is_up` → real: `cluster_layout_node_connected` +- `garage_api_s3_request_counter` → real: `api_s3_request_counter` (use `rate(api_s3_request_counter[5m])`) +- `garage_block_resync_error_count` → real: `block_resync_errored_blocks` + +Check what actually exists before writing queries: +`curl 'http://127.0.0.1:9090/api/v1/label/__name__/values'` filtered by `job=garage`, +or `curl 'http://127.0.0.1:9090/api/v1/query?query={job="garage"}'` and list `__name__`. + +## Provisioning reload semantics (Grafana) + +- Dashboards: re-read automatically (~30s, updateIntervalSeconds). +- Datasources: provisioned ONLY at Grafana startup → after editing + datasources.yml you MUST `docker compose restart grafana`. + +## Readable legends: relabel instance → hostname (not the raw scrape address) + +By default `instance` = scrape address: tproxy shows `127.0.0.1:18081` (the LOCAL +end of the SSH tunnel — NOT "monitoring the local interface", just the transport), +garage shows `10.8.0.x:3903`. Fix at the SOURCE in prometheus.yml relabel_configs, +not per-panel legend, so panels/explore/alerts are consistent: +```yaml +relabel_configs: + - target_label: instance + replacement: vps03:8081 # tproxy + - source_labels: [__address__] # garage / node IP→hostname mapping + regex: 10.8.0.2:3903 + target_label: instance + replacement: bigbox:3903 +``` +Result legends: `vps01:3903 / bigbox:3903 / vps02:3903` (garage), +`vps03:8081` (tproxy), `vps01:9100` etc (node). After changing relabel, restart +Prometheus and re-check `/api/v1/targets` — new labels apply immediately. + +## Alert rules: same phantom-name trap + +alerts.yml had `garage_block_resync_error_count` and `garage_rpc_node_health_is_up` +(never fire). Fixed to `block_resync_errored_blocks > 0` and +`cluster_layout_node_connected == 0`. Validate with +`docker exec prometheus promtool check config /etc/prometheus/prometheus.yml` +→ expect "6 rules found", all rules eval. + +## Grafana sqlite inspection (no CLI auth needed) + +Grafana's admin password may be unknown (changed in UI). Inspect state directly: +``` +docker cp grafana:/var/lib/grafana/grafana.db /tmp/grafana.db +python3 -c "import sqlite3; print(sqlite3.connect('/tmp/grafana.db').execute('SELECT name,uid,url FROM data_source').fetchall())" +# dashboard JSON lives in table dashboard, column data +``` diff --git a/references/metrics-and-monitoring.md b/references/metrics-and-monitoring.md new file mode 100644 index 0000000..991106a --- /dev/null +++ b/references/metrics-and-monitoring.md @@ -0,0 +1,62 @@ +# Garage metrics + monitoring setup (2026-08-30 session) + +## Working recipe: enable Prometheus /metrics on Garage v2.1 + +Root cause of the "Unsupported authorization method" mystery: +Garage v2.x serves the admin API (metrics + admin endpoints) on a SEPARATE port +configured via `[admin] api_bind_addr`. Without it, admin requests collide with the +S3 parser on port 3900 and always return S3-style XML errors — even with a valid +`Authorization: Bearer ` header. The earlier "no working HTTP admin call +was found" section of SKILL.md was WRONG; the HTTP admin API works fine once the +separate port exists. + +Config change per node (in /opt/garage/garage.toml, `[admin]` section): +```toml +[admin] +api_bind_addr = ":3903" # 10.8.0.1 (vps01), 10.8.0.2 (bigbox), 10.8.0.4 (vps02) +metrics_token = "" # equals admin_token in this cluster +admin_token = "" +``` +- `api_bind_addr` in `[s3_api]` is a DIFFERENT key — anchor patches on the `[admin]` + header when replacing, or the s3 entry gets found first. +- Bind to the WG IP (not 0.0.0.0) so metrics don't leak to the internet. + +Verification: +- `garage admin-token list` (via --entrypoint CLI, meta mounted) shows + `metrics_token (from daemon configuration) ... Scope: Metrics`. +- `curl -H "Authorization: Bearer " http://:3903/metrics` → Prometheus text + (~90-110 KB per node), self-documented `# HELP` lines. +- `curl http://:3903/health` → 200 if quorum, 503 otherwise. + +## Rollout procedure (worked) +1. Patch config on all 3 nodes (vps02 needs `sudo`: file owned by root; vps01 too). +2. Restart ONE node at a time with ~8s pauses, order vps02 → vps01 → bigbox, + `docker compose restart garage` (config re-read at container start). +3. After all restarts, `garage status` must show HEALTHY NODES (all 3 up, v2.1.0). + +## Firewall gotchas per host +- vps01: ufw active, policy DROP. New port needed + `sudo ufw allow from 10.8.0.0/24 to any port 3903 proto tcp`. + Symptom without it: WG ping OK but TCP connect hangs (timeout) — RPC/other ports + were already allowed, 3903 wasn't. +- vps02: nftables, INPUT policy ACCEPT — no rule needed. +- bigbox: localhost/WG direct, no firewall issue observed. + +## Monitoring stack targets (prometheus.yml) +- `garage` job: 10.8.0.1:3903, 10.8.0.2:3903, 10.8.0.4:3903, path /metrics, + bearer auth with the shared token. +- `garage_health` job: same targets, path /health. +- `node` job: 10.8.0.2:9100 (bigbox), 10.8.0.1:9100, 10.8.0.4:9100 (node-exporter, + not yet deployed as of session end). +- Grafana on :3001 (3000 occupied by gitea). Loki :3100, promtail via docker socket. +- Alerts (alerts.yml): GarageNodeDown (up==0 for 2m), GarageNoQuorum (<2 of 3 up + for 2m), resync error counter > 0, RPC node health == 0. + +## Environment notes +- bigbox cannot resolve/reach general web (DNS fails for dl.garagehq.de, github.com + raw is reachable; Firecrawl web tools not configured → web_search returns an error + about FIRECRAWL_API_KEY). Authoritative docs: git.deuxfleurs.fr + (Deuxfleurs/garage, branch main-v1) — `doc/book/reference-manual/admin-api.md` and + `doc/book/reference-manual/configuration.md`. +- `docker exec` into garage container impossible (scratch image). Use --entrypoint CLI. +- cli/ dir in /opt/garage is a 10-byte "Not Found" placeholder — ignore. \ No newline at end of file diff --git a/scripts/garage-status.sh b/scripts/garage-status.sh new file mode 100644 index 0000000..497e7a8 --- /dev/null +++ b/scripts/garage-status.sh @@ -0,0 +1,10 @@ +#!/usr/bin/env bash +# Garage cluster health: node table + block manager + usage. +# Works from bigbox; the dxflrs/garage:v2.1.0 image is scratch (no shell), so the +# CLI is invoked via --entrypoint /garage with config + meta mounted. +set -euo pipefail +GE=/opt/garage +exec docker run --rm \ + -v "$GE/garage.toml:/etc/garage.toml:ro" \ + -v "$GE/meta:/var/lib/garage/meta" \ + --entrypoint /garage dxflrs/garage:v2.1.0 status \ No newline at end of file diff --git a/templates/prometheus-garage-scrape.yml b/templates/prometheus-garage-scrape.yml new file mode 100644 index 0000000..2295ee4 --- /dev/null +++ b/templates/prometheus-garage-scrape.yml @@ -0,0 +1,23 @@ +# Scrape config for the Garage cluster (3 nodes over WireGuard). +# Copy into prometheus.yml scrape_configs. Replace with the +# metrics_token from [admin] in garage.toml (all nodes share one in this cluster). + + - job_name: 'garage' + metrics_path: /metrics + scheme: http + authorization: + type: Bearer + credentials: '' + static_configs: + - targets: ['10.8.0.1:3903', '10.8.0.2:3903', '10.8.0.4:3903'] + labels: + cluster: garage + + # Health probe (200 = quorum ok, 503 = lost quorum) + - job_name: 'garage_health' + metrics_path: /health + scheme: http + static_configs: + - targets: ['10.8.0.1:3903', '10.8.0.2:3903', '10.8.0.4:3903'] + labels: + cluster: garage \ No newline at end of file