mirror of
https://gitverse.ru/kpa39l/garage-s3-cluster.git
synced 2026-09-29 09:15:07 +00:00
Initial commit: Hermes skill garage-s3-cluster
This commit is contained in:
@@ -0,0 +1,60 @@
|
|||||||
|
---
|
||||||
|
name: garage-s3-cluster
|
||||||
|
description: Garage S3 cluster status checks on vps01/bigbox/vps02.
|
||||||
|
---
|
||||||
|
|
||||||
|
# Garage S3 cluster (vps01 / bigbox / vps02)
|
||||||
|
|
||||||
|
Garage = self-hosted S3-compatible storage cluster, replication_factor=3, read/write_quorum=2.
|
||||||
|
Runs in Docker on bigbox from /opt/garage. Bucket "obsidian" holds obsidian-vault backups.
|
||||||
|
|
||||||
|
## Topology (WireGuard 10.8.0.0/24)
|
||||||
|
- vps01: 10.8.0.1:3901 (hostname 5599453-kpa39l, zone vps01)
|
||||||
|
- bigbox: 10.8.0.2:3901 (zone home) — /opt/garage, wg0 = 10.8.0.2/24
|
||||||
|
- vps02: 10.8.0.4:3901 (zone vps02)
|
||||||
|
- S3 API (incl. admin API): 0.0.0.0:3900 on every node; RPC: WG IP:3901
|
||||||
|
- Config: /opt/garage/garage.toml (rpc_secret, admin_token, bootstrap_peers inside)
|
||||||
|
|
||||||
|
## Status check — the reliable way
|
||||||
|
Container `dxflrs/garage:v2.1.0` is SCRATCH: no shell, no CLI in PATH.
|
||||||
|
`docker exec garage garage ...` and `docker exec garage sh ...` both fail with
|
||||||
|
"executable file not found". The CLI binary lives at `/garage` INSIDE the image — run it
|
||||||
|
with `--entrypoint`, mounting BOTH config and meta (the node key lives in meta/; skipping
|
||||||
|
the meta mount gives "Unable to read node key. It will be generated..."):
|
||||||
|
|
||||||
|
```
|
||||||
|
docker run --rm \
|
||||||
|
-v /opt/garage/garage.toml:/etc/garage.toml:ro \
|
||||||
|
-v /opt/garage/meta:/var/lib/garage/meta \
|
||||||
|
--entrypoint /garage dxflrs/garage:v2.1.0 status
|
||||||
|
```
|
||||||
|
|
||||||
|
Subcommands that work at top level: `status` (health table — HEALTHY NODES),
|
||||||
|
`stats` (block manager + cluster-wide usage), `bucket list`.
|
||||||
|
Pitfall: `garage cluster status` does NOT exist in v2.1 — "Found argument 'cluster' which
|
||||||
|
wasn't expected". There is no `cluster` subcommand; status/stats are top-level.
|
||||||
|
|
||||||
|
Healthy signals in `garage stats`: "resync queue length: 0", "blocks with resync errors: 0",
|
||||||
|
MklTodo/GcTodo/InsQueue all 0.
|
||||||
|
|
||||||
|
## Admin HTTP API (port 3900) — do not rely on it
|
||||||
|
- `Authorization: Bearer <admin_token>` → "Unsupported authorization method" (S3-style XML error)
|
||||||
|
- `X-Garage-Admin-Token: <admin_token>` → AccessDenied "anonymous access"
|
||||||
|
- `/v2/status`, `/v1/status`, `/cluster/status` all same behavior.
|
||||||
|
- The CLI route above is the dependable path; no working HTTP admin call was found.
|
||||||
|
|
||||||
|
## Files in /opt/garage (bigbox)
|
||||||
|
- docker-compose.yml — service garage, network_mode: host, volumes: garage.toml, meta/, data/
|
||||||
|
- garage.toml — full config; admin_token duplicated here and in admin_token file
|
||||||
|
- admin_token — admin token (65 hex chars); secret — RPC secret (not for admin API)
|
||||||
|
- cli/ — BROKEN: garage-v2.1.0-linux-x86_64.tar.gz is a 10-byte "Not Found" placeholder. Ignore; use the in-image CLI above.
|
||||||
|
- meta/, data/ — LMDB storage (db_engine = "lmdb")
|
||||||
|
|
||||||
|
## Reachability probe (over WG)
|
||||||
|
```
|
||||||
|
for ip in 10.8.0.1 10.8.0.2 10.8.0.4; do
|
||||||
|
timeout 3 bash -c "echo > /dev/tcp/$ip/3901" 2>/dev/null && echo "$ip:3901 OK" || echo "$ip:3901 FAIL"
|
||||||
|
done
|
||||||
|
```
|
||||||
|
|
||||||
|
Script: `scripts/garage-status.sh` — runs the CLI status command with correct mounts.
|
||||||
@@ -0,0 +1,98 @@
|
|||||||
|
# Grafana NO DATA debugging + monitoring hygiene (2026-09-02 session)
|
||||||
|
|
||||||
|
Symptom: dashboard panels show only "NO DATA" while Prometheus itself has data.
|
||||||
|
Three distinct root causes hit in one session — check in this order.
|
||||||
|
|
||||||
|
## 1. Datasource `uid` not pinned → Grafana generates a random one
|
||||||
|
|
||||||
|
If `uid:` is missing from the datasource in
|
||||||
|
`grafana/provisioning/datasources/datasources.yml`, Grafana assigns a RANDOM uid
|
||||||
|
at provisioning time (visible in its sqlite: `SELECT uid FROM data_source` →
|
||||||
|
e.g. `P8E80F9AEF21F6940`). Dashboards reference datasources by `"uid"` in each
|
||||||
|
panel (`"uid": "Prometheus"`), so panels silently find nothing → NO DATA.
|
||||||
|
|
||||||
|
Fix: pin the uid to match the dashboards:
|
||||||
|
```yaml
|
||||||
|
- name: Prometheus
|
||||||
|
type: prometheus
|
||||||
|
uid: Prometheus # must match panel "uid" refs
|
||||||
|
url: http://172.28.0.1:9090
|
||||||
|
isDefault: true
|
||||||
|
```
|
||||||
|
|
||||||
|
## 2. Prometheus on host network → `http://prometheus:9090` unresolvable from Grafana
|
||||||
|
|
||||||
|
Prometheus + blackbox run `network_mode: host` (needed to reach WireGuard
|
||||||
|
10.8.0.0/24), while Grafana is on the docker bridge network. The service name
|
||||||
|
`prometheus` does NOT resolve inside the bridge net
|
||||||
|
(`docker exec grafana getent hosts prometheus` → nothing). Loki, being on the
|
||||||
|
same bridge net, resolves fine — so one datasource works and the other doesn't.
|
||||||
|
|
||||||
|
Fix: point the datasource at the host via the bridge gateway:
|
||||||
|
```yaml
|
||||||
|
url: http://172.28.0.1:9090 # gateway IP = host as seen from Grafana's bridge net
|
||||||
|
```
|
||||||
|
Find the gateway: `docker inspect grafana --format '{{range .NetworkSettings.Networks}}{{.Gateway}} {{end}}'`
|
||||||
|
Verify from inside: `docker exec grafana wget -qO- http://172.28.0.1:9090/-/healthy`
|
||||||
|
Caveat: the bridge subnet is stable per docker daemon, but recreating the network
|
||||||
|
can change it.
|
||||||
|
|
||||||
|
## 3. Panels/alerts reference NON-EXISTENT metric names
|
||||||
|
|
||||||
|
Garage admin-API metrics (:3903) have NO `garage_` prefix. Only
|
||||||
|
`garage_build_info`, `garage_local_disk_avail/total`, `garage_replication_factor`
|
||||||
|
carry the prefix. Everything else is bare: `api_s3_request_counter`,
|
||||||
|
`block_resync_queue_length`, `block_resync_errored_blocks`, `cluster_*`
|
||||||
|
(connected_nodes, healthy, partitions_all_ok, layout_node_connected), `table_*`, `rpc_*`.
|
||||||
|
|
||||||
|
Phantom names that silently render NO DATA (never fire / never plot):
|
||||||
|
- `garage_block_count` → real: `block_resync_queue_length` + `block_resync_errored_blocks`
|
||||||
|
- `garage_rpc_node_health_is_up` → real: `cluster_layout_node_connected`
|
||||||
|
- `garage_api_s3_request_counter` → real: `api_s3_request_counter` (use `rate(api_s3_request_counter[5m])`)
|
||||||
|
- `garage_block_resync_error_count` → real: `block_resync_errored_blocks`
|
||||||
|
|
||||||
|
Check what actually exists before writing queries:
|
||||||
|
`curl 'http://127.0.0.1:9090/api/v1/label/__name__/values'` filtered by `job=garage`,
|
||||||
|
or `curl 'http://127.0.0.1:9090/api/v1/query?query={job="garage"}'` and list `__name__`.
|
||||||
|
|
||||||
|
## Provisioning reload semantics (Grafana)
|
||||||
|
|
||||||
|
- Dashboards: re-read automatically (~30s, updateIntervalSeconds).
|
||||||
|
- Datasources: provisioned ONLY at Grafana startup → after editing
|
||||||
|
datasources.yml you MUST `docker compose restart grafana`.
|
||||||
|
|
||||||
|
## Readable legends: relabel instance → hostname (not the raw scrape address)
|
||||||
|
|
||||||
|
By default `instance` = scrape address: tproxy shows `127.0.0.1:18081` (the LOCAL
|
||||||
|
end of the SSH tunnel — NOT "monitoring the local interface", just the transport),
|
||||||
|
garage shows `10.8.0.x:3903`. Fix at the SOURCE in prometheus.yml relabel_configs,
|
||||||
|
not per-panel legend, so panels/explore/alerts are consistent:
|
||||||
|
```yaml
|
||||||
|
relabel_configs:
|
||||||
|
- target_label: instance
|
||||||
|
replacement: vps03:8081 # tproxy
|
||||||
|
- source_labels: [__address__] # garage / node IP→hostname mapping
|
||||||
|
regex: 10.8.0.2:3903
|
||||||
|
target_label: instance
|
||||||
|
replacement: bigbox:3903
|
||||||
|
```
|
||||||
|
Result legends: `vps01:3903 / bigbox:3903 / vps02:3903` (garage),
|
||||||
|
`vps03:8081` (tproxy), `vps01:9100` etc (node). After changing relabel, restart
|
||||||
|
Prometheus and re-check `/api/v1/targets` — new labels apply immediately.
|
||||||
|
|
||||||
|
## Alert rules: same phantom-name trap
|
||||||
|
|
||||||
|
alerts.yml had `garage_block_resync_error_count` and `garage_rpc_node_health_is_up`
|
||||||
|
(never fire). Fixed to `block_resync_errored_blocks > 0` and
|
||||||
|
`cluster_layout_node_connected == 0`. Validate with
|
||||||
|
`docker exec prometheus promtool check config /etc/prometheus/prometheus.yml`
|
||||||
|
→ expect "6 rules found", all rules eval.
|
||||||
|
|
||||||
|
## Grafana sqlite inspection (no CLI auth needed)
|
||||||
|
|
||||||
|
Grafana's admin password may be unknown (changed in UI). Inspect state directly:
|
||||||
|
```
|
||||||
|
docker cp grafana:/var/lib/grafana/grafana.db /tmp/grafana.db
|
||||||
|
python3 -c "import sqlite3; print(sqlite3.connect('/tmp/grafana.db').execute('SELECT name,uid,url FROM data_source').fetchall())"
|
||||||
|
# dashboard JSON lives in table dashboard, column data
|
||||||
|
```
|
||||||
@@ -0,0 +1,62 @@
|
|||||||
|
# Garage metrics + monitoring setup (2026-08-30 session)
|
||||||
|
|
||||||
|
## Working recipe: enable Prometheus /metrics on Garage v2.1
|
||||||
|
|
||||||
|
Root cause of the "Unsupported authorization method" mystery:
|
||||||
|
Garage v2.x serves the admin API (metrics + admin endpoints) on a SEPARATE port
|
||||||
|
configured via `[admin] api_bind_addr`. Without it, admin requests collide with the
|
||||||
|
S3 parser on port 3900 and always return S3-style XML errors — even with a valid
|
||||||
|
`Authorization: Bearer <admin_token>` header. The earlier "no working HTTP admin call
|
||||||
|
was found" section of SKILL.md was WRONG; the HTTP admin API works fine once the
|
||||||
|
separate port exists.
|
||||||
|
|
||||||
|
Config change per node (in /opt/garage/garage.toml, `[admin]` section):
|
||||||
|
```toml
|
||||||
|
[admin]
|
||||||
|
api_bind_addr = "<WG-IP>:3903" # 10.8.0.1 (vps01), 10.8.0.2 (bigbox), 10.8.0.4 (vps02)
|
||||||
|
metrics_token = "<token>" # equals admin_token in this cluster
|
||||||
|
admin_token = "<token>"
|
||||||
|
```
|
||||||
|
- `api_bind_addr` in `[s3_api]` is a DIFFERENT key — anchor patches on the `[admin]`
|
||||||
|
header when replacing, or the s3 entry gets found first.
|
||||||
|
- Bind to the WG IP (not 0.0.0.0) so metrics don't leak to the internet.
|
||||||
|
|
||||||
|
Verification:
|
||||||
|
- `garage admin-token list` (via --entrypoint CLI, meta mounted) shows
|
||||||
|
`metrics_token (from daemon configuration) ... Scope: Metrics`.
|
||||||
|
- `curl -H "Authorization: Bearer <token>" http://<node>:3903/metrics` → Prometheus text
|
||||||
|
(~90-110 KB per node), self-documented `# HELP` lines.
|
||||||
|
- `curl http://<node>:3903/health` → 200 if quorum, 503 otherwise.
|
||||||
|
|
||||||
|
## Rollout procedure (worked)
|
||||||
|
1. Patch config on all 3 nodes (vps02 needs `sudo`: file owned by root; vps01 too).
|
||||||
|
2. Restart ONE node at a time with ~8s pauses, order vps02 → vps01 → bigbox,
|
||||||
|
`docker compose restart garage` (config re-read at container start).
|
||||||
|
3. After all restarts, `garage status` must show HEALTHY NODES (all 3 up, v2.1.0).
|
||||||
|
|
||||||
|
## Firewall gotchas per host
|
||||||
|
- vps01: ufw active, policy DROP. New port needed
|
||||||
|
`sudo ufw allow from 10.8.0.0/24 to any port 3903 proto tcp`.
|
||||||
|
Symptom without it: WG ping OK but TCP connect hangs (timeout) — RPC/other ports
|
||||||
|
were already allowed, 3903 wasn't.
|
||||||
|
- vps02: nftables, INPUT policy ACCEPT — no rule needed.
|
||||||
|
- bigbox: localhost/WG direct, no firewall issue observed.
|
||||||
|
|
||||||
|
## Monitoring stack targets (prometheus.yml)
|
||||||
|
- `garage` job: 10.8.0.1:3903, 10.8.0.2:3903, 10.8.0.4:3903, path /metrics,
|
||||||
|
bearer auth with the shared token.
|
||||||
|
- `garage_health` job: same targets, path /health.
|
||||||
|
- `node` job: 10.8.0.2:9100 (bigbox), 10.8.0.1:9100, 10.8.0.4:9100 (node-exporter,
|
||||||
|
not yet deployed as of session end).
|
||||||
|
- Grafana on :3001 (3000 occupied by gitea). Loki :3100, promtail via docker socket.
|
||||||
|
- Alerts (alerts.yml): GarageNodeDown (up==0 for 2m), GarageNoQuorum (<2 of 3 up
|
||||||
|
for 2m), resync error counter > 0, RPC node health == 0.
|
||||||
|
|
||||||
|
## Environment notes
|
||||||
|
- bigbox cannot resolve/reach general web (DNS fails for dl.garagehq.de, github.com
|
||||||
|
raw is reachable; Firecrawl web tools not configured → web_search returns an error
|
||||||
|
about FIRECRAWL_API_KEY). Authoritative docs: git.deuxfleurs.fr
|
||||||
|
(Deuxfleurs/garage, branch main-v1) — `doc/book/reference-manual/admin-api.md` and
|
||||||
|
`doc/book/reference-manual/configuration.md`.
|
||||||
|
- `docker exec` into garage container impossible (scratch image). Use --entrypoint CLI.
|
||||||
|
- cli/ dir in /opt/garage is a 10-byte "Not Found" placeholder — ignore.
|
||||||
@@ -0,0 +1,10 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# Garage cluster health: node table + block manager + usage.
|
||||||
|
# Works from bigbox; the dxflrs/garage:v2.1.0 image is scratch (no shell), so the
|
||||||
|
# CLI is invoked via --entrypoint /garage with config + meta mounted.
|
||||||
|
set -euo pipefail
|
||||||
|
GE=/opt/garage
|
||||||
|
exec docker run --rm \
|
||||||
|
-v "$GE/garage.toml:/etc/garage.toml:ro" \
|
||||||
|
-v "$GE/meta:/var/lib/garage/meta" \
|
||||||
|
--entrypoint /garage dxflrs/garage:v2.1.0 status
|
||||||
@@ -0,0 +1,23 @@
|
|||||||
|
# Scrape config for the Garage cluster (3 nodes over WireGuard).
|
||||||
|
# Copy into prometheus.yml scrape_configs. Replace <metrics_token> with the
|
||||||
|
# metrics_token from [admin] in garage.toml (all nodes share one in this cluster).
|
||||||
|
|
||||||
|
- job_name: 'garage'
|
||||||
|
metrics_path: /metrics
|
||||||
|
scheme: http
|
||||||
|
authorization:
|
||||||
|
type: Bearer
|
||||||
|
credentials: '<metrics_token>'
|
||||||
|
static_configs:
|
||||||
|
- targets: ['10.8.0.1:3903', '10.8.0.2:3903', '10.8.0.4:3903']
|
||||||
|
labels:
|
||||||
|
cluster: garage
|
||||||
|
|
||||||
|
# Health probe (200 = quorum ok, 503 = lost quorum)
|
||||||
|
- job_name: 'garage_health'
|
||||||
|
metrics_path: /health
|
||||||
|
scheme: http
|
||||||
|
static_configs:
|
||||||
|
- targets: ['10.8.0.1:3903', '10.8.0.2:3903', '10.8.0.4:3903']
|
||||||
|
labels:
|
||||||
|
cluster: garage
|
||||||
Reference in New Issue
Block a user