mirror of
https://gitverse.ru/kpa39l/garage-s3-cluster.git
synced 2026-09-29 09:15:07 +00:00
Initial commit: Hermes skill garage-s3-cluster
This commit is contained in:
@@ -0,0 +1,98 @@
|
||||
# Grafana NO DATA debugging + monitoring hygiene (2026-09-02 session)
|
||||
|
||||
Symptom: dashboard panels show only "NO DATA" while Prometheus itself has data.
|
||||
Three distinct root causes hit in one session — check in this order.
|
||||
|
||||
## 1. Datasource `uid` not pinned → Grafana generates a random one
|
||||
|
||||
If `uid:` is missing from the datasource in
|
||||
`grafana/provisioning/datasources/datasources.yml`, Grafana assigns a RANDOM uid
|
||||
at provisioning time (visible in its sqlite: `SELECT uid FROM data_source` →
|
||||
e.g. `P8E80F9AEF21F6940`). Dashboards reference datasources by `"uid"` in each
|
||||
panel (`"uid": "Prometheus"`), so panels silently find nothing → NO DATA.
|
||||
|
||||
Fix: pin the uid to match the dashboards:
|
||||
```yaml
|
||||
- name: Prometheus
|
||||
type: prometheus
|
||||
uid: Prometheus # must match panel "uid" refs
|
||||
url: http://172.28.0.1:9090
|
||||
isDefault: true
|
||||
```
|
||||
|
||||
## 2. Prometheus on host network → `http://prometheus:9090` unresolvable from Grafana
|
||||
|
||||
Prometheus + blackbox run `network_mode: host` (needed to reach WireGuard
|
||||
10.8.0.0/24), while Grafana is on the docker bridge network. The service name
|
||||
`prometheus` does NOT resolve inside the bridge net
|
||||
(`docker exec grafana getent hosts prometheus` → nothing). Loki, being on the
|
||||
same bridge net, resolves fine — so one datasource works and the other doesn't.
|
||||
|
||||
Fix: point the datasource at the host via the bridge gateway:
|
||||
```yaml
|
||||
url: http://172.28.0.1:9090 # gateway IP = host as seen from Grafana's bridge net
|
||||
```
|
||||
Find the gateway: `docker inspect grafana --format '{{range .NetworkSettings.Networks}}{{.Gateway}} {{end}}'`
|
||||
Verify from inside: `docker exec grafana wget -qO- http://172.28.0.1:9090/-/healthy`
|
||||
Caveat: the bridge subnet is stable per docker daemon, but recreating the network
|
||||
can change it.
|
||||
|
||||
## 3. Panels/alerts reference NON-EXISTENT metric names
|
||||
|
||||
Garage admin-API metrics (:3903) have NO `garage_` prefix. Only
|
||||
`garage_build_info`, `garage_local_disk_avail/total`, `garage_replication_factor`
|
||||
carry the prefix. Everything else is bare: `api_s3_request_counter`,
|
||||
`block_resync_queue_length`, `block_resync_errored_blocks`, `cluster_*`
|
||||
(connected_nodes, healthy, partitions_all_ok, layout_node_connected), `table_*`, `rpc_*`.
|
||||
|
||||
Phantom names that silently render NO DATA (never fire / never plot):
|
||||
- `garage_block_count` → real: `block_resync_queue_length` + `block_resync_errored_blocks`
|
||||
- `garage_rpc_node_health_is_up` → real: `cluster_layout_node_connected`
|
||||
- `garage_api_s3_request_counter` → real: `api_s3_request_counter` (use `rate(api_s3_request_counter[5m])`)
|
||||
- `garage_block_resync_error_count` → real: `block_resync_errored_blocks`
|
||||
|
||||
Check what actually exists before writing queries:
|
||||
`curl 'http://127.0.0.1:9090/api/v1/label/__name__/values'` filtered by `job=garage`,
|
||||
or `curl 'http://127.0.0.1:9090/api/v1/query?query={job="garage"}'` and list `__name__`.
|
||||
|
||||
## Provisioning reload semantics (Grafana)
|
||||
|
||||
- Dashboards: re-read automatically (~30s, updateIntervalSeconds).
|
||||
- Datasources: provisioned ONLY at Grafana startup → after editing
|
||||
datasources.yml you MUST `docker compose restart grafana`.
|
||||
|
||||
## Readable legends: relabel instance → hostname (not the raw scrape address)
|
||||
|
||||
By default `instance` = scrape address: tproxy shows `127.0.0.1:18081` (the LOCAL
|
||||
end of the SSH tunnel — NOT "monitoring the local interface", just the transport),
|
||||
garage shows `10.8.0.x:3903`. Fix at the SOURCE in prometheus.yml relabel_configs,
|
||||
not per-panel legend, so panels/explore/alerts are consistent:
|
||||
```yaml
|
||||
relabel_configs:
|
||||
- target_label: instance
|
||||
replacement: vps03:8081 # tproxy
|
||||
- source_labels: [__address__] # garage / node IP→hostname mapping
|
||||
regex: 10.8.0.2:3903
|
||||
target_label: instance
|
||||
replacement: bigbox:3903
|
||||
```
|
||||
Result legends: `vps01:3903 / bigbox:3903 / vps02:3903` (garage),
|
||||
`vps03:8081` (tproxy), `vps01:9100` etc (node). After changing relabel, restart
|
||||
Prometheus and re-check `/api/v1/targets` — new labels apply immediately.
|
||||
|
||||
## Alert rules: same phantom-name trap
|
||||
|
||||
alerts.yml had `garage_block_resync_error_count` and `garage_rpc_node_health_is_up`
|
||||
(never fire). Fixed to `block_resync_errored_blocks > 0` and
|
||||
`cluster_layout_node_connected == 0`. Validate with
|
||||
`docker exec prometheus promtool check config /etc/prometheus/prometheus.yml`
|
||||
→ expect "6 rules found", all rules eval.
|
||||
|
||||
## Grafana sqlite inspection (no CLI auth needed)
|
||||
|
||||
Grafana's admin password may be unknown (changed in UI). Inspect state directly:
|
||||
```
|
||||
docker cp grafana:/var/lib/grafana/grafana.db /tmp/grafana.db
|
||||
python3 -c "import sqlite3; print(sqlite3.connect('/tmp/grafana.db').execute('SELECT name,uid,url FROM data_source').fetchall())"
|
||||
# dashboard JSON lives in table dashboard, column data
|
||||
```
|
||||
@@ -0,0 +1,62 @@
|
||||
# Garage metrics + monitoring setup (2026-08-30 session)
|
||||
|
||||
## Working recipe: enable Prometheus /metrics on Garage v2.1
|
||||
|
||||
Root cause of the "Unsupported authorization method" mystery:
|
||||
Garage v2.x serves the admin API (metrics + admin endpoints) on a SEPARATE port
|
||||
configured via `[admin] api_bind_addr`. Without it, admin requests collide with the
|
||||
S3 parser on port 3900 and always return S3-style XML errors — even with a valid
|
||||
`Authorization: Bearer <admin_token>` header. The earlier "no working HTTP admin call
|
||||
was found" section of SKILL.md was WRONG; the HTTP admin API works fine once the
|
||||
separate port exists.
|
||||
|
||||
Config change per node (in /opt/garage/garage.toml, `[admin]` section):
|
||||
```toml
|
||||
[admin]
|
||||
api_bind_addr = "<WG-IP>:3903" # 10.8.0.1 (vps01), 10.8.0.2 (bigbox), 10.8.0.4 (vps02)
|
||||
metrics_token = "<token>" # equals admin_token in this cluster
|
||||
admin_token = "<token>"
|
||||
```
|
||||
- `api_bind_addr` in `[s3_api]` is a DIFFERENT key — anchor patches on the `[admin]`
|
||||
header when replacing, or the s3 entry gets found first.
|
||||
- Bind to the WG IP (not 0.0.0.0) so metrics don't leak to the internet.
|
||||
|
||||
Verification:
|
||||
- `garage admin-token list` (via --entrypoint CLI, meta mounted) shows
|
||||
`metrics_token (from daemon configuration) ... Scope: Metrics`.
|
||||
- `curl -H "Authorization: Bearer <token>" http://<node>:3903/metrics` → Prometheus text
|
||||
(~90-110 KB per node), self-documented `# HELP` lines.
|
||||
- `curl http://<node>:3903/health` → 200 if quorum, 503 otherwise.
|
||||
|
||||
## Rollout procedure (worked)
|
||||
1. Patch config on all 3 nodes (vps02 needs `sudo`: file owned by root; vps01 too).
|
||||
2. Restart ONE node at a time with ~8s pauses, order vps02 → vps01 → bigbox,
|
||||
`docker compose restart garage` (config re-read at container start).
|
||||
3. After all restarts, `garage status` must show HEALTHY NODES (all 3 up, v2.1.0).
|
||||
|
||||
## Firewall gotchas per host
|
||||
- vps01: ufw active, policy DROP. New port needed
|
||||
`sudo ufw allow from 10.8.0.0/24 to any port 3903 proto tcp`.
|
||||
Symptom without it: WG ping OK but TCP connect hangs (timeout) — RPC/other ports
|
||||
were already allowed, 3903 wasn't.
|
||||
- vps02: nftables, INPUT policy ACCEPT — no rule needed.
|
||||
- bigbox: localhost/WG direct, no firewall issue observed.
|
||||
|
||||
## Monitoring stack targets (prometheus.yml)
|
||||
- `garage` job: 10.8.0.1:3903, 10.8.0.2:3903, 10.8.0.4:3903, path /metrics,
|
||||
bearer auth with the shared token.
|
||||
- `garage_health` job: same targets, path /health.
|
||||
- `node` job: 10.8.0.2:9100 (bigbox), 10.8.0.1:9100, 10.8.0.4:9100 (node-exporter,
|
||||
not yet deployed as of session end).
|
||||
- Grafana on :3001 (3000 occupied by gitea). Loki :3100, promtail via docker socket.
|
||||
- Alerts (alerts.yml): GarageNodeDown (up==0 for 2m), GarageNoQuorum (<2 of 3 up
|
||||
for 2m), resync error counter > 0, RPC node health == 0.
|
||||
|
||||
## Environment notes
|
||||
- bigbox cannot resolve/reach general web (DNS fails for dl.garagehq.de, github.com
|
||||
raw is reachable; Firecrawl web tools not configured → web_search returns an error
|
||||
about FIRECRAWL_API_KEY). Authoritative docs: git.deuxfleurs.fr
|
||||
(Deuxfleurs/garage, branch main-v1) — `doc/book/reference-manual/admin-api.md` and
|
||||
`doc/book/reference-manual/configuration.md`.
|
||||
- `docker exec` into garage container impossible (scratch image). Use --entrypoint CLI.
|
||||
- cli/ dir in /opt/garage is a 10-byte "Not Found" placeholder — ignore.
|
||||
Reference in New Issue
Block a user