mirror of
https://gitverse.ru/kpa39l/garage-s3-cluster.git
synced 2026-09-29 09:15:07 +00:00
99 lines
4.6 KiB
Markdown
99 lines
4.6 KiB
Markdown
# Grafana NO DATA debugging + monitoring hygiene (2026-09-02 session)
|
|
|
|
Symptom: dashboard panels show only "NO DATA" while Prometheus itself has data.
|
|
Three distinct root causes hit in one session — check in this order.
|
|
|
|
## 1. Datasource `uid` not pinned → Grafana generates a random one
|
|
|
|
If `uid:` is missing from the datasource in
|
|
`grafana/provisioning/datasources/datasources.yml`, Grafana assigns a RANDOM uid
|
|
at provisioning time (visible in its sqlite: `SELECT uid FROM data_source` →
|
|
e.g. `P8E80F9AEF21F6940`). Dashboards reference datasources by `"uid"` in each
|
|
panel (`"uid": "Prometheus"`), so panels silently find nothing → NO DATA.
|
|
|
|
Fix: pin the uid to match the dashboards:
|
|
```yaml
|
|
- name: Prometheus
|
|
type: prometheus
|
|
uid: Prometheus # must match panel "uid" refs
|
|
url: http://172.28.0.1:9090
|
|
isDefault: true
|
|
```
|
|
|
|
## 2. Prometheus on host network → `http://prometheus:9090` unresolvable from Grafana
|
|
|
|
Prometheus + blackbox run `network_mode: host` (needed to reach WireGuard
|
|
10.8.0.0/24), while Grafana is on the docker bridge network. The service name
|
|
`prometheus` does NOT resolve inside the bridge net
|
|
(`docker exec grafana getent hosts prometheus` → nothing). Loki, being on the
|
|
same bridge net, resolves fine — so one datasource works and the other doesn't.
|
|
|
|
Fix: point the datasource at the host via the bridge gateway:
|
|
```yaml
|
|
url: http://172.28.0.1:9090 # gateway IP = host as seen from Grafana's bridge net
|
|
```
|
|
Find the gateway: `docker inspect grafana --format '{{range .NetworkSettings.Networks}}{{.Gateway}} {{end}}'`
|
|
Verify from inside: `docker exec grafana wget -qO- http://172.28.0.1:9090/-/healthy`
|
|
Caveat: the bridge subnet is stable per docker daemon, but recreating the network
|
|
can change it.
|
|
|
|
## 3. Panels/alerts reference NON-EXISTENT metric names
|
|
|
|
Garage admin-API metrics (:3903) have NO `garage_` prefix. Only
|
|
`garage_build_info`, `garage_local_disk_avail/total`, `garage_replication_factor`
|
|
carry the prefix. Everything else is bare: `api_s3_request_counter`,
|
|
`block_resync_queue_length`, `block_resync_errored_blocks`, `cluster_*`
|
|
(connected_nodes, healthy, partitions_all_ok, layout_node_connected), `table_*`, `rpc_*`.
|
|
|
|
Phantom names that silently render NO DATA (never fire / never plot):
|
|
- `garage_block_count` → real: `block_resync_queue_length` + `block_resync_errored_blocks`
|
|
- `garage_rpc_node_health_is_up` → real: `cluster_layout_node_connected`
|
|
- `garage_api_s3_request_counter` → real: `api_s3_request_counter` (use `rate(api_s3_request_counter[5m])`)
|
|
- `garage_block_resync_error_count` → real: `block_resync_errored_blocks`
|
|
|
|
Check what actually exists before writing queries:
|
|
`curl 'http://127.0.0.1:9090/api/v1/label/__name__/values'` filtered by `job=garage`,
|
|
or `curl 'http://127.0.0.1:9090/api/v1/query?query={job="garage"}'` and list `__name__`.
|
|
|
|
## Provisioning reload semantics (Grafana)
|
|
|
|
- Dashboards: re-read automatically (~30s, updateIntervalSeconds).
|
|
- Datasources: provisioned ONLY at Grafana startup → after editing
|
|
datasources.yml you MUST `docker compose restart grafana`.
|
|
|
|
## Readable legends: relabel instance → hostname (not the raw scrape address)
|
|
|
|
By default `instance` = scrape address: tproxy shows `127.0.0.1:18081` (the LOCAL
|
|
end of the SSH tunnel — NOT "monitoring the local interface", just the transport),
|
|
garage shows `10.8.0.x:3903`. Fix at the SOURCE in prometheus.yml relabel_configs,
|
|
not per-panel legend, so panels/explore/alerts are consistent:
|
|
```yaml
|
|
relabel_configs:
|
|
- target_label: instance
|
|
replacement: vps03:8081 # tproxy
|
|
- source_labels: [__address__] # garage / node IP→hostname mapping
|
|
regex: 10.8.0.2:3903
|
|
target_label: instance
|
|
replacement: bigbox:3903
|
|
```
|
|
Result legends: `vps01:3903 / bigbox:3903 / vps02:3903` (garage),
|
|
`vps03:8081` (tproxy), `vps01:9100` etc (node). After changing relabel, restart
|
|
Prometheus and re-check `/api/v1/targets` — new labels apply immediately.
|
|
|
|
## Alert rules: same phantom-name trap
|
|
|
|
alerts.yml had `garage_block_resync_error_count` and `garage_rpc_node_health_is_up`
|
|
(never fire). Fixed to `block_resync_errored_blocks > 0` and
|
|
`cluster_layout_node_connected == 0`. Validate with
|
|
`docker exec prometheus promtool check config /etc/prometheus/prometheus.yml`
|
|
→ expect "6 rules found", all rules eval.
|
|
|
|
## Grafana sqlite inspection (no CLI auth needed)
|
|
|
|
Grafana's admin password may be unknown (changed in UI). Inspect state directly:
|
|
```
|
|
docker cp grafana:/var/lib/grafana/grafana.db /tmp/grafana.db
|
|
python3 -c "import sqlite3; print(sqlite3.connect('/tmp/grafana.db').execute('SELECT name,uid,url FROM data_source').fetchall())"
|
|
# dashboard JSON lives in table dashboard, column data
|
|
```
|