Files
garage-s3-cluster/references/grafana-nodata-debugging.md
T
2026-09-06 13:50:54 +00:00

4.6 KiB

Grafana NO DATA debugging + monitoring hygiene (2026-09-02 session)

Symptom: dashboard panels show only "NO DATA" while Prometheus itself has data. Three distinct root causes hit in one session — check in this order.

1. Datasource uid not pinned → Grafana generates a random one

If uid: is missing from the datasource in grafana/provisioning/datasources/datasources.yml, Grafana assigns a RANDOM uid at provisioning time (visible in its sqlite: SELECT uid FROM data_source → e.g. P8E80F9AEF21F6940). Dashboards reference datasources by "uid" in each panel ("uid": "Prometheus"), so panels silently find nothing → NO DATA.

Fix: pin the uid to match the dashboards:

- name: Prometheus
  type: prometheus
  uid: Prometheus        # must match panel "uid" refs
  url: http://172.28.0.1:9090
  isDefault: true

2. Prometheus on host network → http://prometheus:9090 unresolvable from Grafana

Prometheus + blackbox run network_mode: host (needed to reach WireGuard 10.8.0.0/24), while Grafana is on the docker bridge network. The service name prometheus does NOT resolve inside the bridge net (docker exec grafana getent hosts prometheus → nothing). Loki, being on the same bridge net, resolves fine — so one datasource works and the other doesn't.

Fix: point the datasource at the host via the bridge gateway:

url: http://172.28.0.1:9090   # gateway IP = host as seen from Grafana's bridge net

Find the gateway: docker inspect grafana --format '{{range .NetworkSettings.Networks}}{{.Gateway}} {{end}}' Verify from inside: docker exec grafana wget -qO- http://172.28.0.1:9090/-/healthy Caveat: the bridge subnet is stable per docker daemon, but recreating the network can change it.

3. Panels/alerts reference NON-EXISTENT metric names

Garage admin-API metrics (:3903) have NO garage_ prefix. Only garage_build_info, garage_local_disk_avail/total, garage_replication_factor carry the prefix. Everything else is bare: api_s3_request_counter, block_resync_queue_length, block_resync_errored_blocks, cluster_* (connected_nodes, healthy, partitions_all_ok, layout_node_connected), table_*, rpc_*.

Phantom names that silently render NO DATA (never fire / never plot):

  • garage_block_count → real: block_resync_queue_length + block_resync_errored_blocks
  • garage_rpc_node_health_is_up → real: cluster_layout_node_connected
  • garage_api_s3_request_counter → real: api_s3_request_counter (use rate(api_s3_request_counter[5m]))
  • garage_block_resync_error_count → real: block_resync_errored_blocks

Check what actually exists before writing queries: curl 'http://127.0.0.1:9090/api/v1/label/__name__/values' filtered by job=garage, or curl 'http://127.0.0.1:9090/api/v1/query?query={job="garage"}' and list __name__.

Provisioning reload semantics (Grafana)

  • Dashboards: re-read automatically (~30s, updateIntervalSeconds).
  • Datasources: provisioned ONLY at Grafana startup → after editing datasources.yml you MUST docker compose restart grafana.

Readable legends: relabel instance → hostname (not the raw scrape address)

By default instance = scrape address: tproxy shows 127.0.0.1:18081 (the LOCAL end of the SSH tunnel — NOT "monitoring the local interface", just the transport), garage shows 10.8.0.x:3903. Fix at the SOURCE in prometheus.yml relabel_configs, not per-panel legend, so panels/explore/alerts are consistent:

relabel_configs:
  - target_label: instance
    replacement: vps03:8081                 # tproxy
  - source_labels: [__address__]            # garage / node IP→hostname mapping
    regex: 10.8.0.2:3903
    target_label: instance
    replacement: bigbox:3903

Result legends: vps01:3903 / bigbox:3903 / vps02:3903 (garage), vps03:8081 (tproxy), vps01:9100 etc (node). After changing relabel, restart Prometheus and re-check /api/v1/targets — new labels apply immediately.

Alert rules: same phantom-name trap

alerts.yml had garage_block_resync_error_count and garage_rpc_node_health_is_up (never fire). Fixed to block_resync_errored_blocks > 0 and cluster_layout_node_connected == 0. Validate with docker exec prometheus promtool check config /etc/prometheus/prometheus.yml → expect "6 rules found", all rules eval.

Grafana sqlite inspection (no CLI auth needed)

Grafana's admin password may be unknown (changed in UI). Inspect state directly:

docker cp grafana:/var/lib/grafana/grafana.db /tmp/grafana.db
python3 -c "import sqlite3; print(sqlite3.connect('/tmp/grafana.db').execute('SELECT name,uid,url FROM data_source').fetchall())"
# dashboard JSON lives in table dashboard, column data