4.6 KiB
Grafana NO DATA debugging + monitoring hygiene (2026-09-02 session)
Symptom: dashboard panels show only "NO DATA" while Prometheus itself has data. Three distinct root causes hit in one session — check in this order.
1. Datasource uid not pinned → Grafana generates a random one
If uid: is missing from the datasource in
grafana/provisioning/datasources/datasources.yml, Grafana assigns a RANDOM uid
at provisioning time (visible in its sqlite: SELECT uid FROM data_source →
e.g. P8E80F9AEF21F6940). Dashboards reference datasources by "uid" in each
panel ("uid": "Prometheus"), so panels silently find nothing → NO DATA.
Fix: pin the uid to match the dashboards:
- name: Prometheus
type: prometheus
uid: Prometheus # must match panel "uid" refs
url: http://172.28.0.1:9090
isDefault: true
2. Prometheus on host network → http://prometheus:9090 unresolvable from Grafana
Prometheus + blackbox run network_mode: host (needed to reach WireGuard
10.8.0.0/24), while Grafana is on the docker bridge network. The service name
prometheus does NOT resolve inside the bridge net
(docker exec grafana getent hosts prometheus → nothing). Loki, being on the
same bridge net, resolves fine — so one datasource works and the other doesn't.
Fix: point the datasource at the host via the bridge gateway:
url: http://172.28.0.1:9090 # gateway IP = host as seen from Grafana's bridge net
Find the gateway: docker inspect grafana --format '{{range .NetworkSettings.Networks}}{{.Gateway}} {{end}}'
Verify from inside: docker exec grafana wget -qO- http://172.28.0.1:9090/-/healthy
Caveat: the bridge subnet is stable per docker daemon, but recreating the network
can change it.
3. Panels/alerts reference NON-EXISTENT metric names
Garage admin-API metrics (:3903) have NO garage_ prefix. Only
garage_build_info, garage_local_disk_avail/total, garage_replication_factor
carry the prefix. Everything else is bare: api_s3_request_counter,
block_resync_queue_length, block_resync_errored_blocks, cluster_*
(connected_nodes, healthy, partitions_all_ok, layout_node_connected), table_*, rpc_*.
Phantom names that silently render NO DATA (never fire / never plot):
garage_block_count→ real:block_resync_queue_length+block_resync_errored_blocksgarage_rpc_node_health_is_up→ real:cluster_layout_node_connectedgarage_api_s3_request_counter→ real:api_s3_request_counter(userate(api_s3_request_counter[5m]))garage_block_resync_error_count→ real:block_resync_errored_blocks
Check what actually exists before writing queries:
curl 'http://127.0.0.1:9090/api/v1/label/__name__/values' filtered by job=garage,
or curl 'http://127.0.0.1:9090/api/v1/query?query={job="garage"}' and list __name__.
Provisioning reload semantics (Grafana)
- Dashboards: re-read automatically (~30s, updateIntervalSeconds).
- Datasources: provisioned ONLY at Grafana startup → after editing
datasources.yml you MUST
docker compose restart grafana.
Readable legends: relabel instance → hostname (not the raw scrape address)
By default instance = scrape address: tproxy shows 127.0.0.1:18081 (the LOCAL
end of the SSH tunnel — NOT "monitoring the local interface", just the transport),
garage shows 10.8.0.x:3903. Fix at the SOURCE in prometheus.yml relabel_configs,
not per-panel legend, so panels/explore/alerts are consistent:
relabel_configs:
- target_label: instance
replacement: vps03:8081 # tproxy
- source_labels: [__address__] # garage / node IP→hostname mapping
regex: 10.8.0.2:3903
target_label: instance
replacement: bigbox:3903
Result legends: vps01:3903 / bigbox:3903 / vps02:3903 (garage),
vps03:8081 (tproxy), vps01:9100 etc (node). After changing relabel, restart
Prometheus and re-check /api/v1/targets — new labels apply immediately.
Alert rules: same phantom-name trap
alerts.yml had garage_block_resync_error_count and garage_rpc_node_health_is_up
(never fire). Fixed to block_resync_errored_blocks > 0 and
cluster_layout_node_connected == 0. Validate with
docker exec prometheus promtool check config /etc/prometheus/prometheus.yml
→ expect "6 rules found", all rules eval.
Grafana sqlite inspection (no CLI auth needed)
Grafana's admin password may be unknown (changed in UI). Inspect state directly:
docker cp grafana:/var/lib/grafana/grafana.db /tmp/grafana.db
python3 -c "import sqlite3; print(sqlite3.connect('/tmp/grafana.db').execute('SELECT name,uid,url FROM data_source').fetchall())"
# dashboard JSON lives in table dashboard, column data