Этап 6: починка NO DATA в Grafana (3 причины)+relabel instance→hostname; исправлены имена garage-метрик в панелях и алертах; опыт в EXPERIENCE.md

This commit is contained in:
kpa39l
2026-09-03 08:39:53 +00:00
parent 7437864b50
commit 66ff07c52a
8 changed files with 927 additions and 189 deletions
+115 -5
View File
@@ -124,6 +124,100 @@ grafana.nixg.ru {
- В логах caddy много ошибок renew для старых доменов — они имеют
уже выпущенные сертификаты в caddy_data, работает всё.
## Опыт: дашборд Grafana пустой (NO DATA) — 3 причины подряд (2026-09-02)
> Дата: 2026-09-02. Ситуация: после этапа 6 (tproxy) пользователь сообщил,
> что дашборд Garage Cluster в Grafana полностью пустой (NO DATA).
### 12. NO DATA #1: у datasource не задан `uid` — Grafana генерит случайный
В `grafana/provisioning/datasources/datasources.yml` у Prometheus-датасорса
НЕ был задан `uid`. Grafana при провиженинге присваивает датасорсу **случайный
UID** (в БД видно для Loki: `P8E80F9AEF21F6940`), а все панели дашборда
ссылались на `"uid": "Prometheus"`. Панели искали датасорс с таким uid — не
находили → NO DATA.
**Решение:** в datasources.yml добавить датасорсу явный `uid: Prometheus`
(совпадение 1-в-1 с uid в панелях). Дашборды провиженятся каждые 30с,
а datasources — **только при старте** Grafana → после правки `docker compose
restart grafana`. Проверка из БД (`docker cp grafana:/var/lib/grafana/grafana.db`):
`SELECT name, uid, url FROM data_source` → uid стал `Prometheus`.
### 13. NO DATA #2: Prometheus на host-сети, а datasource URL `http://prometheus:9090`
Prometheus и blackbox запущены с `network_mode: host` (нужно для WG 10.8.0.0/24),
а Grafana — в bridge-сети. Внутри сети Grafana имя `prometheus` **не
резолвится** (проверено `docker exec grafana getent hosts prometheus` →
NO_RESOLVE), поэтому URL `http://prometheus:9090` недостижим → панели без
данных. Loki в той же bridge-сети — резолвится нормально.
**Решение:** datasource URL заменить на `http://172.28.0.1:9090` — IP хоста
со стороны bridge-сети Grafana (шлюз сети = адрес хоста). Определяется так:
```
docker inspect grafana --format '{{range .NetworkSettings.Networks}}{{.Gateway}} {{end}}'
# → 172.28.0.1
```
Проверено из контейнера: `wget http://172.28.0.1:9090/-/healthy` → OK.
ВАЖНО: шлюз bridge-сети docker стабилен (подсеть фиксированная),
но если пересоздать сеть — IP может смениться.
### 14. NO DATA #3 (частично): панели ссылались на НЕСУЩЕСТВУЮЩИЕ метрики
После починки datasource tproxy-панели ожили, а garage-панели (blocks,
RPC node health, S3 req/sec) остались пустыми. Причина: панели были написаны
под НЕСУЩЕСТВУЮЩИЕ имена метрик, которых в Prometheus нет:
- `garage_block_count` — в реальности `block_resync_queue_length` / `block_resync_errored_blocks`
- `garage_rpc_node_health_is_up` — в реальности `cluster_layout_node_connected`
- `garage_api_s3_request_counter` — в реальности `api_s3_request_counter`
Метрики Garage из admin API (10.8.0.x:3903) **не имеют префикса `garage_`**:
`api_s3_request_counter`, `api_s3_request_duration_*`, `block_resync_*`,
`cluster_*` (connected_nodes, healthy, partitions_all_ok, layout_node_connected),
`table_*`, `rpc_*`. Префикс `garage_` есть только у `garage_build_info`,
`garage_local_disk_avail/total`, `garage_replication_factor`.
Проверка реальных метрик: `curl 'http://127.0.0.1:9090/api/v1/label/__name__/values'`
и фильтр по `job=garage`. Панели переписаны на реальные метрики:
- Garage blocks (resync): `block_resync_queue_length` + `block_resync_errored_blocks`
- Garage node health (layout): `cluster_layout_node_connected` (легенда `{{ role_zone }}`)
- S3 requests /sec: `rate(api_s3_request_counter[5m])`
### 15. Алерты ссылались на те же несуществующие метрики
В alerts.yml алерты GarageResyncErrors и GarageNodeUnstable использовали те же
фантомные имена (`garage_block_resync_error_count`, `garage_rpc_node_health_is_up`),
поэтому никогда не сработали бы. Исправлено на реальные:
`block_resync_errored_blocks > 0` и `cluster_layout_node_connected == 0`.
После правки `promtool check config` → 6 rules found, все eval.
### 16. Relabel `instance` → hostname вместо IP (читаемые легенды)
По умолчанию instance = адрес скрейпа: у tproxy это `127.0.0.1:18081`
(локальный конец SSH-туннеля), у garage `10.8.0.x:3903`, у node `10.8.0.x:9100`.
В Grafana легенды показывали IP — некрасиво и непонятно. Не путать: это НЕ
«мониторинг локального интерфейса», просто лейбл instance = транспорт туннеля.
**Решение (на уровне Prometheus, не панелей)** — relabel_configs в пром.yml:
```yaml
relabel_configs:
- target_label: instance
replacement: vps03:8081 # для job tproxy
- source_labels: [__address__] # для garage/node — маппинг IP → hostname
regex: 10.8.0.2:3903
target_label: instance
replacement: bigbox:3903
```
Теперь легенды: `vps01:3903`, `bigbox:3903`, `vps02:3903` (garage),
`vps03:8081` (tproxy), `vps01:9100` и т.д. (node). Правило: правим источник
(prom.yml relabel), а не легенды в каждой панели — консистентно везде
(панели, explore, алерты).
### 17. Grafana provisioning: дашборды — каждые 30с, datasources — только при старте
Дашборды перечитываются автоматически (updateIntervalSeconds, по умолчанию 30с),
а datasources провиженятся только при старте Grafana. После правки
datasources.yml — обязателен `docker compose restart grafana`.
## Опыт: tproxy-server (WEB Proxy для Telegram Desktop) на vps03
> Дата: 2026-08-31
@@ -163,14 +257,30 @@ backend_dial_failures_total, bytes_up_total, bytes_down_total, limit_hits_total.
Ключевые для алертов: `backend_dial_failures_total` (рост = бэкенд недоступен),
`limit_hits_total` (DDOS/перегруз), `sessions_live` (активность).
### C. Подключение к Prometheus (стек /opt/monitoring, bigbox)
### C. Подключение к Prometheus (стек /opt/monitoring, bigbox) — РЕШЕНО через SSH-туннель
- vps03 НЕ в WG (10.8.0.0/24 = vps01/bigbox/vps02), поэтому прямого доступа
к 8081 с bigbox нет. Prometheus в стеке — `network_mode: host` (видит внешние IP).
- План: nft-правило на vps03 (разрешить TCP 8081 с publIP bigbox 178.176.197.2),
job 'tproxy' в prometheus.yml, targets ['77.67.89.154:8081'].
- ВАЖНО: admin-эндпоинты слушают только loopback. Открывать наружу ТОЛЬКО
по источнику (ip saddr bigbox), не публиковать всем.
- Изначальный план (nft-правило на vps03: разрешить TCP 8081 с publIP bigbox
178.176.197.2) **не сработал**: tproxy-server слушает `admin_listen: 127.0.0.1:8081`,
запрос с bigbox к `77.67.89.154:8081` вернул `Connection refused` (соединение
упёрлось в отсутствующего наружу слушателя, а не в файрвол).
- **Решение:** SSH-туннель bigbox → vps03 (паттерн systemd, как telegram-tunnel):
```ini
# /etc/systemd/system/tproxy-tunnel.service (bigbox)
User=estorozhenko
ExecStart=/usr/bin/ssh -i /home/estorozhenko/.ssh/hostkeyVPS \
-L 127.0.0.1:18081:127.0.0.1:8081 -N \
-o ServerAliveInterval=30 -o ServerAliveCountMax=3 \
-o ExitOnForwardFailure=yes -o StrictHostKeyChecking=accept-new \
root@77.67.89.154
Restart=always
```
Prometheus скрейпит `127.0.0.1:18081`. tproxy-server config/файрвол НЕ трогаются,
метрики остаются на loopback.
- **ВНИМАНИЕ с портом:** сначала взял `8081` — но он на bigbox уже занят ICQ-веб-чатом
(Converse.html на 0.0.0.0:8081), туннель конфликтовал и отдавал HTML чата вместо
метрик. Порт на bigbox выбирать свободный (взял 18081).
- readyz ходит ко ВСЕМ профилям: если добавить профиль с мёртвым бэкендом —
readyz станет 503 (фича, учтена при алертах).
+35 -40
View File
@@ -24,14 +24,14 @@
**Результат:** все 3 ноды отдают метрики, кластер HEALTHY.
## Этап 2. Стек мониторинга на bigbox (в работе 🔄)
## Этап 2. Стек мониторинга на bigbox (выполнено ✅)
- [x] Создать `/opt/monitoring/` с docker-compose.yml, prometheus.yml, alerts.yml
- [x] Prometheus: таргеты garage x3 (:3903), health x3, node-exporters, self
- [ ] Node-exporter на 3 хостах (10.8.0.x:9100)
- [ ] Запустить `docker compose up -d`
- [ ] Проверить Prometheus :9090 (`up{job="garage"}`, targets)
- [ ] Grafana :3001: datasource Prometheus, дашборд Garage, alerts
- [x] Node-exporter на 3 хостах (10.8.0.x:9100)
- [x] Запустить `docker compose up -d`
- [x] Проверить Prometheus :9090 (`up{job="garage"}`, targets)
- [x] Grafana :3001: datasource Prometheus, дашборд Garage, alerts
## Этап 3. Логи (после метрик 🔄)
@@ -78,41 +78,36 @@ Caddy → tproxy-server:8080 → MTProxy:2398. Admin-эндпоинты:
Задача:
- [ ] Обеспечить доступ Prometheus (bigbox) к `:8081` на vps03.
vps03 НЕ в WG-сети (10.8.0.0/24 = vps01/bigbox/vps02).
Варианты:
a) открыть 8081 на vps03 для IP bigbox в nft (правило в
`/etc/tproxy-server/firewall.nft` или отдельный файл) — самый простой;
b) добавить vps03 в WG (если хочется закрытый контур);
c) node-exporter + textfile-коллектор — не наш случай (метрики уже в HTTP).
Рекомендация: (a).
**Конкретные шаги (рекомендуемый вариант a):**
- [ ] На vps03 (77.67.89.154) добавить nft-правило, разрешающее TCP 8081
с публичного IP bigbox **178.176.197.2** (проверить актуальный IP
bigbox перед выполнением: `curl -s https://api.ipify.org`):
```
nft add rule inet filter input ip saddr 178.176.197.2 tcp dport 8081 accept
```
(или внести в `/etc/tproxy-server/firewall.nft`, если он в авто-загрузке)
- [ ] Проверить с bigbox:
`curl -s http://77.67.89.154:8081/metrics | head`
→ должен вернуть метрики tproxy (не timeout/refused).
- [ ] Добавить job 'tproxy' в `/opt/monitoring/prometheus.yml`:
```yaml
- job_name: 'tproxy'
static_configs:
- targets: ['77.67.89.154:8081']
labels:
host: vps03
service: tproxy
```
- [ ] `docker compose restart prometheus` в /opt/monitoring,
проверить `up{job="tproxy"}` на 127.0.0.1:9090 = 1.
- [ ] (Опционально) blackbox-job для внешнего HTTPS-чека vps03.nixg.ru:
`probe_success` — контролирует Caddy+сайт снаружи.
- [ ] Дашборд в Grafana: сессии, трафик (bytes_up/down rate), ошибки бэкенда.
- [ ] Алерт: `tproxy_backend_dial_failures_total` растёт, или `/readyz` 503
(можно ч/з blackbox по HTTP :8081, если открыт).
- [x] **РЕШЕНО через SSH-туннель (не nft!)** — tproxy-server слушает `admin_listen: 127.0.0.1:8081`
только на loopback (осознанно), простое nft-правило из плана НЕ сработало бы:
соединение с bigbox упёрлось бы в `Connection refused` (проверено).
Поэтому добавлен systemd-сервис `tproxy-tunnel.service` на bigbox:
```
ssh -i /home/estorozhenko/.ssh/hostkeyVPS \
-L 127.0.0.1:18081:127.0.0.1:8081 -N \
root@77.67.89.154
```
Порт 18081 (а НЕ 8081 — 8081 на bigbox занят ICQ-веб-чатом!).
Prometheus скрейпит `127.0.0.1:18081`.
- [x] Job 'tproxy' в prometheus.yml: `targets: ['127.0.0.1:18081']` → up=1, метрики скрейпятся.
- [x] Дашборд Garage Cluster дополнен 3 панелями tproxy: live sessions/streams, traffic /sec,
backend errors.
- [x] Алерты TProxyDown (critical) и TProxyBackendErrors (warning) в alerts.yml.
- [x] **Починен существующий баг: alerts.yml не был смонтирован в Prometheus** — garage-алерты
не работали (0 групп правил). Добавлен volume `./alerts.yml:/etc/prometheus/alerts.yml:ro`
в docker-compose.yml. Теперь загружено 6 правил (garage + tproxy).
- [x] **Починены пустые панели (NO DATA, 2026-09-02)**: три причины подряд —
(1) у datasource не был задан `uid` (Grafana генерила случайный), добавлен
`uid: Prometheus` в datasources.yml; (2) URL `http://prometheus:9090` не
резолвился из Grafana (Prometheus на host-сети), заменён на `http://172.28.0.1:9090`
(шлюз bridge-сети Grafana = адрес хоста); (3) панели и алерты ссылались
на несуществующие метрики (`garage_block_count`, `garage_rpc_node_health_is_up`,
`garage_api_s3_request_counter`) — переписаны на реальные (`block_resync_*`,
`cluster_layout_node_connected`, `rate(api_s3_request_counter[5m])`). Подробности —
EXPERIENCE.md п.12-17.
- [x] Relabel `instance` → hostname в prometheus.yml (garage/node/tproxy):
`10.8.0.x` → `vps01|bigbox|vps02`, `127.0.0.1:18081` → `vps03:8081`.
Легенды в Grafana — hostname, а не IP.
Заметки:
- `/metrics` и admin-эндпоинты слушают ТОЛЬКО loopback (127.0.0.1:8081). При
+26 -9
View File
@@ -112,20 +112,37 @@ curl -X POST -H "Authorization: token GITEA_TOK" \
|--------------------|--------------------------------|----------|
| GarageNodeDown | `up{job="garage"} == 0` (2м) | critical |
| GarageNoQuorum | `<2` нод up (2м) | critical |
| GarageResyncErrors | `garage_block_resync_error_count > 0` (10м) | warning |
| GarageNodeUnstable | RPC health down (5м) | warning |
| GarageResyncErrors | `block_resync_errored_blocks > 0` (10м) | warning |
| GarageNodeUnstable | `cluster_layout_node_connected == 0` (5м) | warning |
| TProxyDown | `up{job="tproxy"} == 0` (2м) | critical |
| TProxyBackendErrors| `increase(tproxy_backend_dial_failures_total[5m]) > 0` (10м) | warning |
## tproxy-server (vps03) — метрики WEB Proxy (этап 6, в работе)
Примечание: метрики Garage из admin API (:3903) НЕ имеют префикса `garage_` —
это `api_s3_request_counter`, `block_resync_*`, `cluster_*`. Префикс `garage_`
только у `garage_build_info`, `garage_local_disk_*`, `garage_replication_factor`.
## tproxy-server (vps03) — метрики WEB Proxy (этап 6, РЕШЕНО ✅)
tproxy-server (Telegram Desktop WEB Proxy) развёрнут на **vps03** (77.67.89.154),
admin-эндпоинты на loopback :8081: `/healthz`, `/readyz`, `/metrics`
(Prometheus-формат, 13 счётчиков `tproxy_*`). vps03 НЕ в WG-сети, поэтому
для scrape с bigbox нужно:
(Prometheus-формат, 13 счётчиков `tproxy_*`). vps03 НЕ в WG-сети, а :8081
слушает ТОЛЬКО loopback — поэтому вместо открытия порта наружу сделан
**SSH-туннель bigbox → vps03**:
1. nft-правило на vps03: разрешить TCP 8081 с publIP bigbox **178.176.197.2**
(`nft add rule inet filter input ip saddr 178.176.197.2 tcp dport 8081 accept`)
2. job `tproxy` в prometheus.yml: `targets: ['77.67.89.154:8081']`
3. `docker compose restart prometheus`, проверить `up{job="tproxy"}`
```
systemd: tproxy-tunnel.service (bigbox)
ssh -i /home/estorozhenko/.ssh/hostkeyVPS \
-L 127.0.0.1:18081:127.0.0.1:8081 -N root@77.67.89.154
```
- Локальный порт **18081** (не 8081 — тот на bigbox занят ICQ-веб-чатом!)
- Prometheus скрейпит `127.0.0.1:18081` (job `tproxy`, host=vps03), но через
relabel_configs в prometheus.yml лейбл `instance` заменён на **`vps03:8081`**,
чтобы в Grafana легенды показывали hostname, а не IP туннеля.
Аналогично relabel сделан для garage (`10.8.0.x:3903` → `vps01|bigbox|vps02:3903`)
и node (`10.8.0.x:9100` → hostname).
- Дашборд: 3 панели tproxy (sessions/streams, traffic /sec, backend errors)
- Алерты: TProxyDown (critical), TProxyBackendErrors (warning)
Подробности — в PLAN.md (этап 6) и EXPERIENCE.md.
+46 -33
View File
@@ -1,34 +1,47 @@
groups:
- name: garage
rules:
- alert: GarageNodeDown
expr: up{job="garage"} == 0
for: 2m
labels:
severity: critical
annotations:
summary: "Garage node {{ $labels.instance }} is down"
- alert: GarageNoQuorum
expr: (count(up{job="garage"} == 1) < 2)
for: 2m
labels:
severity: critical
annotations:
summary: "Garage cluster lost quorum (<2 nodes up)"
- alert: GarageResyncErrors
expr: garage_block_resync_error_count > 0
for: 10m
labels:
severity: warning
annotations:
summary: "Garage resync errors on {{ $labels.instance }}"
- alert: GarageNodeUnstable
expr: garage_rpc_node_health_is_up == 0
for: 5m
labels:
severity: warning
annotations:
summary: "Garage RPC health of some node is down"
- name: garage
rules:
- alert: GarageNodeDown
expr: up{job="garage"} == 0
for: 2m
labels:
severity: critical
annotations:
summary: Garage node {{ $labels.instance }} is down
- alert: GarageNoQuorum
expr: (count(up{job="garage"} == 1) < 2)
for: 2m
labels:
severity: critical
annotations:
summary: Garage cluster lost quorum (<2 nodes up)
- alert: GarageResyncErrors
expr: block_resync_errored_blocks > 0
for: 10m
labels:
severity: warning
annotations:
summary: Garage resync errors on {{ $labels.instance }}
- alert: GarageNodeUnstable
expr: cluster_layout_node_connected == 0
for: 5m
labels:
severity: warning
annotations:
summary: Garage RPC health of some node is down
- name: tproxy
rules:
- alert: TProxyDown
expr: up{job="tproxy"} == 0
for: 2m
labels:
severity: critical
annotations:
summary: tproxy-server (vps03) is down or unreachable
- alert: TProxyBackendErrors
expr: increase(tproxy_backend_dial_failures_total[5m]) > 0
for: 10m
labels:
severity: warning
annotations:
summary: tproxy-server backend dial failures (MTProxy unreachable)
+1
View File
@@ -9,6 +9,7 @@ services:
restart: unless-stopped
volumes:
- ./prometheus.yml:/etc/prometheus/prometheus.yml:ro
- ./alerts.yml:/etc/prometheus/alerts.yml:ro
- ./prometheus-data:/prometheus
command:
- '--config.file=/etc/prometheus/prometheus.yml'
+598 -45
View File
@@ -9,136 +9,689 @@
"links": [],
"panels": [
{
"datasource": { "type": "prometheus", "uid": "Prometheus" },
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": { "mode": "thresholds" },
"color": {
"mode": "thresholds"
},
"mappings": [
{ "options": { "0": { "color": "red", "text": "DOWN" }, "1": { "color": "green", "text": "UP" } }, "type": "value" }
{
"options": {
"0": {
"color": "red",
"text": "DOWN"
},
"1": {
"color": "green",
"text": "UP"
}
},
"type": "value"
}
],
"thresholds": { "mode": "absolute", "steps": [ { "color": "red", "value": null }, { "color": "green", "value": 1 } ] }
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "red",
"value": null
},
{
"color": "green",
"value": 1
}
]
}
},
"overrides": []
},
"gridPos": { "h": 8, "w": 6, "x": 0, "y": 0 },
"gridPos": {
"h": 8,
"w": 6,
"x": 0,
"y": 0
},
"id": 1,
"options": {
"colorMode": "background",
"graphMode": "none",
"justifyMode": "auto",
"orientation": "auto",
"reduceOptions": { "calcs": [ "lastNotNull" ], "fields": "", "values": false },
"reduceOptions": {
"calcs": [
"lastNotNull"
],
"fields": "",
"values": false
},
"textMode": "auto"
},
"pluginVersion": "11.1.0",
"targets": [
{ "datasource": { "type": "prometheus", "uid": "Prometheus" }, "expr": "up{job=\"garage\"}", "legendFormat": "{{ instance }}", "refId": "A" }
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "up{job=\"garage\"}",
"legendFormat": "{{ instance }}",
"refId": "A"
}
],
"title": "Garage nodes up",
"type": "stat"
},
{
"datasource": { "type": "prometheus", "uid": "Prometheus" },
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": { "mode": "palette-classic" },
"custom": { "axisCenteredZero": false, "axisColorMode": "text", "axisLabel": "", "axisPlacement": "auto", "drawStyle": "line", "fillOpacity": 10, "gradientMode": "none", "hideFrom": { "legend": false, "tooltip": false, "viz": false }, "lineInterpolation": "linear", "lineWidth": 1, "pointSize": 5, "scaleDistribution": { "type": "linear" }, "showPoints": "never", "spanNulls": false, "stacking": { "group": "A", "mode": "none" }, "thresholdsStyle": { "mode": "off" } },
"color": {
"mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": { "mode": "absolute", "steps": [ { "color": "green", "value": null } ] }
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": { "h": 8, "w": 6, "x": 6, "y": 0 },
"gridPos": {
"h": 8,
"w": 6,
"x": 6,
"y": 0
},
"id": 2,
"options": {
"legend": { "calcs": [], "displayMode": "list", "placement": "bottom", "showLegend": true },
"tooltip": { "mode": "multi", "sort": "none" }
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"targets": [
{ "datasource": { "type": "prometheus", "uid": "Prometheus" }, "expr": "node_filesystem_avail_bytes{fstype!~\"tmpfs|overlay|squashfs\"} / 1024 / 1024 / 1024", "legendFormat": "{{ instance }}", "refId": "A" }
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "node_filesystem_avail_bytes{fstype!~\"tmpfs|overlay|squashfs\"} / 1024 / 1024 / 1024",
"legendFormat": "{{ instance }}",
"refId": "A"
}
],
"title": "Node disk free (GB)",
"type": "timeseries"
},
{
"datasource": { "type": "prometheus", "uid": "Prometheus" },
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": { "mode": "palette-classic" },
"custom": { "axisCenteredZero": false, "axisColorMode": "text", "axisLabel": "", "axisPlacement": "auto", "drawStyle": "line", "fillOpacity": 10, "gradientMode": "none", "hideFrom": { "legend": false, "tooltip": false, "viz": false }, "lineInterpolation": "linear", "lineWidth": 1, "pointSize": 5, "scaleDistribution": { "type": "linear" }, "showPoints": "never", "spanNulls": false, "stacking": { "group": "A", "mode": "none" }, "thresholdsStyle": { "mode": "off" } },
"color": {
"mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": { "mode": "absolute", "steps": [ { "color": "green", "value": null } ] }
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": { "h": 8, "w": 12, "x": 12, "y": 0 },
"gridPos": {
"h": 8,
"w": 12,
"x": 12,
"y": 0
},
"id": 3,
"options": {
"legend": { "calcs": [], "displayMode": "list", "placement": "bottom", "showLegend": true },
"tooltip": { "mode": "multi", "sort": "none" }
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"targets": [
{ "datasource": { "type": "prometheus", "uid": "Prometheus" }, "expr": "garage_block_count", "legendFormat": "blocks {{ instance }}", "refId": "A" },
{ "datasource": { "type": "prometheus", "uid": "Prometheus" }, "expr": "garage_block_resync_queue_length", "legendFormat": "resync queue {{ instance }}", "refId": "B" }
{
"expr": "block_resync_queue_length",
"legendFormat": "resync queue {{ instance }}",
"refId": "A"
},
{
"expr": "block_resync_errored_blocks",
"legendFormat": "errored {{ instance }}",
"refId": "B"
}
],
"title": "Garage blocks",
"title": "Garage blocks (resync)",
"type": "timeseries"
},
{
"datasource": { "type": "prometheus", "uid": "Prometheus" },
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": { "mode": "thresholds" },
"color": {
"mode": "thresholds"
},
"mappings": [],
"thresholds": { "mode": "absolute", "steps": [ { "color": "green", "value": null }, { "color": "red", "value": 1 } ] }
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
},
{
"color": "red",
"value": 1
}
]
}
},
"overrides": []
},
"gridPos": { "h": 8, "w": 12, "x": 0, "y": 8 },
"gridPos": {
"h": 8,
"w": 12,
"x": 0,
"y": 8
},
"id": 4,
"options": {
"legend": { "calcs": [], "displayMode": "list", "placement": "bottom", "showLegend": true },
"tooltip": { "mode": "multi", "sort": "none" }
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"targets": [
{ "datasource": { "type": "prometheus", "uid": "Prometheus" }, "expr": "garage_rpc_node_health_is_up", "legendFormat": "{{ instance }}", "refId": "A" }
{
"expr": "cluster_layout_node_connected",
"legendFormat": "{{ role_zone }} {{ instance }}",
"refId": "A"
}
],
"title": "Garage RPC node health",
"title": "Garage node health (layout)",
"type": "timeseries"
},
{
"datasource": { "type": "prometheus", "uid": "Prometheus" },
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": { "mode": "palette-classic" },
"custom": { "axisCenteredZero": false, "axisColorMode": "text", "axisLabel": "", "axisPlacement": "auto", "drawStyle": "line", "fillOpacity": 10, "gradientMode": "none", "hideFrom": { "legend": false, "tooltip": false, "viz": false }, "lineInterpolation": "linear", "lineWidth": 1, "pointSize": 5, "scaleDistribution": { "type": "linear" }, "showPoints": "never", "spanNulls": false, "stacking": { "group": "A", "mode": "none" }, "thresholdsStyle": { "mode": "off" } },
"color": {
"mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": { "mode": "absolute", "steps": [ { "color": "green", "value": null } ] }
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": { "h": 8, "w": 12, "x": 12, "y": 8 },
"gridPos": {
"h": 8,
"w": 12,
"x": 12,
"y": 8
},
"id": 5,
"options": {
"legend": { "calcs": [], "displayMode": "list", "placement": "bottom", "showLegend": true },
"tooltip": { "mode": "multi", "sort": "none" }
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"targets": [
{ "datasource": { "type": "prometheus", "uid": "Prometheus" }, "expr": "rate(garage_api_s3_request_counter[5m])", "legendFormat": "{{ instance }} {{ api_endpoint }}", "refId": "A" }
{
"expr": "rate(api_s3_request_counter[5m])",
"legendFormat": "{{ instance }} {{ api_endpoint }}",
"refId": "A"
}
],
"title": "S3 requests /sec",
"type": "timeseries"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 6,
"x": 0,
"y": 16
},
"id": 6,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"title": "tproxy live sessions/streams",
"type": "timeseries",
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "tproxy_sessions_live",
"legendFormat": "sessions {{ instance }}",
"refId": "A"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "tproxy_streams_live",
"legendFormat": "streams {{ instance }}",
"refId": "B"
}
]
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 6,
"x": 6,
"y": 16
},
"id": 7,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"title": "tproxy traffic /sec",
"type": "timeseries",
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "rate(tproxy_bytes_up_total[5m])",
"legendFormat": "up {{ instance }}",
"refId": "A"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "rate(tproxy_bytes_down_total[5m])",
"legendFormat": "down {{ instance }}",
"refId": "B"
}
]
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "thresholds"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 12,
"x": 12,
"y": 16
},
"id": 8,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"title": "tproxy backend errors",
"type": "timeseries",
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "tproxy_backend_dial_failures_total",
"legendFormat": "dial failures {{ instance }}",
"refId": "A"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "tproxy_streams_rejected_total",
"legendFormat": "rejected {{ instance }}",
"refId": "B"
}
]
}
],
"refresh": "30s",
"schemaVersion": 39,
"tags": [ "garage" ],
"templating": { "list": [] },
"time": { "from": "now-6h", "to": "now" },
"tags": [
"garage"
],
"templating": {
"list": []
},
"time": {
"from": "now-6h",
"to": "now"
},
"timepicker": {},
"timezone": "browser",
"title": "Garage Cluster",
"uid": "garage-cluster",
"version": 1,
"version": 2,
"weekStart": ""
}
@@ -4,7 +4,8 @@ datasources:
- name: Prometheus
type: prometheus
access: proxy
url: http://prometheus:9090
uid: Prometheus
url: http://172.28.0.1:9090
isDefault: true
editable: true
+104 -56
View File
@@ -2,61 +2,109 @@ global:
scrape_interval: 15s
evaluation_interval: 15s
external_labels:
monitor: 'garage-cluster'
monitor: garage-cluster
rule_files:
- /etc/prometheus/alerts.yml
- /etc/prometheus/alerts.yml
scrape_configs:
# Garage node metrics (admin API on WG) — Bearer auth
- job_name: 'garage'
metrics_path: /metrics
scheme: http
authorization:
type: Bearer
credentials: c213debe864061cefe501f7791d557b6dba701710b41791b4a4a0fb036690d1d
static_configs:
- targets: ['10.8.0.1:3903', '10.8.0.2:3903', '10.8.0.4:3903']
labels:
cluster: garage
# HTTP health probe via blackbox-exporter (returns probe_success)
- job_name: 'garage_health'
metrics_path: /probe
params:
module: [http_2xx]
static_configs:
- targets: ['http://10.8.0.1:3903/health']
labels:
instance: vps01
- targets: ['http://10.8.0.2:3903/health']
labels:
instance: bigbox
- targets: ['http://10.8.0.4:3903/health']
labels:
instance: vps02
relabel_configs:
- source_labels: [__address__]
target_label: __param_target
- source_labels: [__param_target]
target_label: target
- target_label: __address__
replacement: 127.0.0.1:9115
# Node exporter on each host (system metrics via WG)
- job_name: 'node'
static_configs:
- targets: ['10.8.0.2:9100']
labels:
host: bigbox
- targets: ['10.8.0.1:9100']
labels:
host: vps01
- targets: ['10.8.0.4:9100']
labels:
host: vps02
# Prometheus self
- job_name: 'prometheus'
static_configs:
- targets: ['localhost:9090']
- job_name: garage
metrics_path: /metrics
scheme: http
authorization:
type: Bearer
credentials: c213debe864061cefe501f7791d557b6dba701710b41791b4a4a0fb036690d1d
static_configs:
- targets:
- 10.8.0.1:3903
- 10.8.0.2:3903
- 10.8.0.4:3903
labels:
cluster: garage
relabel_configs:
- source_labels:
- __address__
regex: 10.8.0.1:3903
target_label: instance
replacement: vps01:3903
- source_labels:
- __address__
regex: 10.8.0.2:3903
target_label: instance
replacement: bigbox:3903
- source_labels:
- __address__
regex: 10.8.0.4:3903
target_label: instance
replacement: vps02:3903
- job_name: garage_health
metrics_path: /probe
params:
module:
- http_2xx
static_configs:
- targets:
- http://10.8.0.1:3903/health
labels:
instance: vps01
- targets:
- http://10.8.0.2:3903/health
labels:
instance: bigbox
- targets:
- http://10.8.0.4:3903/health
labels:
instance: vps02
relabel_configs:
- source_labels:
- __address__
target_label: __param_target
- source_labels:
- __param_target
target_label: target
- target_label: __address__
replacement: 127.0.0.1:9115
- job_name: node
static_configs:
- targets:
- 10.8.0.2:9100
labels:
host: bigbox
- targets:
- 10.8.0.1:9100
labels:
host: vps01
- targets:
- 10.8.0.4:9100
labels:
host: vps02
relabel_configs:
- source_labels:
- __address__
regex: 10.8.0.2:9100
target_label: instance
replacement: bigbox:9100
- source_labels:
- __address__
regex: 10.8.0.1:9100
target_label: instance
replacement: vps01:9100
- source_labels:
- __address__
regex: 10.8.0.4:9100
target_label: instance
replacement: vps02:9100
- job_name: prometheus
static_configs:
- targets:
- localhost:9090
- job_name: tproxy
static_configs:
- targets:
- 127.0.0.1:18081
labels:
host: vps03
service: tproxy
relabel_configs:
- target_label: instance
replacement: vps03:8081
- target_label: host
replacement: vps03