diff --git a/EXPERIENCE.md b/EXPERIENCE.md index 024b64c..3a14520 100644 --- a/EXPERIENCE.md +++ b/EXPERIENCE.md @@ -124,6 +124,100 @@ grafana.nixg.ru { - В логах caddy много ошибок renew для старых доменов — они имеют уже выпущенные сертификаты в caddy_data, работает всё. +## Опыт: дашборд Grafana пустой (NO DATA) — 3 причины подряд (2026-09-02) + +> Дата: 2026-09-02. Ситуация: после этапа 6 (tproxy) пользователь сообщил, +> что дашборд Garage Cluster в Grafana полностью пустой (NO DATA). + +### 12. NO DATA #1: у datasource не задан `uid` — Grafana генерит случайный + +В `grafana/provisioning/datasources/datasources.yml` у Prometheus-датасорса +НЕ был задан `uid`. Grafana при провиженинге присваивает датасорсу **случайный +UID** (в БД видно для Loki: `P8E80F9AEF21F6940`), а все панели дашборда +ссылались на `"uid": "Prometheus"`. Панели искали датасорс с таким uid — не +находили → NO DATA. + +**Решение:** в datasources.yml добавить датасорсу явный `uid: Prometheus` +(совпадение 1-в-1 с uid в панелях). Дашборды провиженятся каждые 30с, +а datasources — **только при старте** Grafana → после правки `docker compose +restart grafana`. Проверка из БД (`docker cp grafana:/var/lib/grafana/grafana.db`): +`SELECT name, uid, url FROM data_source` → uid стал `Prometheus`. + +### 13. NO DATA #2: Prometheus на host-сети, а datasource URL `http://prometheus:9090` + +Prometheus и blackbox запущены с `network_mode: host` (нужно для WG 10.8.0.0/24), +а Grafana — в bridge-сети. Внутри сети Grafana имя `prometheus` **не +резолвится** (проверено `docker exec grafana getent hosts prometheus` → +NO_RESOLVE), поэтому URL `http://prometheus:9090` недостижим → панели без +данных. Loki в той же bridge-сети — резолвится нормально. + +**Решение:** datasource URL заменить на `http://172.28.0.1:9090` — IP хоста +со стороны bridge-сети Grafana (шлюз сети = адрес хоста). Определяется так: +``` +docker inspect grafana --format '{{range .NetworkSettings.Networks}}{{.Gateway}} {{end}}' +# → 172.28.0.1 +``` +Проверено из контейнера: `wget http://172.28.0.1:9090/-/healthy` → OK. +ВАЖНО: шлюз bridge-сети docker стабилен (подсеть фиксированная), +но если пересоздать сеть — IP может смениться. + +### 14. NO DATA #3 (частично): панели ссылались на НЕСУЩЕСТВУЮЩИЕ метрики + +После починки datasource tproxy-панели ожили, а garage-панели (blocks, +RPC node health, S3 req/sec) остались пустыми. Причина: панели были написаны +под НЕСУЩЕСТВУЮЩИЕ имена метрик, которых в Prometheus нет: +- `garage_block_count` — в реальности `block_resync_queue_length` / `block_resync_errored_blocks` +- `garage_rpc_node_health_is_up` — в реальности `cluster_layout_node_connected` +- `garage_api_s3_request_counter` — в реальности `api_s3_request_counter` + +Метрики Garage из admin API (10.8.0.x:3903) **не имеют префикса `garage_`**: +`api_s3_request_counter`, `api_s3_request_duration_*`, `block_resync_*`, +`cluster_*` (connected_nodes, healthy, partitions_all_ok, layout_node_connected), +`table_*`, `rpc_*`. Префикс `garage_` есть только у `garage_build_info`, +`garage_local_disk_avail/total`, `garage_replication_factor`. + +Проверка реальных метрик: `curl 'http://127.0.0.1:9090/api/v1/label/__name__/values'` +и фильтр по `job=garage`. Панели переписаны на реальные метрики: +- Garage blocks (resync): `block_resync_queue_length` + `block_resync_errored_blocks` +- Garage node health (layout): `cluster_layout_node_connected` (легенда `{{ role_zone }}`) +- S3 requests /sec: `rate(api_s3_request_counter[5m])` + +### 15. Алерты ссылались на те же несуществующие метрики + +В alerts.yml алерты GarageResyncErrors и GarageNodeUnstable использовали те же +фантомные имена (`garage_block_resync_error_count`, `garage_rpc_node_health_is_up`), +поэтому никогда не сработали бы. Исправлено на реальные: +`block_resync_errored_blocks > 0` и `cluster_layout_node_connected == 0`. +После правки `promtool check config` → 6 rules found, все eval. + +### 16. Relabel `instance` → hostname вместо IP (читаемые легенды) + +По умолчанию instance = адрес скрейпа: у tproxy это `127.0.0.1:18081` +(локальный конец SSH-туннеля), у garage `10.8.0.x:3903`, у node `10.8.0.x:9100`. +В Grafana легенды показывали IP — некрасиво и непонятно. Не путать: это НЕ +«мониторинг локального интерфейса», просто лейбл instance = транспорт туннеля. + +**Решение (на уровне Prometheus, не панелей)** — relabel_configs в пром.yml: +```yaml +relabel_configs: + - target_label: instance + replacement: vps03:8081 # для job tproxy + - source_labels: [__address__] # для garage/node — маппинг IP → hostname + regex: 10.8.0.2:3903 + target_label: instance + replacement: bigbox:3903 +``` +Теперь легенды: `vps01:3903`, `bigbox:3903`, `vps02:3903` (garage), +`vps03:8081` (tproxy), `vps01:9100` и т.д. (node). Правило: правим источник +(prom.yml relabel), а не легенды в каждой панели — консистентно везде +(панели, explore, алерты). + +### 17. Grafana provisioning: дашборды — каждые 30с, datasources — только при старте + +Дашборды перечитываются автоматически (updateIntervalSeconds, по умолчанию 30с), +а datasources провиженятся только при старте Grafana. После правки +datasources.yml — обязателен `docker compose restart grafana`. + ## Опыт: tproxy-server (WEB Proxy для Telegram Desktop) на vps03 > Дата: 2026-08-31 @@ -163,14 +257,30 @@ backend_dial_failures_total, bytes_up_total, bytes_down_total, limit_hits_total. Ключевые для алертов: `backend_dial_failures_total` (рост = бэкенд недоступен), `limit_hits_total` (DDOS/перегруз), `sessions_live` (активность). -### C. Подключение к Prometheus (стек /opt/monitoring, bigbox) +### C. Подключение к Prometheus (стек /opt/monitoring, bigbox) — РЕШЕНО через SSH-туннель - vps03 НЕ в WG (10.8.0.0/24 = vps01/bigbox/vps02), поэтому прямого доступа к 8081 с bigbox нет. Prometheus в стеке — `network_mode: host` (видит внешние IP). -- План: nft-правило на vps03 (разрешить TCP 8081 с publIP bigbox 178.176.197.2), - job 'tproxy' в prometheus.yml, targets ['77.67.89.154:8081']. -- ВАЖНО: admin-эндпоинты слушают только loopback. Открывать наружу ТОЛЬКО - по источнику (ip saddr bigbox), не публиковать всем. +- Изначальный план (nft-правило на vps03: разрешить TCP 8081 с publIP bigbox + 178.176.197.2) **не сработал**: tproxy-server слушает `admin_listen: 127.0.0.1:8081`, + запрос с bigbox к `77.67.89.154:8081` вернул `Connection refused` (соединение + упёрлось в отсутствующего наружу слушателя, а не в файрвол). +- **Решение:** SSH-туннель bigbox → vps03 (паттерн systemd, как telegram-tunnel): + ```ini + # /etc/systemd/system/tproxy-tunnel.service (bigbox) + User=estorozhenko + ExecStart=/usr/bin/ssh -i /home/estorozhenko/.ssh/hostkeyVPS \ + -L 127.0.0.1:18081:127.0.0.1:8081 -N \ + -o ServerAliveInterval=30 -o ServerAliveCountMax=3 \ + -o ExitOnForwardFailure=yes -o StrictHostKeyChecking=accept-new \ + root@77.67.89.154 + Restart=always + ``` + Prometheus скрейпит `127.0.0.1:18081`. tproxy-server config/файрвол НЕ трогаются, + метрики остаются на loopback. +- **ВНИМАНИЕ с портом:** сначала взял `8081` — но он на bigbox уже занят ICQ-веб-чатом + (Converse.html на 0.0.0.0:8081), туннель конфликтовал и отдавал HTML чата вместо + метрик. Порт на bigbox выбирать свободный (взял 18081). - readyz ходит ко ВСЕМ профилям: если добавить профиль с мёртвым бэкендом — readyz станет 503 (фича, учтена при алертах). diff --git a/PLAN.md b/PLAN.md index bdaa438..d0d7308 100644 --- a/PLAN.md +++ b/PLAN.md @@ -24,14 +24,14 @@ **Результат:** все 3 ноды отдают метрики, кластер HEALTHY. -## Этап 2. Стек мониторинга на bigbox (в работе 🔄) +## Этап 2. Стек мониторинга на bigbox (выполнено ✅) - [x] Создать `/opt/monitoring/` с docker-compose.yml, prometheus.yml, alerts.yml - [x] Prometheus: таргеты garage x3 (:3903), health x3, node-exporters, self -- [ ] Node-exporter на 3 хостах (10.8.0.x:9100) -- [ ] Запустить `docker compose up -d` -- [ ] Проверить Prometheus :9090 (`up{job="garage"}`, targets) -- [ ] Grafana :3001: datasource Prometheus, дашборд Garage, alerts +- [x] Node-exporter на 3 хостах (10.8.0.x:9100) +- [x] Запустить `docker compose up -d` +- [x] Проверить Prometheus :9090 (`up{job="garage"}`, targets) +- [x] Grafana :3001: datasource Prometheus, дашборд Garage, alerts ## Этап 3. Логи (после метрик 🔄) @@ -78,41 +78,36 @@ Caddy → tproxy-server:8080 → MTProxy:2398. Admin-эндпоинты: Задача: -- [ ] Обеспечить доступ Prometheus (bigbox) к `:8081` на vps03. - vps03 НЕ в WG-сети (10.8.0.0/24 = vps01/bigbox/vps02). - Варианты: - a) открыть 8081 на vps03 для IP bigbox в nft (правило в - `/etc/tproxy-server/firewall.nft` или отдельный файл) — самый простой; - b) добавить vps03 в WG (если хочется закрытый контур); - c) node-exporter + textfile-коллектор — не наш случай (метрики уже в HTTP). - Рекомендация: (a). - **Конкретные шаги (рекомендуемый вариант a):** - - [ ] На vps03 (77.67.89.154) добавить nft-правило, разрешающее TCP 8081 - с публичного IP bigbox **178.176.197.2** (проверить актуальный IP - bigbox перед выполнением: `curl -s https://api.ipify.org`): - ``` - nft add rule inet filter input ip saddr 178.176.197.2 tcp dport 8081 accept - ``` - (или внести в `/etc/tproxy-server/firewall.nft`, если он в авто-загрузке) - - [ ] Проверить с bigbox: - `curl -s http://77.67.89.154:8081/metrics | head` - → должен вернуть метрики tproxy (не timeout/refused). - - [ ] Добавить job 'tproxy' в `/opt/monitoring/prometheus.yml`: - ```yaml - - job_name: 'tproxy' - static_configs: - - targets: ['77.67.89.154:8081'] - labels: - host: vps03 - service: tproxy - ``` - - [ ] `docker compose restart prometheus` в /opt/monitoring, - проверить `up{job="tproxy"}` на 127.0.0.1:9090 = 1. -- [ ] (Опционально) blackbox-job для внешнего HTTPS-чека vps03.nixg.ru: - `probe_success` — контролирует Caddy+сайт снаружи. -- [ ] Дашборд в Grafana: сессии, трафик (bytes_up/down rate), ошибки бэкенда. -- [ ] Алерт: `tproxy_backend_dial_failures_total` растёт, или `/readyz` 503 - (можно ч/з blackbox по HTTP :8081, если открыт). +- [x] **РЕШЕНО через SSH-туннель (не nft!)** — tproxy-server слушает `admin_listen: 127.0.0.1:8081` + только на loopback (осознанно), простое nft-правило из плана НЕ сработало бы: + соединение с bigbox упёрлось бы в `Connection refused` (проверено). + Поэтому добавлен systemd-сервис `tproxy-tunnel.service` на bigbox: + ``` + ssh -i /home/estorozhenko/.ssh/hostkeyVPS \ + -L 127.0.0.1:18081:127.0.0.1:8081 -N \ + root@77.67.89.154 + ``` + Порт 18081 (а НЕ 8081 — 8081 на bigbox занят ICQ-веб-чатом!). + Prometheus скрейпит `127.0.0.1:18081`. +- [x] Job 'tproxy' в prometheus.yml: `targets: ['127.0.0.1:18081']` → up=1, метрики скрейпятся. +- [x] Дашборд Garage Cluster дополнен 3 панелями tproxy: live sessions/streams, traffic /sec, + backend errors. +- [x] Алерты TProxyDown (critical) и TProxyBackendErrors (warning) в alerts.yml. +- [x] **Починен существующий баг: alerts.yml не был смонтирован в Prometheus** — garage-алерты + не работали (0 групп правил). Добавлен volume `./alerts.yml:/etc/prometheus/alerts.yml:ro` + в docker-compose.yml. Теперь загружено 6 правил (garage + tproxy). +- [x] **Починены пустые панели (NO DATA, 2026-09-02)**: три причины подряд — + (1) у datasource не был задан `uid` (Grafana генерила случайный), добавлен + `uid: Prometheus` в datasources.yml; (2) URL `http://prometheus:9090` не + резолвился из Grafana (Prometheus на host-сети), заменён на `http://172.28.0.1:9090` + (шлюз bridge-сети Grafana = адрес хоста); (3) панели и алерты ссылались + на несуществующие метрики (`garage_block_count`, `garage_rpc_node_health_is_up`, + `garage_api_s3_request_counter`) — переписаны на реальные (`block_resync_*`, + `cluster_layout_node_connected`, `rate(api_s3_request_counter[5m])`). Подробности — + EXPERIENCE.md п.12-17. +- [x] Relabel `instance` → hostname в prometheus.yml (garage/node/tproxy): + `10.8.0.x` → `vps01|bigbox|vps02`, `127.0.0.1:18081` → `vps03:8081`. + Легенды в Grafana — hostname, а не IP. Заметки: - `/metrics` и admin-эндпоинты слушают ТОЛЬКО loopback (127.0.0.1:8081). При diff --git a/README.md b/README.md index 0ec9afb..b8d0da9 100644 --- a/README.md +++ b/README.md @@ -112,20 +112,37 @@ curl -X POST -H "Authorization: token GITEA_TOK" \ |--------------------|--------------------------------|----------| | GarageNodeDown | `up{job="garage"} == 0` (2м) | critical | | GarageNoQuorum | `<2` нод up (2м) | critical | -| GarageResyncErrors | `garage_block_resync_error_count > 0` (10м) | warning | -| GarageNodeUnstable | RPC health down (5м) | warning | +| GarageResyncErrors | `block_resync_errored_blocks > 0` (10м) | warning | +| GarageNodeUnstable | `cluster_layout_node_connected == 0` (5м) | warning | +| TProxyDown | `up{job="tproxy"} == 0` (2м) | critical | +| TProxyBackendErrors| `increase(tproxy_backend_dial_failures_total[5m]) > 0` (10м) | warning | -## tproxy-server (vps03) — метрики WEB Proxy (этап 6, в работе) +Примечание: метрики Garage из admin API (:3903) НЕ имеют префикса `garage_` — +это `api_s3_request_counter`, `block_resync_*`, `cluster_*`. Префикс `garage_` +только у `garage_build_info`, `garage_local_disk_*`, `garage_replication_factor`. + +## tproxy-server (vps03) — метрики WEB Proxy (этап 6, РЕШЕНО ✅) tproxy-server (Telegram Desktop WEB Proxy) развёрнут на **vps03** (77.67.89.154), admin-эндпоинты на loopback :8081: `/healthz`, `/readyz`, `/metrics` -(Prometheus-формат, 13 счётчиков `tproxy_*`). vps03 НЕ в WG-сети, поэтому -для scrape с bigbox нужно: +(Prometheus-формат, 13 счётчиков `tproxy_*`). vps03 НЕ в WG-сети, а :8081 +слушает ТОЛЬКО loopback — поэтому вместо открытия порта наружу сделан +**SSH-туннель bigbox → vps03**: -1. nft-правило на vps03: разрешить TCP 8081 с publIP bigbox **178.176.197.2** - (`nft add rule inet filter input ip saddr 178.176.197.2 tcp dport 8081 accept`) -2. job `tproxy` в prometheus.yml: `targets: ['77.67.89.154:8081']` -3. `docker compose restart prometheus`, проверить `up{job="tproxy"}` +``` +systemd: tproxy-tunnel.service (bigbox) +ssh -i /home/estorozhenko/.ssh/hostkeyVPS \ + -L 127.0.0.1:18081:127.0.0.1:8081 -N root@77.67.89.154 +``` + +- Локальный порт **18081** (не 8081 — тот на bigbox занят ICQ-веб-чатом!) +- Prometheus скрейпит `127.0.0.1:18081` (job `tproxy`, host=vps03), но через + relabel_configs в prometheus.yml лейбл `instance` заменён на **`vps03:8081`**, + чтобы в Grafana легенды показывали hostname, а не IP туннеля. + Аналогично relabel сделан для garage (`10.8.0.x:3903` → `vps01|bigbox|vps02:3903`) + и node (`10.8.0.x:9100` → hostname). +- Дашборд: 3 панели tproxy (sessions/streams, traffic /sec, backend errors) +- Алерты: TProxyDown (critical), TProxyBackendErrors (warning) Подробности — в PLAN.md (этап 6) и EXPERIENCE.md. diff --git a/alerts.yml b/alerts.yml index 977df0a..62530bc 100644 --- a/alerts.yml +++ b/alerts.yml @@ -1,34 +1,47 @@ groups: - - name: garage - rules: - - alert: GarageNodeDown - expr: up{job="garage"} == 0 - for: 2m - labels: - severity: critical - annotations: - summary: "Garage node {{ $labels.instance }} is down" - - - alert: GarageNoQuorum - expr: (count(up{job="garage"} == 1) < 2) - for: 2m - labels: - severity: critical - annotations: - summary: "Garage cluster lost quorum (<2 nodes up)" - - - alert: GarageResyncErrors - expr: garage_block_resync_error_count > 0 - for: 10m - labels: - severity: warning - annotations: - summary: "Garage resync errors on {{ $labels.instance }}" - - - alert: GarageNodeUnstable - expr: garage_rpc_node_health_is_up == 0 - for: 5m - labels: - severity: warning - annotations: - summary: "Garage RPC health of some node is down" \ No newline at end of file +- name: garage + rules: + - alert: GarageNodeDown + expr: up{job="garage"} == 0 + for: 2m + labels: + severity: critical + annotations: + summary: Garage node {{ $labels.instance }} is down + - alert: GarageNoQuorum + expr: (count(up{job="garage"} == 1) < 2) + for: 2m + labels: + severity: critical + annotations: + summary: Garage cluster lost quorum (<2 nodes up) + - alert: GarageResyncErrors + expr: block_resync_errored_blocks > 0 + for: 10m + labels: + severity: warning + annotations: + summary: Garage resync errors on {{ $labels.instance }} + - alert: GarageNodeUnstable + expr: cluster_layout_node_connected == 0 + for: 5m + labels: + severity: warning + annotations: + summary: Garage RPC health of some node is down +- name: tproxy + rules: + - alert: TProxyDown + expr: up{job="tproxy"} == 0 + for: 2m + labels: + severity: critical + annotations: + summary: tproxy-server (vps03) is down or unreachable + - alert: TProxyBackendErrors + expr: increase(tproxy_backend_dial_failures_total[5m]) > 0 + for: 10m + labels: + severity: warning + annotations: + summary: tproxy-server backend dial failures (MTProxy unreachable) diff --git a/docker-compose.yml b/docker-compose.yml index 6594713..0e8acec 100644 --- a/docker-compose.yml +++ b/docker-compose.yml @@ -9,6 +9,7 @@ services: restart: unless-stopped volumes: - ./prometheus.yml:/etc/prometheus/prometheus.yml:ro + - ./alerts.yml:/etc/prometheus/alerts.yml:ro - ./prometheus-data:/prometheus command: - '--config.file=/etc/prometheus/prometheus.yml' diff --git a/grafana/dashboards/garage-cluster.json b/grafana/dashboards/garage-cluster.json index 0de5352..2ea1503 100644 --- a/grafana/dashboards/garage-cluster.json +++ b/grafana/dashboards/garage-cluster.json @@ -9,136 +9,689 @@ "links": [], "panels": [ { - "datasource": { "type": "prometheus", "uid": "Prometheus" }, + "datasource": { + "type": "prometheus", + "uid": "Prometheus" + }, "fieldConfig": { "defaults": { - "color": { "mode": "thresholds" }, + "color": { + "mode": "thresholds" + }, "mappings": [ - { "options": { "0": { "color": "red", "text": "DOWN" }, "1": { "color": "green", "text": "UP" } }, "type": "value" } + { + "options": { + "0": { + "color": "red", + "text": "DOWN" + }, + "1": { + "color": "green", + "text": "UP" + } + }, + "type": "value" + } ], - "thresholds": { "mode": "absolute", "steps": [ { "color": "red", "value": null }, { "color": "green", "value": 1 } ] } + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "red", + "value": null + }, + { + "color": "green", + "value": 1 + } + ] + } }, "overrides": [] }, - "gridPos": { "h": 8, "w": 6, "x": 0, "y": 0 }, + "gridPos": { + "h": 8, + "w": 6, + "x": 0, + "y": 0 + }, "id": 1, "options": { "colorMode": "background", "graphMode": "none", "justifyMode": "auto", "orientation": "auto", - "reduceOptions": { "calcs": [ "lastNotNull" ], "fields": "", "values": false }, + "reduceOptions": { + "calcs": [ + "lastNotNull" + ], + "fields": "", + "values": false + }, "textMode": "auto" }, "pluginVersion": "11.1.0", "targets": [ - { "datasource": { "type": "prometheus", "uid": "Prometheus" }, "expr": "up{job=\"garage\"}", "legendFormat": "{{ instance }}", "refId": "A" } + { + "datasource": { + "type": "prometheus", + "uid": "Prometheus" + }, + "expr": "up{job=\"garage\"}", + "legendFormat": "{{ instance }}", + "refId": "A" + } ], "title": "Garage nodes up", "type": "stat" }, { - "datasource": { "type": "prometheus", "uid": "Prometheus" }, + "datasource": { + "type": "prometheus", + "uid": "Prometheus" + }, "fieldConfig": { "defaults": { - "color": { "mode": "palette-classic" }, - "custom": { "axisCenteredZero": false, "axisColorMode": "text", "axisLabel": "", "axisPlacement": "auto", "drawStyle": "line", "fillOpacity": 10, "gradientMode": "none", "hideFrom": { "legend": false, "tooltip": false, "viz": false }, "lineInterpolation": "linear", "lineWidth": 1, "pointSize": 5, "scaleDistribution": { "type": "linear" }, "showPoints": "never", "spanNulls": false, "stacking": { "group": "A", "mode": "none" }, "thresholdsStyle": { "mode": "off" } }, + "color": { + "mode": "palette-classic" + }, + "custom": { + "axisCenteredZero": false, + "axisColorMode": "text", + "axisLabel": "", + "axisPlacement": "auto", + "drawStyle": "line", + "fillOpacity": 10, + "gradientMode": "none", + "hideFrom": { + "legend": false, + "tooltip": false, + "viz": false + }, + "lineInterpolation": "linear", + "lineWidth": 1, + "pointSize": 5, + "scaleDistribution": { + "type": "linear" + }, + "showPoints": "never", + "spanNulls": false, + "stacking": { + "group": "A", + "mode": "none" + }, + "thresholdsStyle": { + "mode": "off" + } + }, "mappings": [], - "thresholds": { "mode": "absolute", "steps": [ { "color": "green", "value": null } ] } + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + } }, "overrides": [] }, - "gridPos": { "h": 8, "w": 6, "x": 6, "y": 0 }, + "gridPos": { + "h": 8, + "w": 6, + "x": 6, + "y": 0 + }, "id": 2, "options": { - "legend": { "calcs": [], "displayMode": "list", "placement": "bottom", "showLegend": true }, - "tooltip": { "mode": "multi", "sort": "none" } + "legend": { + "calcs": [], + "displayMode": "list", + "placement": "bottom", + "showLegend": true + }, + "tooltip": { + "mode": "multi", + "sort": "none" + } }, "targets": [ - { "datasource": { "type": "prometheus", "uid": "Prometheus" }, "expr": "node_filesystem_avail_bytes{fstype!~\"tmpfs|overlay|squashfs\"} / 1024 / 1024 / 1024", "legendFormat": "{{ instance }}", "refId": "A" } + { + "datasource": { + "type": "prometheus", + "uid": "Prometheus" + }, + "expr": "node_filesystem_avail_bytes{fstype!~\"tmpfs|overlay|squashfs\"} / 1024 / 1024 / 1024", + "legendFormat": "{{ instance }}", + "refId": "A" + } ], "title": "Node disk free (GB)", "type": "timeseries" }, { - "datasource": { "type": "prometheus", "uid": "Prometheus" }, + "datasource": { + "type": "prometheus", + "uid": "Prometheus" + }, "fieldConfig": { "defaults": { - "color": { "mode": "palette-classic" }, - "custom": { "axisCenteredZero": false, "axisColorMode": "text", "axisLabel": "", "axisPlacement": "auto", "drawStyle": "line", "fillOpacity": 10, "gradientMode": "none", "hideFrom": { "legend": false, "tooltip": false, "viz": false }, "lineInterpolation": "linear", "lineWidth": 1, "pointSize": 5, "scaleDistribution": { "type": "linear" }, "showPoints": "never", "spanNulls": false, "stacking": { "group": "A", "mode": "none" }, "thresholdsStyle": { "mode": "off" } }, + "color": { + "mode": "palette-classic" + }, + "custom": { + "axisCenteredZero": false, + "axisColorMode": "text", + "axisLabel": "", + "axisPlacement": "auto", + "drawStyle": "line", + "fillOpacity": 10, + "gradientMode": "none", + "hideFrom": { + "legend": false, + "tooltip": false, + "viz": false + }, + "lineInterpolation": "linear", + "lineWidth": 1, + "pointSize": 5, + "scaleDistribution": { + "type": "linear" + }, + "showPoints": "never", + "spanNulls": false, + "stacking": { + "group": "A", + "mode": "none" + }, + "thresholdsStyle": { + "mode": "off" + } + }, "mappings": [], - "thresholds": { "mode": "absolute", "steps": [ { "color": "green", "value": null } ] } + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + } }, "overrides": [] }, - "gridPos": { "h": 8, "w": 12, "x": 12, "y": 0 }, + "gridPos": { + "h": 8, + "w": 12, + "x": 12, + "y": 0 + }, "id": 3, "options": { - "legend": { "calcs": [], "displayMode": "list", "placement": "bottom", "showLegend": true }, - "tooltip": { "mode": "multi", "sort": "none" } + "legend": { + "calcs": [], + "displayMode": "list", + "placement": "bottom", + "showLegend": true + }, + "tooltip": { + "mode": "multi", + "sort": "none" + } }, "targets": [ - { "datasource": { "type": "prometheus", "uid": "Prometheus" }, "expr": "garage_block_count", "legendFormat": "blocks {{ instance }}", "refId": "A" }, - { "datasource": { "type": "prometheus", "uid": "Prometheus" }, "expr": "garage_block_resync_queue_length", "legendFormat": "resync queue {{ instance }}", "refId": "B" } + { + "expr": "block_resync_queue_length", + "legendFormat": "resync queue {{ instance }}", + "refId": "A" + }, + { + "expr": "block_resync_errored_blocks", + "legendFormat": "errored {{ instance }}", + "refId": "B" + } ], - "title": "Garage blocks", + "title": "Garage blocks (resync)", "type": "timeseries" }, { - "datasource": { "type": "prometheus", "uid": "Prometheus" }, + "datasource": { + "type": "prometheus", + "uid": "Prometheus" + }, "fieldConfig": { "defaults": { - "color": { "mode": "thresholds" }, + "color": { + "mode": "thresholds" + }, "mappings": [], - "thresholds": { "mode": "absolute", "steps": [ { "color": "green", "value": null }, { "color": "red", "value": 1 } ] } + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + }, + { + "color": "red", + "value": 1 + } + ] + } }, "overrides": [] }, - "gridPos": { "h": 8, "w": 12, "x": 0, "y": 8 }, + "gridPos": { + "h": 8, + "w": 12, + "x": 0, + "y": 8 + }, "id": 4, "options": { - "legend": { "calcs": [], "displayMode": "list", "placement": "bottom", "showLegend": true }, - "tooltip": { "mode": "multi", "sort": "none" } + "legend": { + "calcs": [], + "displayMode": "list", + "placement": "bottom", + "showLegend": true + }, + "tooltip": { + "mode": "multi", + "sort": "none" + } }, "targets": [ - { "datasource": { "type": "prometheus", "uid": "Prometheus" }, "expr": "garage_rpc_node_health_is_up", "legendFormat": "{{ instance }}", "refId": "A" } + { + "expr": "cluster_layout_node_connected", + "legendFormat": "{{ role_zone }} {{ instance }}", + "refId": "A" + } ], - "title": "Garage RPC node health", + "title": "Garage node health (layout)", "type": "timeseries" }, { - "datasource": { "type": "prometheus", "uid": "Prometheus" }, + "datasource": { + "type": "prometheus", + "uid": "Prometheus" + }, "fieldConfig": { "defaults": { - "color": { "mode": "palette-classic" }, - "custom": { "axisCenteredZero": false, "axisColorMode": "text", "axisLabel": "", "axisPlacement": "auto", "drawStyle": "line", "fillOpacity": 10, "gradientMode": "none", "hideFrom": { "legend": false, "tooltip": false, "viz": false }, "lineInterpolation": "linear", "lineWidth": 1, "pointSize": 5, "scaleDistribution": { "type": "linear" }, "showPoints": "never", "spanNulls": false, "stacking": { "group": "A", "mode": "none" }, "thresholdsStyle": { "mode": "off" } }, + "color": { + "mode": "palette-classic" + }, + "custom": { + "axisCenteredZero": false, + "axisColorMode": "text", + "axisLabel": "", + "axisPlacement": "auto", + "drawStyle": "line", + "fillOpacity": 10, + "gradientMode": "none", + "hideFrom": { + "legend": false, + "tooltip": false, + "viz": false + }, + "lineInterpolation": "linear", + "lineWidth": 1, + "pointSize": 5, + "scaleDistribution": { + "type": "linear" + }, + "showPoints": "never", + "spanNulls": false, + "stacking": { + "group": "A", + "mode": "none" + }, + "thresholdsStyle": { + "mode": "off" + } + }, "mappings": [], - "thresholds": { "mode": "absolute", "steps": [ { "color": "green", "value": null } ] } + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + } }, "overrides": [] }, - "gridPos": { "h": 8, "w": 12, "x": 12, "y": 8 }, + "gridPos": { + "h": 8, + "w": 12, + "x": 12, + "y": 8 + }, "id": 5, "options": { - "legend": { "calcs": [], "displayMode": "list", "placement": "bottom", "showLegend": true }, - "tooltip": { "mode": "multi", "sort": "none" } + "legend": { + "calcs": [], + "displayMode": "list", + "placement": "bottom", + "showLegend": true + }, + "tooltip": { + "mode": "multi", + "sort": "none" + } }, "targets": [ - { "datasource": { "type": "prometheus", "uid": "Prometheus" }, "expr": "rate(garage_api_s3_request_counter[5m])", "legendFormat": "{{ instance }} {{ api_endpoint }}", "refId": "A" } + { + "expr": "rate(api_s3_request_counter[5m])", + "legendFormat": "{{ instance }} {{ api_endpoint }}", + "refId": "A" + } ], "title": "S3 requests /sec", "type": "timeseries" + }, + { + "datasource": { + "type": "prometheus", + "uid": "Prometheus" + }, + "fieldConfig": { + "defaults": { + "color": { + "mode": "palette-classic" + }, + "custom": { + "axisCenteredZero": false, + "axisColorMode": "text", + "axisLabel": "", + "axisPlacement": "auto", + "drawStyle": "line", + "fillOpacity": 10, + "gradientMode": "none", + "hideFrom": { + "legend": false, + "tooltip": false, + "viz": false + }, + "lineInterpolation": "linear", + "lineWidth": 1, + "pointSize": 5, + "scaleDistribution": { + "type": "linear" + }, + "showPoints": "never", + "spanNulls": false, + "stacking": { + "group": "A", + "mode": "none" + }, + "thresholdsStyle": { + "mode": "off" + } + }, + "mappings": [], + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + } + }, + "overrides": [] + }, + "gridPos": { + "h": 8, + "w": 6, + "x": 0, + "y": 16 + }, + "id": 6, + "options": { + "legend": { + "calcs": [], + "displayMode": "list", + "placement": "bottom", + "showLegend": true + }, + "tooltip": { + "mode": "multi", + "sort": "none" + } + }, + "title": "tproxy live sessions/streams", + "type": "timeseries", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "Prometheus" + }, + "expr": "tproxy_sessions_live", + "legendFormat": "sessions {{ instance }}", + "refId": "A" + }, + { + "datasource": { + "type": "prometheus", + "uid": "Prometheus" + }, + "expr": "tproxy_streams_live", + "legendFormat": "streams {{ instance }}", + "refId": "B" + } + ] + }, + { + "datasource": { + "type": "prometheus", + "uid": "Prometheus" + }, + "fieldConfig": { + "defaults": { + "color": { + "mode": "palette-classic" + }, + "custom": { + "axisCenteredZero": false, + "axisColorMode": "text", + "axisLabel": "", + "axisPlacement": "auto", + "drawStyle": "line", + "fillOpacity": 10, + "gradientMode": "none", + "hideFrom": { + "legend": false, + "tooltip": false, + "viz": false + }, + "lineInterpolation": "linear", + "lineWidth": 1, + "pointSize": 5, + "scaleDistribution": { + "type": "linear" + }, + "showPoints": "never", + "spanNulls": false, + "stacking": { + "group": "A", + "mode": "none" + }, + "thresholdsStyle": { + "mode": "off" + } + }, + "mappings": [], + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + } + }, + "overrides": [] + }, + "gridPos": { + "h": 8, + "w": 6, + "x": 6, + "y": 16 + }, + "id": 7, + "options": { + "legend": { + "calcs": [], + "displayMode": "list", + "placement": "bottom", + "showLegend": true + }, + "tooltip": { + "mode": "multi", + "sort": "none" + } + }, + "title": "tproxy traffic /sec", + "type": "timeseries", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "Prometheus" + }, + "expr": "rate(tproxy_bytes_up_total[5m])", + "legendFormat": "up {{ instance }}", + "refId": "A" + }, + { + "datasource": { + "type": "prometheus", + "uid": "Prometheus" + }, + "expr": "rate(tproxy_bytes_down_total[5m])", + "legendFormat": "down {{ instance }}", + "refId": "B" + } + ] + }, + { + "datasource": { + "type": "prometheus", + "uid": "Prometheus" + }, + "fieldConfig": { + "defaults": { + "color": { + "mode": "thresholds" + }, + "custom": { + "axisCenteredZero": false, + "axisColorMode": "text", + "axisLabel": "", + "axisPlacement": "auto", + "drawStyle": "line", + "fillOpacity": 10, + "gradientMode": "none", + "hideFrom": { + "legend": false, + "tooltip": false, + "viz": false + }, + "lineInterpolation": "linear", + "lineWidth": 1, + "pointSize": 5, + "scaleDistribution": { + "type": "linear" + }, + "showPoints": "never", + "spanNulls": false, + "stacking": { + "group": "A", + "mode": "none" + }, + "thresholdsStyle": { + "mode": "off" + } + }, + "mappings": [], + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + } + }, + "overrides": [] + }, + "gridPos": { + "h": 8, + "w": 12, + "x": 12, + "y": 16 + }, + "id": 8, + "options": { + "legend": { + "calcs": [], + "displayMode": "list", + "placement": "bottom", + "showLegend": true + }, + "tooltip": { + "mode": "multi", + "sort": "none" + } + }, + "title": "tproxy backend errors", + "type": "timeseries", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "Prometheus" + }, + "expr": "tproxy_backend_dial_failures_total", + "legendFormat": "dial failures {{ instance }}", + "refId": "A" + }, + { + "datasource": { + "type": "prometheus", + "uid": "Prometheus" + }, + "expr": "tproxy_streams_rejected_total", + "legendFormat": "rejected {{ instance }}", + "refId": "B" + } + ] } ], "refresh": "30s", "schemaVersion": 39, - "tags": [ "garage" ], - "templating": { "list": [] }, - "time": { "from": "now-6h", "to": "now" }, + "tags": [ + "garage" + ], + "templating": { + "list": [] + }, + "time": { + "from": "now-6h", + "to": "now" + }, "timepicker": {}, "timezone": "browser", "title": "Garage Cluster", "uid": "garage-cluster", - "version": 1, + "version": 2, "weekStart": "" } \ No newline at end of file diff --git a/grafana/provisioning/datasources/datasources.yml b/grafana/provisioning/datasources/datasources.yml index 54f1c69..e806c4f 100644 --- a/grafana/provisioning/datasources/datasources.yml +++ b/grafana/provisioning/datasources/datasources.yml @@ -4,7 +4,8 @@ datasources: - name: Prometheus type: prometheus access: proxy - url: http://prometheus:9090 + uid: Prometheus + url: http://172.28.0.1:9090 isDefault: true editable: true diff --git a/prometheus.yml b/prometheus.yml index a041aef..b4af32d 100644 --- a/prometheus.yml +++ b/prometheus.yml @@ -2,61 +2,109 @@ global: scrape_interval: 15s evaluation_interval: 15s external_labels: - monitor: 'garage-cluster' - + monitor: garage-cluster rule_files: - - /etc/prometheus/alerts.yml - +- /etc/prometheus/alerts.yml scrape_configs: - # Garage node metrics (admin API on WG) — Bearer auth - - job_name: 'garage' - metrics_path: /metrics - scheme: http - authorization: - type: Bearer - credentials: c213debe864061cefe501f7791d557b6dba701710b41791b4a4a0fb036690d1d - static_configs: - - targets: ['10.8.0.1:3903', '10.8.0.2:3903', '10.8.0.4:3903'] - labels: - cluster: garage - - # HTTP health probe via blackbox-exporter (returns probe_success) - - job_name: 'garage_health' - metrics_path: /probe - params: - module: [http_2xx] - static_configs: - - targets: ['http://10.8.0.1:3903/health'] - labels: - instance: vps01 - - targets: ['http://10.8.0.2:3903/health'] - labels: - instance: bigbox - - targets: ['http://10.8.0.4:3903/health'] - labels: - instance: vps02 - relabel_configs: - - source_labels: [__address__] - target_label: __param_target - - source_labels: [__param_target] - target_label: target - - target_label: __address__ - replacement: 127.0.0.1:9115 - - # Node exporter on each host (system metrics via WG) - - job_name: 'node' - static_configs: - - targets: ['10.8.0.2:9100'] - labels: - host: bigbox - - targets: ['10.8.0.1:9100'] - labels: - host: vps01 - - targets: ['10.8.0.4:9100'] - labels: - host: vps02 - - # Prometheus self - - job_name: 'prometheus' - static_configs: - - targets: ['localhost:9090'] \ No newline at end of file +- job_name: garage + metrics_path: /metrics + scheme: http + authorization: + type: Bearer + credentials: c213debe864061cefe501f7791d557b6dba701710b41791b4a4a0fb036690d1d + static_configs: + - targets: + - 10.8.0.1:3903 + - 10.8.0.2:3903 + - 10.8.0.4:3903 + labels: + cluster: garage + relabel_configs: + - source_labels: + - __address__ + regex: 10.8.0.1:3903 + target_label: instance + replacement: vps01:3903 + - source_labels: + - __address__ + regex: 10.8.0.2:3903 + target_label: instance + replacement: bigbox:3903 + - source_labels: + - __address__ + regex: 10.8.0.4:3903 + target_label: instance + replacement: vps02:3903 +- job_name: garage_health + metrics_path: /probe + params: + module: + - http_2xx + static_configs: + - targets: + - http://10.8.0.1:3903/health + labels: + instance: vps01 + - targets: + - http://10.8.0.2:3903/health + labels: + instance: bigbox + - targets: + - http://10.8.0.4:3903/health + labels: + instance: vps02 + relabel_configs: + - source_labels: + - __address__ + target_label: __param_target + - source_labels: + - __param_target + target_label: target + - target_label: __address__ + replacement: 127.0.0.1:9115 +- job_name: node + static_configs: + - targets: + - 10.8.0.2:9100 + labels: + host: bigbox + - targets: + - 10.8.0.1:9100 + labels: + host: vps01 + - targets: + - 10.8.0.4:9100 + labels: + host: vps02 + relabel_configs: + - source_labels: + - __address__ + regex: 10.8.0.2:9100 + target_label: instance + replacement: bigbox:9100 + - source_labels: + - __address__ + regex: 10.8.0.1:9100 + target_label: instance + replacement: vps01:9100 + - source_labels: + - __address__ + regex: 10.8.0.4:9100 + target_label: instance + replacement: vps02:9100 +- job_name: prometheus + static_configs: + - targets: + - localhost:9090 +- job_name: tproxy + static_configs: + - targets: + - 127.0.0.1:18081 + labels: + host: vps03 + service: tproxy + relabel_configs: + - target_label: instance + replacement: vps03:8081 + - target_label: host + replacement: vps03