Files
openspec-lab/openspec/specs/vinograd-wan-monitoring/spec.md
T

85 lines
3.3 KiB
Markdown

# vinograd-wan-monitoring Specification
## Purpose
TBD - created by archiving change vinograd-rostelecom-channel-monitoring. Update Purpose after archive.
## Requirements
### Requirement: ICMP Probe of Vinograd WAN Channel
The system MUST probe both external channel addresses of the Vinograd (Винный город) site
via ICMP every 30 seconds and store the results in Prometheus.
| Address | Role |
|---|---|
| 83.239.50.145 | Gateway (шлюз Ростелеком) |
| 83.239.50.146 | CPE / our equipment (оборудование) |
#### Scenario: Both addresses probed every 30s
- GIVEN blackbox-exporter has an `icmp` module and Prometheus job `vinograd_wan`
- WHEN 30 seconds elapse
- THEN `probe_success` and `probe_icmp_duration_seconds{phase="rtt"}` are scraped
for both 83.239.50.145 and 83.239.50.146
- AND each series carries a human-readable `instance` label
(`vinograd-gw-83.239.50.145`, `vinograd-cpe-83.239.50.146`)
#### Scenario: Probe failure
- GIVEN an address does not answer ICMP (e.g. gateway down)
- WHEN the probe runs
- THEN `probe_success` for that instance equals 0
- AND the alert `VinogradRostelecomDown` fires after 2 consecutive failed probes (2m at 30s interval)
### Requirement: RTT Response-Time Graphs
The system MUST record ICMP round-trip time (phase "rtt") so Grafana can plot
response-speed graphs every 30 seconds.
#### Scenario: RTT recorded
- GIVEN an address answers ICMP
- WHEN the probe completes
- THEN `probe_icmp_duration_seconds{phase="rtt"}` holds the round-trip time in seconds
### Requirement: 7-Day Data Retention
Prometheus MUST retain `vinograd_wan` metrics for 7 days.
#### Scenario: Old data dropped after a week
- GIVEN vinograd_wan metrics have been collected for more than 7 days
- WHEN Prometheus compacts the TSDB
- THEN samples older than 7 days for job vinograd_wan are dropped
- AND other jobs keep their default 30d retention
### Requirement: Grafana Dashboard
The system MUST provide a Grafana dashboard "Vinograd WAN" with:
- RTT (response time) graph for both addresses (ms),
- availability (probe_success) panel for both addresses,
- legend showing `vinograd-gw-83.239.50.145` / `vinograd-cpe-83.239.50.146`.
#### Scenario: Dashboard shows data
- GIVEN Grafana has the Vinograd WAN dashboard provisioned
- WHEN a user opens it
- THEN it shows the RTT graph and availability of both channel addresses
### Requirement: Vinograd WAN dashboard opens without datasource errors
The Vinograd WAN dashboard (`/d/vinograd-wan/vinograd-wan`) MUST open and render
all panels WITHOUT the error "Datasource __grafana__ was not found".
- The dashboard JSON MUST NOT reference the built-in `__grafana__` datasource in
its `annotations.list` (it is not registered in this Grafana's database).
- `annotations.list` MUST be empty (`[]`), matching the working
`garage-cluster.json` dashboard.
#### Scenario: Dashboard renders without datasource error
- **WHEN** a user opens `https://grafana.nixg.ru/d/vinograd-wan/vinograd-wan`
- **THEN** the dashboard loads without the error "Datasource __grafana__ was not found"
- **AND** all panels render metric data from Prometheus (`uid: Prometheus`)
#### Scenario: Dashboard file stores no __grafana__ reference
- **WHEN** the file `grafana/dashboards/vinograd-wan.json` is parsed
- **THEN** `annotations.list` is `[]` OR contains no item whose
`datasource.uid` equals `__grafana__`