mirror of
https://gitverse.ru/kpa39l/openspec-lab.git
synced 2026-09-29 09:15:01 +00:00
vinograd-rostelecom-channel-monitoring: archived (Vinograd WAN ICMP monitoring, /opt/monitoring)
This commit is contained in:
+59
@@ -0,0 +1,59 @@
|
||||
# Delta for vinograd-wan-monitoring
|
||||
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: ICMP Probe of Vinograd WAN Channel
|
||||
|
||||
The system MUST probe both external channel addresses of the Vinograd (Винный город) site
|
||||
via ICMP every 30 seconds and store the results in Prometheus.
|
||||
|
||||
| Address | Role |
|
||||
|---|---|
|
||||
| 83.239.50.145 | Gateway (шлюз Ростелеком) |
|
||||
| 83.239.50.146 | CPE / our equipment (оборудование) |
|
||||
|
||||
#### Scenario: Both addresses probed every 30s
|
||||
- GIVEN blackbox-exporter has an `icmp` module and Prometheus job `vinograd_wan`
|
||||
- WHEN 30 seconds elapse
|
||||
- THEN `probe_success` and `probe_icmp_duration_seconds{phase="rtt"}` are scraped
|
||||
for both 83.239.50.145 and 83.239.50.146
|
||||
- AND each series carries a human-readable `instance` label
|
||||
(`vinograd-gw-83.239.50.145`, `vinograd-cpe-83.239.50.146`)
|
||||
|
||||
#### Scenario: Probe failure
|
||||
- GIVEN an address does not answer ICMP (e.g. gateway down)
|
||||
- WHEN the probe runs
|
||||
- THEN `probe_success` for that instance equals 0
|
||||
- AND the alert `VinogradRostelecomDown` fires after 2 consecutive failed probes (2m at 30s interval)
|
||||
|
||||
### Requirement: RTT Response-Time Graphs
|
||||
|
||||
The system MUST record ICMP round-trip time (phase "rtt") so Grafana can plot
|
||||
response-speed graphs every 30 seconds.
|
||||
|
||||
#### Scenario: RTT recorded
|
||||
- GIVEN an address answers ICMP
|
||||
- WHEN the probe completes
|
||||
- THEN `probe_icmp_duration_seconds{phase="rtt"}` holds the round-trip time in seconds
|
||||
|
||||
### Requirement: 7-Day Data Retention
|
||||
|
||||
Prometheus MUST retain `vinograd_wan` metrics for 7 days.
|
||||
|
||||
#### Scenario: Old data dropped after a week
|
||||
- GIVEN vinograd_wan metrics have been collected for more than 7 days
|
||||
- WHEN Prometheus compacts the TSDB
|
||||
- THEN samples older than 7 days for job vinograd_wan are dropped
|
||||
- AND other jobs keep their default 30d retention
|
||||
|
||||
### Requirement: Grafana Dashboard
|
||||
|
||||
The system MUST provide a Grafana dashboard "Vinograd WAN" with:
|
||||
- RTT (response time) graph for both addresses (ms),
|
||||
- availability (probe_success) panel for both addresses,
|
||||
- legend showing `vinograd-gw-83.239.50.145` / `vinograd-cpe-83.239.50.146`.
|
||||
|
||||
#### Scenario: Dashboard shows data
|
||||
- GIVEN Grafana has the Vinograd WAN dashboard provisioned
|
||||
- WHEN a user opens it
|
||||
- THEN it shows the RTT graph and availability of both channel addresses
|
||||
Reference in New Issue
Block a user