Compare commits

..

16 Commits

Author SHA1 Message Date
kpa39l ee9249697d feat(gotosocial): Grafana dashboard GoToSocial (20 panels); fix dashboard provisioning provider overlap; openspec: gotosocial-monitoring 2026-09-19 18:44:24 +00:00
kpa39l 9cde035e5a chore(openspec): archive 2026-09-13-vesti-alerts → specs/vesti-alerts 2026-09-13 15:40:29 +00:00
kpa39l 62a4f1a0db feat(openspec): vesti-alerts — 6 алертов на компоненты VESTI (web/publisher/tunnel/db/stale); OpenSpec change 2026-09-13 15:34:47 +00:00
kpa39l df1ebed76d docs: убрать ошибочное упоминание зависания systemd1; зафиксировать таймер vesti-metrics как активный 2026-09-13 15:23:59 +00:00
kpa39l 847fcfad9c feat: VESTI dashboard — textfile-коллектор vesti_* метрик + дашборд (доступность компонентов) 2026-09-13 15:06:10 +00:00
kpa39l 3c5e9e5a71 OpenSpec: разнести openspec по проектам 2026-09-11 17:31:01 +00:00
kpa39l 9f2619039a docs: закрытие сессии (вечер) — STATUS/TODO/WALKTHROUGH/README: подпапки дашбордов 2026-09-08 18:26:42 +00:00
kpa39l 4b0539d8da refactor: дашборды в подпапки nodes/ и vinogorod/ + provisioning folder nodes/vinogorod 2026-09-08 18:08:50 +00:00
kpa39l 4382f9c0ff docs: закрытие сессии 2026-09-08 — STATUS/TODO/WALKTHROUGH/PRD 2026-09-08 16:58:45 +00:00
kpa39l eaafb1ff3a docs: EXPERIENCE.md — грабли per-node дашбордов (node_network_*, speed=-125000 на виртуалках) 2026-09-08 14:26:32 +00:00
kpa39l 6e48fd81fa feat: per-node dashboards (vps01/vps02/vps03/bigbox) с RX/TX + утилизацией канала по физическому интерфейсу 2026-09-08 14:26:10 +00:00
kpa39l 1b2a2357d1 feat: vps03 в WireGuard (10.8.0.3) + node-exporter + дашборд Nodes (диски/память/сеть/доступность) 2026-09-08 14:12:27 +00:00
Evgeny Storozhenko 5ea6fd290a Fix vinograd-wan dashboard: убрана ссылка на __grafana__ в annotations (Datasource not found); опыт #24 2026-09-08 13:41:13 +00:00
Evgeny Storozhenko da1746c455 Grafana: read-only пользователь it@vinogorod.ru (Viewer); опыт: provisioning users в OSS 11.1 не работает — только UI 2026-09-08 13:26:21 +00:00
estorozhenko b4f4493961 Vinograd WAN (Ростелеком): ICMP-мониторинг канала Винный город — шлюз 83.239.50.145 + оборудование 83.239.50.146, scrape 30s, RTT графики, алерт VinogradRostelecomDown, дашборд Vinograd WAN 2026-09-08 06:01:01 +00:00
kpa39l 66ff07c52a Этап 6: починка NO DATA в Grafana (3 причины)+relabel instance→hostname; исправлены имена garage-метрик в панелях и алертах; опыт в EXPERIENCE.md 2026-09-03 08:39:53 +00:00
59 changed files with 9398 additions and 195 deletions
@@ -0,0 +1,188 @@
---
name: openspec-apply-change
description: Implement tasks from an OpenSpec change. Use when the user wants to start implementing, continue implementation, or work through tasks.
allowed-tools: Bash(openspec:*)
license: MIT
compatibility: Requires openspec CLI.
metadata:
author: openspec
version: "1.0"
generatedBy: "1.12.0"
---
Implement tasks from an OpenSpec change.
**Store selection:** If the user names a store (a store is a standalone OpenSpec repo registered on this machine) or the work lives in one, run `openspec store list --json` to discover registered store ids, then pass `--store <id>` on the commands that read or write specs and changes (`new change`, `status`, `instructions`, `list`, `show`, `validate`, `archive`, `doctor`, `context`, `schemas`, `view`). Once selected, treat `--store <id>` as sticky for the rest of the workflow. Every unscoped example of those commands below is shorthand: before running it, append the flag. For example, run `openspec status --change "<name>" --json --store "<id>"`, not the unscoped form shown below. Other commands do not take the flag. Hints printed by commands already carry the flag; keep it on follow-ups. Without a store, commands act on the nearest local `openspec/` root.
**Input**: Optionally specify a change name (e.g., `/openspec-apply-change add-auth`). If omitted, check if it can be inferred from conversation context. If vague or ambiguous you MUST prompt for available changes.
**Steps**
1. **Select the change**
If a name is provided, use it. Otherwise:
- Infer from conversation context if the user mentioned a change
- Auto-select if only one active change exists
- If ambiguous, run `openspec list --json` to get available changes and ask the user to select one
Always announce: "Using change: <name>" and how to override (e.g., `/openspec-apply-change <other>`).
2. **Check status to understand the schema**
```bash
openspec status --change "<name>" --json
```
Parse the JSON to understand:
- `schemaName`: The workflow being used (e.g., "spec-driven")
- `planningHome`, `changeRoot`, and `actionContext`: planning scope and edit constraints
- Which artifact contains the tasks (typically "tasks" for spec-driven, check status for others)
3. **Get apply instructions**
```bash
openspec instructions apply --change "<name>" --json
```
This returns:
- `contextFiles`: artifact ID -> array of concrete file paths (varies by schema - could be proposal/specs/design/tasks or spec/tests/implementation/docs)
- Progress (total, complete, remaining)
- Task list with status
- Dynamic instruction based on current state
- Optional `context`: current required project instruction input from the selected root
- Optional `operationGuidance`: current advisory guidance for apply
**Handle states:**
- If `state: "blocked"` (missing artifacts): show message, suggest using `/openspec-continue-change` (if it is not installed, run `openspec status --change "<name>" --json` to see the next artifact and `openspec instructions <artifact-id> --change "<name>" --json` for how to create it)
- If `state: "all_done"`: congratulate, suggest archive
- Otherwise: proceed to implementation
Treat `context` as a required prompt-level input. Read and consider it, and
apply relevant project facts, conventions, and constraints while implementing.
Treat `operationGuidance` as optional additive advice. Read and consider every
entry, and follow entries that are applicable and compatible with the built-in
workflow.
Keep both fields separate from CLI-returned state, missing artifacts, tasks,
progress, `contextFiles`, and the built-in `instruction`. They are not
evidence of task completion, do not replace the built-in instruction, and do
not permit bypassing a blocked state. If context conflicts with the built-in
instruction, an explicit user choice, or a CLI-controlled value, report the
conflict and preserve the controlling value. If guidance is inapplicable or
conflicts with those controlling inputs, do not follow it and explain why.
These are prompt-level behavior contracts, not enforceable checks.
4. **Read context files**
Read every file path listed under `contextFiles` from the apply instructions output.
The files depend on the schema being used:
- **spec-driven**: proposal, specs, design, tasks
- Other schemas: follow the contextFiles from CLI output
Do not copy `context` or `operationGuidance` verbatim into implementation
files or planning artifacts unless the user separately asks for that content.
5. **Show current progress**
Display:
- Schema being used
- Progress: "N/M tasks complete"
- Remaining tasks overview
- Dynamic instruction from CLI
6. **Implement tasks (loop until done or blocked)**
For each pending task:
- Show which task is being worked on
- Make the code changes required
- Keep changes minimal and focused
- Mark task complete in the tasks file: `- [ ]` → `- [x]`
- Continue to next task
**Pause if:**
- Task is unclear → ask for clarification
- Implementation reveals a design issue → suggest updating artifacts
- A task needs work beyond what the spec and tasks describe, or you are tempted to drop, narrow, defer, or accept exceptions to specified behavior to make it fit → surface the added scope and ask; do not absorb it silently
- Error or blocker encountered → report and wait for guidance
- User interrupts
7. **On completion or pause, show status**
Display:
- Tasks completed this session
- Overall progress: "N/M tasks complete"
- If all done: suggest archive
- If paused: explain why and wait for guidance
**Output During Implementation**
```
## Implementing: <change-name> (schema: <schema-name>)
Working on task 3/7: <task description>
[...implementation happening...]
✓ Task complete
Working on task 4/7: <task description>
[...implementation happening...]
✓ Task complete
```
**Output On Completion**
```
## Implementation Complete
**Change:** <change-name>
**Schema:** <schema-name>
**Progress:** 7/7 tasks complete ✓
### Completed This Session
- [x] Task 1
- [x] Task 2
...
All tasks complete! You can archive this change with `/openspec-archive-change`.
```
**Output On Pause (Issue Encountered)**
```
## Implementation Paused
**Change:** <change-name>
**Schema:** <schema-name>
**Progress:** 4/7 tasks complete
### Issue Encountered
<description of the issue>
**Options:**
1. <option 1>
2. <option 2>
3. Other approach
What would you like to do?
```
**Guardrails**
- Keep going through tasks until done or blocked
- Always read context files before starting (from the apply instructions output)
- If task is ambiguous, pause and ask before implementing
- If implementation reveals issues, pause and suggest artifact updates
- Keep code changes minimal and scoped to each task
- Update task checkbox immediately after completing each task
- Pause on errors, blockers, or unclear requirements - don't guess
- When a task needs work beyond what the spec describes, surface the added scope and pause - never silently narrow, defer, or simplify away specified behavior
- Only mark a task `- [x]` when its specified behavior is fully implemented, not when it is partially done or deferred
- Use contextFiles from CLI output, don't assume specific file names
- Do not use context or operation guidance as proof that a task is complete
- Apply relevant project context; report conflicts with controlling workflow inputs
- Consider every guidance entry; explain any inapplicable or conflicting advice
- Do not copy runtime context or operation guidance into implementation files or planning artifacts
- Preserve CLI-controlled blocked/ready/all-done behavior and completion criteria
**Fluid Workflow Integration**
This skill supports the "actions on a change" model:
- **Can be invoked anytime**: Before all artifacts are done (if tasks exist), after partial implementation, interleaved with other actions
- **Allows artifact updates**: If implementation reveals design issues, suggest updating artifacts - not phase-locked, work fluidly
@@ -0,0 +1,182 @@
---
name: openspec-archive-change
description: Archive a completed change in the experimental workflow. Use when the user wants to finalize and archive a change after implementation is complete.
allowed-tools: Bash(openspec:*)
license: MIT
compatibility: Requires openspec CLI.
metadata:
author: openspec
version: "1.0"
generatedBy: "1.12.0"
---
Archive a completed change in the experimental workflow.
**Store selection:** If the user names a store (a store is a standalone OpenSpec repo registered on this machine) or the work lives in one, run `openspec store list --json` to discover registered store ids, then pass `--store <id>` on the commands that read or write specs and changes (`new change`, `status`, `instructions`, `list`, `show`, `validate`, `archive`, `doctor`, `context`, `schemas`, `view`). Once selected, treat `--store <id>` as sticky for the rest of the workflow. Every unscoped example of those commands below is shorthand: before running it, append the flag. For example, run `openspec status --change "<name>" --json --store "<id>"`, not the unscoped form shown below. Other commands do not take the flag. Hints printed by commands already carry the flag; keep it on follow-ups. Without a store, commands act on the nearest local `openspec/` root.
`<capability-path>` is the spec directory relative to `specs/` (for example, `user-auth` or `identity/user-auth`). Preserve the full path from each delta spec when resolving its main spec.
**Input**: Optionally specify a change name. If omitted, check if it can be inferred from conversation context. If vague or ambiguous you MUST prompt for available changes.
**Steps**
1. **Select the change**
If a name is provided, use it. Otherwise:
- Infer from conversation context if the user mentioned a change
- Auto-select if only one active change exists
- If ambiguous, run `openspec list --json` to get available changes and ask the user to select one
When prompting, show only active changes (not already archived).
Include the schema used for each change if available.
Always announce: "Using change: <name>" and how to override (e.g., `/openspec-archive-change <other>`).
**Load current archive inputs before the existing archive checks:**
After resolving the selected change and planning root, run:
```bash
openspec instructions archive --change "<name>" --json
```
Keep the same selected-root flags on this command. This lookup is advisory and
optional: it only supplies extra prompt inputs, so it must never block archiving.
If it exits non-zero or returns invalid JSON — for example on an older CLI that
does not support this command yet — continue the archive workflow with no
context and no operation guidance. Do not report an error and do not stop.
A successful response may omit both optional fields. Treat `context` as a
required prompt-level input: read and consider it, and apply relevant project
facts, conventions, and constraints. Treat `operationGuidance` as optional
additive advice: read and consider every entry, and follow entries that are
applicable and compatible with the built-in archive workflow.
Keep both fields separate from built-in steps, explicit user choices, resolved
paths, CLI checks, and command contracts. If context conflicts with one of those
controlling inputs, report the conflict and preserve the controlling value. If
guidance is inapplicable or conflicts with a controlling input, do not follow it
and explain why. Do not infer replacement paths, skipped prompts, or flags from
either field, and do not copy their text verbatim into specs, change artifacts,
or archive summaries unless the user separately asks for it. These are
prompt-level behavior contracts, not enforceable checks.
2. **Check artifact completion status**
Run `openspec status --change "<name>" --json` to check artifact completion.
Parse the JSON to understand:
- `schemaName`: The workflow being used
- `planningHome`, `changeRoot`, `artifactPaths`, and `actionContext`: path and scope context
- `artifacts`: List of artifacts with their status (`done`, `skipped`, or other)
**If any artifacts are neither `done` nor `skipped`** (skipped artifacts satisfy the requirement - the change declares skip_specs):
- Display warning listing incomplete artifacts
- Ask the user to confirm they want to proceed
- Proceed if user confirms
3. **Check task completion status**
Read the tasks file (typically `tasks.md`) to check for incomplete tasks.
Count tasks marked with `- [ ]` (incomplete) vs `- [x]` (complete).
**If incomplete tasks found:**
- Display warning showing count of incomplete tasks
- Ask the user to confirm they want to proceed
- Proceed if user confirms
**If no tasks file exists:** Proceed without task-related warning.
4. **Assess delta spec sync state**
Use `artifactPaths.specs.existingOutputPaths` from status JSON as the only
delta-spec source. If the `specs` entry is missing or
`existingOutputPaths` is empty, proceed without a sync prompt and do not infer
delta specs from other artifacts.
**If delta specs exist:**
- Compare each delta spec with its corresponding main spec at `<planningHome.root>/openspec/specs/<capability-path>/spec.md` (use the store-aware `planningHome.root` from step 2, not a hardcoded repo path)
- Determine what changes would be applied (adds, modifications, removals, renames)
- Show a combined summary before prompting
**Prompt options:**
- If changes needed: "Sync now (recommended)", "Archive without syncing"
- If already synced: "Archive now", "Sync anyway", "Cancel"
Route on the answer:
- "Cancel" — stop, do not archive
- "Archive without syncing" or "Archive now" — proceed to archive
- "Sync now" or "Sync anyway" — sync, then verify (below)
- Anything else — ask again rather than archiving
Before a selected sync writes any main spec, run
`openspec instructions specs --change "<name>" --json` once with the same
selected-root flags. Require a zero exit status and valid artifact-instruction
JSON. If the lookup fails or returns invalid JSON, report the error and stop
before writing any main spec or moving the change. A valid response with omitted
`rules` is the no-rules case. Apply returned `rules` only to the content and
form of main specs produced by this merge; do not use them as archive guidance,
change CLI behavior, or copy the rule text into any output file.
Then run the `openspec-sync-specs` workflow inline (agent-driven intelligent merge) for change '<name>', passing the delta spec analysis and the fetched specs-rule snapshot from above, and wait for it to finish. The inline sync must reuse that snapshot without fetching `specs` instructions again. Do not delegate it to a background task — step 5 would move `changeRoot` out from under a sync that is still reading it, leaving the change archived and the main specs never updated. If your agent can only run it by delegation, delegate synchronously and wait for the result.
Then re-run the comparison from the top of this step against every capability that has a delta spec in `artifactPaths.specs.existingOutputPaths` — not only the ones the sync reports it touched. A successful sync leaves nothing left to apply, so each capability must now read as already synced:
- ADDED requirements present
- MODIFIED requirements carrying the scenario and description changes named in the delta, with their other scenarios intact
- REMOVED requirements gone — and where this sync retired a capability (removed its last requirement, leaving `## Requirements` empty), its main spec deleted rather than left empty; a spec the sync deliberately kept and reported is also a match
- RENAMED requirements present under the new name and absent under the old one
If the sync failed, or any capability does not match, report what differs and stop — do not archive. Nothing has moved and `changeRoot` is intact, so the user can fix the mismatch or re-run the sync and start the archive again.
5. **Perform the archive**
Create an `archive` directory under `planningHome.changesDir` if it doesn't exist:
```bash
mkdir -p "<planningHome.changesDir>/archive"
```
Generate the target name: use the change name as-is when it already starts with a `YYYY-MM-DD-` prefix; otherwise prepend the current date as `YYYY-MM-DD-<change-name>`. Never stack a second date (same rule as `openspec archive`).
**Check if target already exists:**
- If yes: Fail with error, suggest renaming existing archive or using different date
- If no: Move `changeRoot` to the archive directory
```bash
mv "<changeRoot>" "<planningHome.changesDir>/archive/<target-name>"
```
6. **Display summary**
Show archive completion summary including:
- Change name
- Schema that was used
- Archive location
- Whether specs were synced (if applicable)
- Note about any warnings (incomplete artifacts/tasks)
**Output On Success**
```markdown
## Archive Complete
**Change:** <change-name>
**Schema:** <schema-name>
**Archived to:** the archive path derived from `planningHome.changesDir`/<target-name>/
**Specs:** <"✓ Synced to main specs" only if the step 4 verification passed; otherwise "No delta specs" or "Sync skipped">
<"All artifacts complete. All tasks complete." — or, if archived with warnings, list them instead (e.g. "Archived with 2 incomplete tasks")>
```
**Guardrails**
- Announce the selected change; prompt for selection when it is ambiguous
- Use artifact graph (openspec status --json) for completion checking
- Don't block archive on warnings - just inform and confirm
- Preserve .openspec.yaml when moving to archive (it moves with the directory)
- Show clear summary of what happened
- If sync is requested, run the `openspec-sync-specs` workflow inline (agent-driven)
- Never archive while a spec sync is still in flight — run the sync inline and verify the main specs before moving `changeRoot`
- If delta specs exist, always run the sync assessment and show the combined summary before prompting
- Apply relevant runtime context and report conflicts; operation guidance remains advisory
- Consider every guidance entry and explain any inapplicable or conflicting advice
- Existing CLI checks, resolved paths, prompts, and command contracts are unchanged
- Artifact rules constrain only the specs being written and are never operation guidance
- Never copy runtime context, operation guidance, or artifact-rule text verbatim into output files
+335
View File
@@ -0,0 +1,335 @@
---
name: openspec-explore
description: Enter explore mode - a thinking partner for exploring ideas, investigating problems, and clarifying requirements. Use when the user wants to think through something before or during a change.
allowed-tools: Bash(openspec:*)
license: MIT
compatibility: Requires openspec CLI.
metadata:
author: openspec
version: "1.0"
generatedBy: "1.12.0"
---
Enter explore mode. Think deeply. Visualize freely. Follow the conversation wherever it goes.
**IMPORTANT: Explore mode is for thinking, not implementing.** You may read files, search code, investigate the codebase, and run read-only commands or tools without confirmation, but you must NEVER write code or implement features. If the user asks you to implement something, remind them to exit explore mode first and create a change proposal. You MAY create or update OpenSpec change artifacts (proposals, designs, specs) within a confirmed scope—that's capturing thinking, not implementing. Answering design or clarifying questions is never consent to write. Before the first write-capable action, name the artifacts or files you would change and what you would do, ask a direct yes/no question, and wait for the user's confirmation in a separate message. Confirmation covers only the scope you described; ask again before expanding it. For a new change, scaffold it first as described below.
**This is a stance, not a workflow.** There are no fixed steps, no required sequence, no mandatory outputs. You're a thinking partner helping the user explore.
**Store selection:** If the user names a store (a store is a standalone OpenSpec repo registered on this machine) or the work lives in one, run `openspec store list --json` to discover registered store ids, then pass `--store <id>` on the commands that read or write specs and changes (`new change`, `status`, `instructions`, `list`, `show`, `validate`, `archive`, `doctor`, `context`, `schemas`, `view`). Once selected, treat `--store <id>` as sticky for the rest of the workflow. Every unscoped example of those commands below is shorthand: before running it, append the flag. For example, run `openspec status --change "<name>" --json --store "<id>"`, not the unscoped form shown below. Other commands do not take the flag. Hints printed by commands already carry the flag; keep it on follow-ups. Without a store, commands act on the nearest local `openspec/` root.
---
## The Stance
- **Curious, not prescriptive** - Ask questions that emerge naturally, don't follow a script
- **Open threads, not interrogations** - Surface multiple interesting directions and let the user follow what resonates. Don't funnel them through a single path of questions.
- **Visual** - Use ASCII diagrams liberally when they'd help clarify thinking
- **Adaptive** - Follow interesting threads, pivot when new information emerges
- **Patient** - Don't rush to conclusions, let the shape of the problem emerge
- **Grounded** - Explore the actual codebase when relevant, don't just theorize
---
## Planning a Change
When the user is planning a change, guide them toward shared understanding with focused discovery questions. For open-ended discussion, follow the conversation without imposing an interview or a required output.
Before asking a factual question, follow the context discovery below and inspect relevant OpenSpec artifacts, source, tests, docs, and configuration. Do not ask the user to repeat facts you can verify. Summarize relevant findings without reproducing private context or rules. If evidence is missing, conflicting, or inaccessible, state that limitation and ask only for the clarification needed to proceed.
- **Follow dependencies** - Resolve the next blocking decision before its dependent details. For example, clarify the user's outcome and scope before choosing an API or data model. Revisit downstream assumptions when an earlier answer changes. Skip branches that do not matter to this goal.
- **Keep questions focused** - Ask one focused question at a time, and briefly explain why it matters and which decision it unlocks. Batch questions only if the user asks for a batch; keep them small and group related decisions.
- **Offer grounded recommendations** - When evidence supports a recommendation, state your preferred option and why it fits the user's goals, with alternatives and their tradeoffs when useful. Do not invent intent, priorities, or external constraints: ask the user when only they can answer. Avoid a fixed question format.
- **Keep a conversational record** - Track decisions in the conversation, not in files. Separate confirmed decisions from proposed defaults and unresolved questions. Silence is not acceptance. Accepting an answer or a batch of recommendations is not permission to write. Keep file-write confirmation separate from discovery questions and follow the guardrails below.
Stop asking when the user has enough clarity. Let them pause, pivot, or defer a decision; do not exhaust every branch or force a proposal.
For example, after inspecting the relevant code:
```text
The CLI already uses SQLite and has no remote service. Is sharing state
across devices in scope? That determines whether local storage is enough.
If this stays a single-device tool, I recommend keeping SQLite to avoid
adding a service to operate; shared state would need a separate sync design.
```
---
## What You Might Do
Depending on what the user brings, you might:
**Explore the problem space**
- Ask clarifying questions that emerge from what they said
- Challenge assumptions
- Reframe the problem
- Find analogies
**Investigate the codebase**
- Map existing architecture relevant to the discussion
- Find integration points
- Identify patterns already in use
- Surface hidden complexity
**Compare options**
- Brainstorm multiple approaches
- Build comparison tables
- Sketch tradeoffs
- Recommend a path (if asked)
**Visualize**
```
+------------------------------------------+
| Use ASCII diagrams liberally |
+------------------------------------------+
| |
| [State A] -------> [State B] |
| | |
| v |
| [State C] |
| |
| System diagrams, state machines, |
| data flows, architecture sketches, |
| dependency graphs, comparison tables |
| |
+------------------------------------------+
```
**Draw with plain ASCII only** — borders `+` `-` `|`, arrows `-->` `<--` `^` `v`, markers `*` `x`.
Unicode diagram glyphs can render at different widths across terminals, fonts, and locales, so padded boxes and aligned tables can drift. Keep every diagram character ASCII.
**Surface risks and unknowns**
- Identify what could go wrong
- Find gaps in understanding
- Suggest spikes or investigations
---
## OpenSpec Awareness
You have full context of the OpenSpec system. Use it naturally, don't force it.
### Check for context
At the start, quickly check what exists:
```bash
openspec list --json
```
This tells you:
- If there are active changes
- Their names, schemas, and status
- What the user might be working on
Then read the project's own context from the resolved root - `<root.path>/openspec/config.yaml` (or `config.yml`). Use the `root.path` returned above, and skip this if neither file exists:
- `context`: project background - tech stack, conventions, constraints
- `rules`: keyed by artifact id - the entries for an artifact apply only when you write that artifact
Ground your thinking in these. They are constraints for you to follow, not content to reproduce: do NOT copy them into the conversation or into any artifact you create.
### When no change exists
Think freely. When insights crystallize, you might offer:
- "This feels solid enough to start a change. Want me to create a proposal?"
- Or keep exploring - no pressure to formalize
If the user asks you to capture the exploration as a new change, transition seamlessly into the requested capture:
1. Run `openspec new change "<name>"` (with `--store <id>` when applicable) before creating any artifacts. Never create a new change directory under `openspec/changes/` by hand; the CLI scaffold creates required metadata such as `.openspec.yaml`. Keep the selected `--store <id>` on every applicable follow-up `status` and `instructions` command.
2. Run `openspec status --change "<name>" --json` (append the confirmed `--store "<id>"` only for a registered standalone store), then process the requested artifacts in dependency order. For each requested artifact that is `ready`, run `openspec instructions "<artifact-id>" --change "<name>" --json` (append the confirmed `--store "<id>"` only for a registered standalone store). Before creating a requested artifact, evaluate any condition in its own `instruction` against the explored change; record a deliberate skip instead when the condition does not apply. If a requested artifact is blocked by a direct prerequisite the user did not request, run `openspec instructions "<prerequisite-id>" --change "<name>" --json` (append the confirmed `--store "<id>"` only for a registered standalone store) for that prerequisite whether it is `ready` or `blocked`. If its own `instruction` states a condition, evaluate that condition against the explored change and record a deliberate skip only when the condition does not apply. If the condition applies, or the prerequisite is not conditional, treat it as a normal prerequisite and ask before expanding the capture. Do not create an unrequested prerequisite unless the user approves.
3. Follow the returned `template` and `instruction` fields. Read completed dependency files listed in `dependencies`, and apply `context` and `rules` as constraints without copying them into the artifact. If the instruction delegates creation to a specific skill or command, invoke it; otherwise write the artifact to `resolvedOutputPath`, using the instruction to choose a concrete path when it is a glob. Verify that the selected concrete output exists.
4. After creating each artifact, re-run `openspec status --change "<name>" --json` (append the confirmed `--store "<id>"` only for a registered standalone store) and continue until every requested artifact is `done`, `skipped`, or was deliberately skipped because its own `instruction` stated a condition that did not apply. Tell the user about a deliberate conditional skip, remember it, and do not reconsider it. Dependencies are enablers, not gates: if a requested artifact is still `blocked` only because you deliberately skipped a conditional prerequisite, run `openspec instructions "<artifact-id>" --change "<name>" --json` (append the confirmed `--store "<id>"` only for a registered standalone store) despite the blocked status, then create it using step 3 only when those recorded conditional skips are its sole missing dependencies. If a requested artifact is blocked by a prerequisite the user did not ask to capture and cannot be conditionally skipped, explain that dependency and ask before expanding the capture.
Capture the artifact(s) the user requested without asking them to invoke another workflow command. If they asked only to start a change, stop after scaffolding and show its status.
### When a change exists
If the user mentions a change or you detect one is relevant:
1. **Resolve and read existing artifacts for context**
- Run `openspec status --change "<name>" --json`.
- Use `changeRoot`, `artifactPaths`, and `actionContext` from the status JSON.
- Read existing files from `artifactPaths.<artifact>.existingOutputPaths`.
2. **Reference them naturally in conversation**
- "Your design mentions using Redis, but we just realized SQLite fits better..."
- "The proposal scopes this to premium users, but we're now thinking everyone..."
3. **Offer to capture when decisions are made**
`<capability-path>` is the spec directory relative to `specs/` (for example, `user-auth` or `identity/user-auth`). Preserve an existing capability's full path and follow the project's established organization for new capabilities.
| Insight Type | Where to Capture |
|----------------------------|-------------------------------------|
| New requirement discovered | `specs/<capability-path>/spec.md` |
| Requirement changed | `specs/<capability-path>/spec.md` |
| Design decision made | `design.md` |
| Scope changed | `proposal.md` |
| New work identified | `tasks.md` |
| Assumption invalidated | Relevant artifact |
Example offers:
- "That's a design decision. Capture it in design.md?"
- "This is a new requirement. Add it to specs?"
- "This changes scope. Update the proposal?"
4. **The user decides** - Offer and move on. Don't pressure. Don't auto-capture.
---
## What You Don't Have To Do
- Follow a script
- Ask the same questions every time
- Produce a specific artifact
- Reach a conclusion
- Stay on topic if a tangent is valuable
- Be brief (this is thinking time)
---
## Handling Different Entry Points
**User brings a vague idea:**
```
User: I'm thinking about adding real-time collaboration
You: Real-time collab is a big space. Let me think about this...
COLLABORATION SPECTRUM
============================================
Awareness Coordination Sync
| | |
v v v
+--------+ +--------+ +--------+
|Presence| |Cursors | | CRDT |
| "3 | | Multi | |Conflict|
|online" | | select | | free |
+--------+ +--------+ +--------+
| | |
trivial moderate complex
Where's your head at?
```
**User brings a specific problem:**
```
User: The auth system is a mess
You: [reads codebase]
+---------------------------------------------+
| CURRENT AUTH FLOW |
+---------------------------------------------+
|
+-------------+-------------+
v v v
+---------+ +---------+ +---------+
| Google | | GitHub | | Email |
| OAuth | | OAuth | | Magic |
+----+----+ +----+----+ +----+----+
| | |
+-------------+-------------+
v
+-----------+
| Session |
+-----+-----+
|
v
+-----------+
| Perms |
+-----------+
I see three tangles. Which one's burning?
```
**User is stuck mid-implementation:**
```
User: /openspec-explore add-auth-system
The OAuth integration is more complex than expected
You: [reads change artifacts]
You're on task 4: "Implement OAuth flow"
Let me trace what's involved...
[draws diagram, explores options, suggests paths]
Want to update the design to reflect this?
Or add a spike task to investigate?
```
**User wants to compare options:**
```
User: Should we use Postgres or SQLite?
You: Generic answer is boring. What's the context?
User: A CLI tool that tracks local dev environments
You: That changes everything.
+-------------------------------------------------+
| CLI TOOL DATA STORAGE |
+-------------------------------------------------+
Key constraints:
- No daemon running
- Must work offline
- Single user
SQLite Postgres
Deployment embedded needs server
Offline yes no
Single file yes no
SQLite. Not even close.
Unless... is there a sync component?
```
---
## Ending Discovery
There's no required ending. Discovery might:
- **Flow into a proposal**: "Ready to start? I can create a change proposal."
- **Result in artifact updates**: "Updated design.md with these decisions"
- **Just provide clarity**: User has what they need, moves on
- **Continue later**: "We can pick this up anytime"
When it feels like things are crystallizing, you might summarize:
```
## What We Figured Out
**The problem**: [crystallized understanding]
**The approach**: [if one emerged]
**Open questions**: [if any remain]
**Next steps** (if ready):
- Create a change proposal
- Keep exploring: just keep talking
```
But this summary is optional. Sometimes the thinking IS the value.
---
## Guardrails
- **Don't implement** - Never write code or implement features. Workflow configuration counts too: creating or editing schemas, templates, or `openspec/config.yaml` is a change, not thinking. Creating or updating OpenSpec change artifacts within the confirmed scope is fine, writing anything else is not.
- **Don't fake understanding** - If something is unclear, dig deeper
- **Don't rush** - Discovery is thinking time, not task time
- **Don't force structure** - Let patterns emerge naturally
- **Don't auto-capture** - Offer to save insights, don't just do it. Read-only commands and tools need no confirmation. Before the first write-capable action—including `openspec new change` or another command that writes files—name the artifacts or files and proposed changes, ask a direct yes/no question, and wait for explicit confirmation in a separate user message. That confirmation covers only the described scope; ask again before expanding it. Answers to design or clarifying questions are never consent to write.
- **Don't manually scaffold changes** - Never create a new change directory under `openspec/changes/` by hand. Always use `openspec new change "<name>"` (with `--store <id>` when applicable) so required metadata such as `.openspec.yaml` is created before writing artifacts.
- **Do visualize** - A good diagram is worth many paragraphs
- **Do explore the codebase** - Ground discussions in reality
- **Do question assumptions** - Including the user's and your own
+153
View File
@@ -0,0 +1,153 @@
---
name: openspec-propose
description: Propose a new change with all artifacts generated in one step. Use when the user wants to quickly describe what they want to build and get a complete proposal with design, specs, and tasks ready for implementation.
allowed-tools: Bash(openspec:*)
license: MIT
compatibility: Requires openspec CLI.
metadata:
author: openspec
version: "1.0"
generatedBy: "1.12.0"
---
Propose a new change - create the change and generate all artifacts in one step.
**Planning boundary**: This workflow creates planning artifacts only. The user request that selected or triggered this workflow authorizes planning only, even if it asks to build or fix something. Do not edit project code. After the planning artifacts are complete, stop. Do not start implementation in the same response, even if the initial request asks for it. Wait for a new user request after the artifacts are presented; then start the apply workflow.
I'll create a change with the artifacts your schema defines. With the default spec-driven schema that is:
- proposal.md (what & why)
- `specs/<capability-path>/spec.md` (what the system must do - a delta, not the main spec)
- design.md (how)
- tasks.md (implementation steps)
`<capability-path>` is the spec directory relative to `specs/` (for example, `user-auth` or `identity/user-auth`). Preserve an existing capability's full path and follow the project's established organization for new capabilities.
When the user is ready to implement, they must start the apply workflow explicitly.
---
**Store selection:** If the user names a store (a store is a standalone OpenSpec repo registered on this machine) or the work lives in one, run `openspec store list --json` to discover registered store ids, then pass `--store <id>` on the commands that read or write specs and changes (`new change`, `status`, `instructions`, `list`, `show`, `validate`, `archive`, `doctor`, `context`, `schemas`, `view`). Once selected, treat `--store <id>` as sticky for the rest of the workflow. Every unscoped example of those commands below is shorthand: before running it, append the flag. For example, run `openspec status --change "<name>" --json --store "<id>"`, not the unscoped form shown below. Other commands do not take the flag. Hints printed by commands already carry the flag; keep it on follow-ups. Without a store, commands act on the nearest local `openspec/` root.
**Input**: The user's request should include a change name (kebab-case) OR a description of what they want to build.
**Steps**
1. **Understand the request and clarify material ambiguity**
If no clear input is provided, ask the user (open-ended, no preset options):
> "What change do you want to work on? Describe what you want to build or fix."
From their description, derive a kebab-case name (e.g., "add user authentication" → `add-user-auth`).
**IMPORTANT**: Do NOT proceed without understanding what the user wants to build.
If the request contains ambiguity that would materially affect scope, externally observable behavior, compatibility, or acceptance criteria, ask the user before creating the change. For minor details, make a reasonable assumption and record it in the planning artifacts.
2. **Determine the workflow schema**
Use the configured default schema unless the user explicitly requests a different workflow.
**Use a different schema only if the user:**
- Explicitly requests a specific schema by name → use `--schema <schema-name>`
- Asks to "show workflows" or asks "what workflows" exist → resolve the authoritative root by running `openspec context --json` from the current working directory. If the user explicitly selected a registered store, use `openspec context --json --store "<store-id>"`. Then run `openspec schemas --json` with its working directory set to the returned `root.path` and let them choose. This preserves roots selected by a local `store:` pointer or the global `defaultStore`; when a registered store was explicitly selected, append `--store "<store-id>"` to `openspec schemas --json` as well. If context reports only `no_openspec_root`, run `openspec schemas --json` from the current working directory instead. Do not use this fallback for invalid or unavailable stores.
Otherwise, omit `--schema` to preserve the configured default.
3. **Create the change directory**
Choose one schema form below. If a registered store is selected, append `--store "<store-id>"` to that command and each later OpenSpec command shown below that accepts `--store`.
Using the configured default:
```bash
openspec new change "<name>"
```
Using an explicitly requested schema:
```bash
openspec new change "<name>" --schema "<schema-name>"
```
This creates a scaffolded change in the planning home resolved by the CLI with `.openspec.yaml`.
4. **Get the artifact build order**
```bash
openspec status --change "<name>" --json
```
Parse the JSON to get:
- `applyRequires`: array of artifact IDs needed before implementation (e.g., `["tasks"]`)
- `artifacts`: list of all artifacts, each with its `status` and its `requires` edges (the artifact IDs it directly depends on)
- `planningHome`, `changeRoot`, `artifactPaths`, and `actionContext`: path and scope context. Use these instead of assuming repo-local paths.
5. **Create every artifact in the required set**
Use a todo list to track progress through the artifacts.
Loop through artifacts in dependency order (artifacts with no pending dependencies first):
a. **For each artifact that is `ready` (dependencies satisfied)**:
- Get instructions:
```bash
openspec instructions <artifact-id> --change "<name>" --json
```
- The instructions JSON includes:
- `context`: Project background (constraints for you - do NOT include in output)
- `rules`: Artifact-specific rules (constraints for you - do NOT include in output)
- `template`: The structure to use for your output file
- `instruction`: Schema-specific guidance for this artifact type
- `skipped`/`warning`: present when the change declares skip_specs and this artifact must NOT be created - stop and pick another artifact
- `resolvedOutputPath`: Resolved path or pattern to write the artifact
- `dependencies`: Completed artifacts to read for context
- Read any completed dependency files for context - always re-read them from disk, even if you saw them earlier in the conversation (the user may have edited them)
- **Inspect the relevant project before drafting**: Read `context` and `rules` first, then inspect relevant implementation, nearby tests, configuration, and documentation outside `openspec/`. Keep inspection read-only and proportional to the change; reuse findings for later artifacts and inspect more only as needed.
- Identify the target project from the request and project context; the planning home may be separate from the code. If the target is unclear, ask. For greenfield or non-code changes, inspect the available structure and relevant documents. If source is unavailable, state the limitation and ask when it materially affects the plan.
- Ground scope, approach, and tasks in what you find. Distinguish observed behavior from assumptions and proposed additions; surface conflicts with existing specs instead of silently deciding which is correct.
- Do this discovery now, rather than leaving generic "explore the codebase" or "make a plan" tasks for implementation. Keep any necessary follow-up investigation specific to an unresolved question.
- If the `instruction` field delegates creation to a specific skill or command, invoke it to produce the artifact instead of writing the file yourself, then verify the artifact file exists at `resolvedOutputPath`
- Otherwise create the artifact file using `template` as the structure and write it to `resolvedOutputPath`. If `resolvedOutputPath` is a glob, follow `instruction` to choose the concrete file path
- Apply `context` and `rules` as constraints - but do NOT copy them into the file
- Show brief progress: "Created <artifact-id>"
b. **Continue until every artifact in the required set exists (not just `apply.requires`)**
- After creating each artifact, re-run `openspec status --change "<name>" --json`
- The required set is `applyRequires` plus every artifact reachable from those by following the `requires` edges in `status --json` - walk them transitively (spec-driven closes over proposal, specs, design, tasks). Leave artifacts outside that set alone
- `status` is file-existence only, so an `applyRequires` artifact reading `done` does NOT mean its dependencies exist - writing `tasks.md` early marks `tasks` done while `specs` was never written. Use each artifact's `requires` edges, not its `status`, to build the required set: a `done` artifact still lists what it depends on
- An artifact already reading `status: "skipped"` is satisfied: the change declares `skip_specs` in `.openspec.yaml`, so its files must NOT exist. Never try to create one
- Create every artifact in the required set that is missing, then re-check - creating one can unblock others
- Skip one only when `status` already reports it `skipped`, or when its own `instruction` says it is conditional: run `openspec instructions <artifact-id> --change "<name>" --json` and skip only if its `instruction` field marks it optional (e.g. "create only if..."). Spec-driven's `design.md` qualifies; `specs` qualifies only via the `skipped` status above, never by your own judgment. Tell the user, and do not reconsider it
- Dependencies are enablers, not gates: if a required artifact is still `blocked` only because you skipped a conditional dependency, write it anyway
- Stop when every artifact in the required set is `done`, `skipped`, or was deliberately skipped
c. **If an artifact requires user input** (unclear context):
- Ask the user to clarify
- Then continue with creation
6. **Show final status**
```bash
openspec status --change "<name>"
```
**Output**
After completing all artifacts, summarize:
- Change name and location
- List of artifacts created with brief descriptions, plus any conditional artifact you skipped and why
- What's ready: "All artifacts needed for implementation are ready."
- Prompt: "The artifacts are ready for review. When you are ready, run `/openspec-apply-change` or ask me to apply this change."
**Artifact Creation Guidelines**
- Follow the `instruction` field from `openspec instructions` for each artifact type - it is the authoritative guidance, even for familiar artifact names
- If the `instruction` field directs you to use a specific skill or command to create the artifact, invoke it instead of writing the artifact directly
- The schema defines what each artifact should contain - follow it
- Read dependency artifacts for context before creating new ones
- Use `template` as the structure for your output file - fill in its sections
- **IMPORTANT**: `context` and `rules` are constraints for YOU, not content for the file
- Do NOT copy `<context>`, `<rules>`, `<project_context>` blocks into the artifact
- These guide what you write, but should never appear in the output
**Guardrails**
- The request that invoked this workflow authorizes planning only. Any implementation or apply instruction in that request does not carry forward. Do NOT implement the change, start the apply workflow, or edit project code during this workflow. After presenting the artifacts, stop and wait for a new user request to start the apply workflow
- Create every artifact the apply phase transitively depends on, not just the ids listed in `apply.requires`
- Always read dependency artifacts before creating a new one - re-read from disk, not from conversation memory (files may have changed since you last saw them)
- Ask about ambiguities that would materially change scope, externally observable behavior, compatibility, or acceptance criteria; for minor details, make reasonable assumptions and record them
- If a change with that name already exists, ask if user wants to continue it or create a new one
- Verify each artifact file exists after writing before proceeding to next
+262
View File
@@ -0,0 +1,262 @@
---
name: openspec-sync-specs
description: Sync delta specs from a change to main specs. Use when the user wants to update main specs with changes from a delta spec, without archiving the change.
allowed-tools: Bash(openspec:*)
license: MIT
compatibility: Requires openspec CLI.
metadata:
author: openspec
version: "1.0"
generatedBy: "1.12.0"
---
Sync delta specs from a change to main specs.
This is an **agent-driven** operation - you will read delta specs and directly edit main specs to apply the changes. This allows intelligent merging (e.g., adding a scenario without copying the entire requirement).
**Store selection:** If the user names a store (a store is a standalone OpenSpec repo registered on this machine) or the work lives in one, run `openspec store list --json` to discover registered store ids, then pass `--store <id>` on the commands that read or write specs and changes (`new change`, `status`, `instructions`, `list`, `show`, `validate`, `archive`, `doctor`, `context`, `schemas`, `view`). Once selected, treat `--store <id>` as sticky for the rest of the workflow. Every unscoped example of those commands below is shorthand: before running it, append the flag. For example, run `openspec status --change "<name>" --json --store "<id>"`, not the unscoped form shown below. Other commands do not take the flag. Hints printed by commands already carry the flag; keep it on follow-ups. Without a store, commands act on the nearest local `openspec/` root.
`<capability-path>` is the spec directory relative to `specs/` (for example, `user-auth` or `identity/user-auth`). Preserve the full path from each delta spec when resolving its main spec.
**Input**: Optionally specify a change name. If omitted, check if it can be inferred from conversation context. If vague or ambiguous you MUST prompt for available changes.
**Steps**
1. **Select the change**
If a name is provided, use it. Otherwise:
- Infer from conversation context if the user mentioned a change
- Auto-select if only one active change exists
- If ambiguous, run `openspec list --json` to get available changes and ask the user to select one
When prompting, show changes that have delta specs (under `specs/` directory).
Always announce: "Using change: <name>" and how to override (e.g., `/openspec-sync-specs <other>`).
2. **Resolve change context**
Run:
```bash
openspec status --change "<name>" --json
```
The JSON includes `planningHome.root`. Main specs live under `<planningHome.root>/openspec/specs/` — use that (store-aware) root for every main-spec path below, not a hardcoded repo path. When a store is selected it points at the store, not the current repository.
3. **Find delta specs**
Use `artifactPaths.specs.existingOutputPaths` from the status JSON as the
only source of delta spec paths. If the `specs` entry is missing or
`existingOutputPaths` is empty, report that there are no delta specs to sync,
do not infer them from other artifacts, and stop without requesting artifact
instructions or writing a main spec.
Sync every path in `existingOutputPaths` unless the caller narrowed the set.
A caller narrows it by naming an explicit list of complete entries from
`existingOutputPaths` — copy those absolute values verbatim. Archive does
this inline, and a user can too (for example, by selecting the entry ending
in `/specs/billing/invoices/spec.md`).
Then sync only the named paths and leave the remaining delta specs untouched:
bulk archive excludes a delta whose implementation it could not find, and
syncing it anyway would write a main spec the caller deliberately withheld.
Carry that narrowed selection through step 4; never widen it back to the full
list. If a named path is not in `existingOutputPaths`, do not sync it —
report it and stop, rather than dropping it silently. If the named list is
empty, report that there is nothing to sync and stop without writing a main
spec.
Each delta spec file contains sections like:
- `## ADDED Requirements` - New requirements to add
- `## MODIFIED Requirements` - Changes to existing requirements
- `## REMOVED Requirements` - Requirements to remove
- `## RENAMED Requirements` - Requirements to rename (FROM:/TO: format)
If no delta specs found, inform user and stop.
4. **For each delta spec, apply changes to main specs**
Before the first main-spec write, obtain one current specs-rule snapshot:
- If archive invoked this workflow inline and supplied a valid snapshot from
`openspec instructions specs --change "<name>" --json`, reuse it and do not
fetch the same instructions again.
- Otherwise run that command once now with the same selected-root flags.
- If the direct lookup exits non-zero or returns invalid artifact-instruction
JSON, report the error and stop before writing any main spec. Do not treat the
failure as an absent rule set.
- A valid response with omitted `rules` means no artifact rules are configured
and the existing semantic merge continues.
Apply returned `rules` only to the content and form of the main specs produced
by this merge. Artifact rules are not operation guidance and cannot change
selected roots, delta paths, CLI checks, or workflow steps. Use their text as
constraints without copying it verbatim into a main spec or summary.
For each capability delta spec path selected in step 3 — the full `existingOutputPaths` list, or the narrowed subset when a caller supplied one (these may belong to a selected store, not the repo):
a. **Read the delta spec** to understand the intended changes
b. **Read the main spec** at `<planningHome.root>/openspec/specs/<capability-path>/spec.md` (may not exist yet)
c. **Apply changes intelligently**:
**ADDED Requirements:**
- If requirement doesn't exist in main spec → add it
- If requirement already exists → update it to match (treat as implicit MODIFIED)
**MODIFIED Requirements:**
- Find the requirement in main spec
- Apply the changes - this can be:
- Adding new scenarios the main spec does not have yet
- Modifying existing scenarios
- Changing the requirement description
- Preserve scenarios/content not mentioned in the delta
**REMOVED Requirements:**
- Remove the entire requirement block from main spec
- Retiring the capability. Delete the whole `spec.md` - and the directory once
nothing else is left in it - only when ALL of these hold:
1. removing the requirements *this run* left no requirement blocks;
2. the rest of the spec is well-formed (it still has a `## Purpose`);
3. the main spec was not already empty before this sync - if you removed
nothing, change nothing;
4. every other nonblank line in the whole file is accounted for as the
title, Purpose, Requirements header, or a canonical requirement's
statement, scenarios, or fenced examples;
5. the change's `.openspec.yaml` declares `retire_capabilities: true`;
6. the `spec.md` resolves inside the real specs root (do not follow a
capability-directory symlink to delete an external file).
If removing the selected requirements would leave no requirement blocks and
any retirement condition is not satisfied, do not modify the main spec. Stop
the sync for that capability, report the blocking condition, and tell the user
how to resolve it. Never write or leave an empty `## Requirements` section.
When only the marker is missing, say that too - it is the one thing the user
can add to make the retirement go through.
- Deleting the file also deletes its `## Purpose`; any other section blocks
retirement. Name Purpose when you report the retirement. Include a pasteable
`git checkout` only when the spec lived in the caller's checkout;
otherwise give checkout-scoped recovery guidance.
**RENAMED Requirements:**
- Find the FROM requirement, rename to TO
**`## Purpose` in the delta:**
- The main spec already has one and it is authoritative - leave it alone
(this is what `openspec archive` does; it warns and moves on)
d. **Create new main spec** if capability doesn't exist yet:
- Create `<planningHome.root>/openspec/specs/<capability-path>/spec.md`
- Add Purpose section: copy the delta's `## Purpose` body verbatim when it has one
(this is what `openspec archive` does); only write a brief TBD placeholder when it does not
- Add Requirements section with the ADDED requirements
- Follow the **Main Spec Format Reference** below
5. **Validate updated main specs**
Run `openspec validate --specs` with the same selected-root flags used earlier.
If validation fails, report the problems and do not claim the sync succeeded.
6. **Show summary**
After applying all changes, summarize:
- Which capabilities were updated
- What changes were made (requirements added/modified/removed/renamed)
- Any new main spec left with a TBD Purpose placeholder, so it gets written
now rather than lingering
- Any capability retired, naming the deleted `spec.md`, its Purpose, and
either a pasteable `git checkout` or checkout-scoped recovery guidance
**Delta Spec Format Reference**
```markdown
## Purpose
Only on a delta that introduces a brand-new capability. Seeds the new main spec.
## ADDED Requirements
### Requirement: New Feature
The system SHALL do something new.
#### Scenario: Basic case
- **WHEN** user does X
- **THEN** system does Y
## MODIFIED Requirements
### Requirement: Existing Feature
The system SHALL keep doing the existing thing, now also handling A.
#### Scenario: Scenario the main spec already has
- **WHEN** user does X
- **THEN** system does Y
#### Scenario: New scenario to add
- **WHEN** user does A
- **THEN** system does B
## REMOVED Requirements
### Requirement: Deprecated Feature
## RENAMED Requirements
- FROM: `### Requirement: Old Name`
- TO: `### Requirement: New Name`
```
**Main Spec Format Reference**
Main specs are what the delta merges INTO. They must never contain delta operation headers (`## ADDED/MODIFIED/REMOVED/RENAMED Requirements`) - after syncing, every requirement lives under a single `## Requirements` section:
```markdown
# <capability> Specification
## Purpose
Short description of what this capability does and why it exists.
## Requirements
### Requirement: New Feature
The system SHALL do something new.
#### Scenario: Basic case
- **WHEN** user does X
- **THEN** system does Y
```
**Key Principle: Intelligent Merging**
Unlike programmatic merging, you merge rather than overwrite:
- A MODIFIED block carries the whole requirement - body plus every scenario that survives the change. `openspec validate` and `openspec archive` both reject one that drops a scenario the main spec still has.
- Keep anything the delta does not mention, in the main spec's existing order
- Use your judgment to merge changes sensibly
**Output On Success**
```markdown
## Specs Synced: <change-name>
Updated main specs:
**<capability-1>**:
- Added requirement: "New Feature"
- Modified requirement: "Existing Feature" (added 1 scenario)
**<capability-2>**:
- Created new spec file
- Added requirement: "Another Feature"
Main specs are now updated. The change remains active - archive when implementation is complete.
```
**Guardrails**
- Read both delta and main specs before making changes
- Preserve existing content not mentioned in delta
- Never copy a delta file into a main spec as-is - merge its content so the main spec keeps the Main Spec Format Reference structure, with no delta operation headers
- If something is unclear, ask for clarification
- Show what you're changing as you go
- The operation should be idempotent - running twice should give same result
- Use only `artifactPaths.specs.existingOutputPaths`; never infer delta specs from unrelated artifacts
- Honor a caller-supplied subset of `existingOutputPaths`; never widen it back to the full list
- Fetch specs instructions once for direct sync, or reuse the archive-supplied snapshot inline
- Stop before every main-spec write on a non-zero or invalid JSON specs-instruction response
- Artifact rules constrain only the specs being written and are never copied into output files
@@ -0,0 +1,91 @@
---
name: openspec-update-change
description: Update an OpenSpec change by revising its existing planning artifacts and keeping them coherent with one another. Use when the user wants to revise a change's plan, fold new decisions into it, or reconcile its artifacts after an edit. Never edits code.
allowed-tools: Bash(openspec:*)
license: MIT
compatibility: Requires openspec CLI.
metadata:
author: openspec
version: "1.0"
generatedBy: "1.12.0"
---
Revise a change's existing planning artifacts and keep them coherent. Never edit code.
**Store selection:** If the user names a store (a store is a standalone OpenSpec repo registered on this machine) or the work lives in one, run `openspec store list --json` to discover registered store ids, then pass `--store <id>` on the commands that read or write specs and changes (`new change`, `status`, `instructions`, `list`, `show`, `validate`, `archive`, `doctor`, `context`, `schemas`, `view`). Once selected, treat `--store <id>` as sticky for the rest of the workflow. Every unscoped example of those commands below is shorthand: before running it, append the flag. For example, run `openspec status --change "<name>" --json --store "<id>"`, not the unscoped form shown below. Other commands do not take the flag. Hints printed by commands already carry the flag; keep it on follow-ups. Without a store, commands act on the nearest local `openspec/` root.
**Input**: Optionally specify a change name. If omitted, check if it can be inferred from conversation context. If vague or ambiguous you MUST prompt for available changes.
`/openspec-continue-change` is an optional workflow and may not be installed. Before suggesting it anywhere below, verify that it is available. If it is unavailable, `openspec status --change "<name>" --json` shows the next artifact and `openspec instructions "<artifact-id>" --change "<name>" --json` explains how to create it.
**Steps**
1. **Select the change**
If a name is provided, use it. Otherwise:
- Infer from conversation context if the user mentioned a change
- Auto-select if only one active change exists
- If ambiguous, run `openspec list --json` to get available changes sorted by most recently modified, and ask the user to select one
When prompting, present the top 3-4 most recently modified changes as options, showing:
- Change name
- Schema (from `schema` field if present, otherwise "spec-driven")
- Status (e.g., "0/5 tasks", "complete", "no tasks")
- How recently it was modified (from `lastModified` field)
Mark the most recently modified change as "(Recommended)" since it's likely what the user wants to update.
Always announce: "Using change: <name>" and how to override (e.g., `/openspec-update-change <other>`).
2. **Get the change's artifacts**
```bash
openspec status --change "<name>" --json
```
Parse the JSON to understand current state. The response includes:
- `schemaName`: The workflow schema being used (e.g., "spec-driven")
- `artifacts`: Array of artifacts with their status ("done", "skipped", "ready", "blocked")
- `isPlanningComplete`: Boolean indicating if all planning artifacts are complete. Older CLI versions expose the same value as `isComplete`.
- `planningHome`, `changeRoot`, `artifactPaths`, and `actionContext`: path and scope context. Use these instead of assuming repo-local paths.
The artifact ids and paths come from the active schema - do NOT assume them, and do NOT branch on hardcoded artifact names. Custom schemas must work unchanged.
The files to edit are `artifactPaths.<id>.existingOutputPaths` - the concrete files that exist on disk, already glob-expanded for glob artifacts (e.g. `specs/**/*.md`). Do NOT write to `resolvedOutputPath`: for a glob artifact it is still the glob pattern, not a real file.
3. **Understand the request**
- If the user asked for a specific revision ("the design now uses X"), that is the starting edit.
- If they only said "update" / "make this coherent", treat it as a coherence review: read the existing artifacts and check them against each other for contradictions, gaps, and duplication.
4. **Read and reconcile**
- Read the artifact(s) the request touches and the change's other existing artifacts.
- Apply the requested edit. Then check every other existing artifact against it - in ANY direction: an edit to a later artifact may require revising an earlier one, not only the other way around. Build order is a useful reading order, not a constraint on which artifacts may be revised.
- Note everything that is now inconsistent, missing, or contradictory.
- Revise only files that already exist (`existingOutputPaths`). Do NOT create artifacts that don't exist yet, and do NOT invent new files under a glob artifact - note them and point the user to `/openspec-continue-change` to create them.
- If the change is already coherent, say so and make no edits.
5. **Confirm and apply, one artifact at a time**
- Show each proposed revision and why. Write only after the user confirms.
- If the user rejects a revision, do not write it - leave that artifact unchanged.
- When a substantial rewrite is needed, get that artifact's rules and template first:
```bash
openspec instructions "<artifact-id>" --change "<name>" --json
```
6. **Point to the next step (guidance only - NEVER act on it)**
- Artifacts still missing -> suggest `/openspec-continue-change` to create them.
- Change already implemented (tasks checked off / already applied) -> the code may no longer match the revised plan; suggest `/openspec-apply-change` to carry the delta into code.
- Everything done and implemented -> suggest `/openspec-archive-change`.
**Output**
After each invocation, show:
- Which artifacts were revised (and which proposed revisions were rejected)
- Anything deferred to `/openspec-continue-change` (not-yet-created artifacts or files)
- Where the change stands and the recommended next command
**Guardrails**
- Planning artifacts only - NEVER edit implementation code. If the revised plan implies code changes, stop and point to `/openspec-apply-change`.
- Use the artifact ids and paths reported by `openspec status`; never branch on hardcoded artifact names.
- Edit only the concrete files in `existingOutputPaths`; never write to a glob `resolvedOutputPath`.
- Do not advance the build frontier: no new artifacts, no new files under glob artifacts - that is `/openspec-continue-change`'s job.
- Confirm every edit with the user before writing.
- If the request changes the change's *intent* rather than refining it, first verify whether the optional `/openspec-new-change` workflow is available. If it is, recommend starting fresh with `/openspec-new-change` (the "Update vs. Start Fresh" heuristic). If it is unavailable, ask for a distinct unused change name and recommend `openspec new change "<new-change-name>"` instead.
+325 -6
View File
@@ -4,6 +4,134 @@
> Ситуация: кластер Garage v2.1 (RF=3) на vps01 + bigbox + vps02, WireGuard 10.8.0.0/24. > Ситуация: кластер Garage v2.1 (RF=3) на vps01 + bigbox + vps02, WireGuard 10.8.0.0/24.
> Задача: вывести статус кластера в браузер (Grafana + Prometheus + Loki). > Задача: вывести статус кластера в браузер (Grafana + Prometheus + Loki).
## Опыт: Vinograd WAN (Ростелеком) — ICMP-мониторинг внешнего канала (2026-09-08)
> Ситуация: UptimeKuma алертил про 100% потерю пингов на шлюз 83.239.50.145
> (канал «Винный город», РТК). Задача — мониторить ОБА адреса канала (шлюз +
> наше оборудование) в нашем стеке с графиками RTT каждые 30с.
### 18. ICMP-пробы через blackbox-exporter — модуль `icmp`
blackbox-exporter поддерживает ICMP-пробы (prober: icmp). Метрики:
- `probe_success` — 1/0 (успех пробы)
- `probe_icmp_duration_seconds{phase="rtt"}` — RTT в секундах
- `probe_icmp_reply_hop_limit` — TTL ответа
Нюансы:
- В контейнере (host-network, root) ICMP работает без доп. настроек — проверил
`docker exec blackbox-exporter id` → root. В не-root окружении нужен
`setcap cap_net_raw+ep` или `net.ipv4.ping_group_range`.
- **Важно про YAML:** в `static_configs` таргеты — это список, `labels` относится
к списку целиком, а НЕ к каждому элементу отдельно. Ошибка синтаксиса ловится
`promtool check config`.
### 19. Scrape job с интервалом 30s и relabel instance
```yaml
- job_name: vinograd_wan
scrape_interval: 30s
metrics_path: /probe
params:
module: [icmp]
static_configs:
- targets: [83.239.50.145, 83.239.50.146]
relabel_configs:
# __address__ → instance: человекочитаемые имена для легенд Grafana
- source_labels: [__address__]
regex: '83\.239\.50\.145.*'
target_label: instance
replacement: vinograd-gw-83.239.50.145
...
# __address__ → реальный адрес blackbox (multi-target exporter pattern)
- target_label: __address__
replacement: 127.0.0.1:9115
```
- `scrape_interval: 30s` на уровне job — работает (проверено: точки каждые 30с).
- Regex с точками надо экранировать (`\.`), иначе 83.239.50.145 совпадёт с .146.
- relabel применяется по-порядку; сначала маппим instance, потом __address__ → blackbox.
- Проверка таргетов: `curl http://127.0.0.1:9090/api/v1/targets` → vinograd_wan 2 targets UP.
### 20. Алерт на probe_success
```yaml
- name: vinograd
rules:
- alert: VinogradRostelecomDown
expr: probe_success{job="vinograd_wan"} == 0
for: 2m
labels: {severity: critical}
```
- `for: 2m` при scrape 30s ≈ 4 пробы подряд. `promtool check config` → 7 rules found.
### 21. Grafana dashboard provisioning и ретеншн
- Дашборд кладём в `grafana/dashboards/vinograd-wan.json` — provisioner
(updateIntervalSeconds: 30) сам импортирует в фолдере Garage; рестарт не нужен.
Проверка: в grafana.db появился dashboard с uid=vinograd-wan.
- **Ретеншн «неделя»:** retention в Prometheus глобальный (--storage.tsdb.retention.time=30d
в этом стеке). Для 7 дней ровно нужен отдельный инстанс — здесь оставили 30d
(перекрывает неделю с запасом). Не пытаться задать retention per-job — его нет.
### 22. Наблюдение: шлюз РТК не пингуется, но оборудование пингуется
- 83.239.50.146 (наше оборудование) — probe_success=1, RTT ~13ms.
- 83.239.50.145 (шлюз) — probe_success=0 (не отвечает на ICMP). Совпадает с
алертом UptimeKuma. Это реальная авария, а не ошибка конфига: blackbox
корректно видит недоступность шлюза.
### 23. Read-only пользователь Grafana: provisioning НЕ работает (OSS), только UI (2026-09-08)
Задача: дать сотруднику IT Винограда read-only доступ к дашбордам
(`it@vinogorod.ru`, роль Viewer). Попытки автоматизировать — провалились:
- **Файловое provisioning пользователей** (`grafana/provisioning/access-control/users.yml`)
в Grafana 11 OSS **не обрабатывается**: в логах старта только
dashboards/datasources/alerting/plugins; access-control — фича
Enterprise/Cloud (`security.provisioning`). Файл молча игнорируется, даже с
валидным YAML.
- **API create**: `POST /api/users` → 404 (в OSS недоступно). `POST /api/login`
(JSON) → 401 даже при верном пароле; а **Basic auth работает**
(`curl -u estorozhenko:пароль /api/user` → 200).
- Итог: пользователя можно создать **только в UI** (Administration → Users →
New user, роль Viewer). Пароль задаётся при создании.
Проверка после создания:
```bash
curl -s -u 'it@vinogorod.ru:1qazXSW2' http://127.0.0.1:3001/api/user # → 200
curl -s -u 'estorozhenko:ПАРОЛЬ' http://127.0.0.1:3001/api/orgs/1/users # role=Viewer
curl -s -u 'it@vinogorod.ru:1qazXSW2' http://127.0.0.1:3001/api/users # → 403 (read-only)
```
### 24. Ошибка "Datasource __grafana__ was not found" при открытии дашборда (2026-09-08)
Симптом: на `https://grafana.nixg.ru/d/vinograd-wan/vinograd-wan` выскакивало
окно "Failed to retrieve datasource / Datasource __grafana__ was not found".
Панели при этом в порядке (Prometheus uid есть), а **секция `annotations`**
в JSON дашборда ссылалась на встроенный датасорс:
```json
"annotations": { "list": [ { "builtIn": 1,
"datasource": {"type": "grafana", "uid": "__grafana__"}, ... } ] }
```
- `__grafana__` — встроенный datasource Grafana (аннотации/алерты). В нашей
БД `data_source` его НЕТ (только Prometheus и Loki) → Grafana 11 OSS не
может его найти и показывает ошибку. Панели не используют аннотации —
секция добавляется в JSON автоматически при создании (шаблон),
но в файле она бесполезна.
- **Фикс:** очистить `annotations.list` в файле дашборда:
```bash
jq '.annotations.list = []' grafana/dashboards/vinograd-wan.json > /tmp/vw.json \
&& mv /tmp/vw.json grafana/dashboards/vinograd-wan.json
```
Провайдер дашбордов перечитывает файл **каждые 30с** (рестарт не нужен),
в БД появляется version 2 с `"list": []` — ошибка исчезает.
- Эталон: рабочий `garage-cluster.json` всегда имеет `"annotations": {"list": []}`.
- Проверка из БД: `docker cp grafana:/var/lib/grafana/grafana.db /tmp/gf.db` →
`SELECT data FROM dashboard WHERE uid='vinograd-wan'` → `__grafana__` отсутствует.
## Ключевые находки / грабли ## Ключевые находки / грабли
### 1. Admin API Garage v2.1 слушает ОТДЕЛЬНЫЙ порт (`[admin] api_bind_addr`) ### 1. Admin API Garage v2.1 слушает ОТДЕЛЬНЫЙ порт (`[admin] api_bind_addr`)
@@ -124,6 +252,138 @@ grafana.nixg.ru {
- В логах caddy много ошибок renew для старых доменов — они имеют - В логах caddy много ошибок renew для старых доменов — они имеют
уже выпущенные сертификаты в caddy_data, работает всё. уже выпущенные сертификаты в caddy_data, работает всё.
## Опыт: VESTI — textfile-метрики + встроенный Prometheus Alerting (2026-09-13)
> Ситуация: добавлен дашборд и алерты на компоненты /opt/vesti. Проект
> управляется через OpenSpec: все изменения — только через openspec/change.
### 22. Textfile-коллектор node-exporter для метрик сервиса
node-exporter умеет отдавать произвольные метрики из файлов каталога
(`--collector.textfile.directory=...` ← `ARGS` в /etc/default/prometheus-node-exporter).
Скрипт раз в минуту пишет файл в Prometheus-формате (1 метрика = 1 строка
`name{labels} value`), node-exporter отдаёт их в /metrics — и Prometheus
тащит их обычным job'ом `node` (тот же 9100). Для сервиса достаточно:
`/opt/vesti/scripts/vesti-metrics.sh` → `/var/lib/node_exporter/textfile_collector/vesti.prom`.
Грабля: каталог textfile принадлежит пользователю `prometheus`, скрипт от
root должен делать `chown prometheus:prometheus` (иначе файл появится, но
node-exporter его не прочитает — права!). Запуск: cron.d (root) — каждую минуту.
### 23. Prometheus-алерты: файл alerts.yml + `/api/v1/rules`
В этом стеке алерты — НЕ Grafana, а встроенный механизм Prometheus:
`prometheus.yml` → `rule_files: /etc/prometheus/alerts.yml` (группы/правила в
YAML). Правила видны в `http://127.0.0.1:9090/api/v1/rules` (JSON: groups,
state, query). Перечитываются ТОЛЬКО рестартом прометеуса
(`docker restart prometheus`), НЕ автоматически.
Грабля: `promtool check config /etc/prometheus/prometheus.yml` считает
alerts.yml правило-файлом: «FAILED: field groups not found» при прямом
проверке alerts.yml — это НОРМАЛЬНО (это не самостоятельный конфиг;
проверять через рrometheus.yml, который находит 13 rules).
### 24. OpenSpec для /opt/monitoring
Все изменения проекта — ТОЛЬКО через `openspec/change/` (proposal/design/tasks/
specs/). Формат spec: `## ADDED Requirements` + `### Requirement: X` +
`#### Scenario:` (иначе validate ругается). После внесения: `openspec validate`,
применить, `openspec archive --yes` + обновить STATUS/README/WALKTHROUGH.
## Опыт: дашборд Grafana пустой (NO DATA) — 3 причины подряд (2026-09-02)
> Дата: 2026-09-02. Ситуация: после этапа 6 (tproxy) пользователь сообщил,
> что дашборд Garage Cluster в Grafana полностью пустой (NO DATA).
### 12. NO DATA #1: у datasource не задан `uid` — Grafana генерит случайный
В `grafana/provisioning/datasources/datasources.yml` у Prometheus-датасорса
НЕ был задан `uid`. Grafana при провиженинге присваивает датасорсу **случайный
UID** (в БД видно для Loki: `P8E80F9AEF21F6940`), а все панели дашборда
ссылались на `"uid": "Prometheus"`. Панели искали датасорс с таким uid — не
находили → NO DATA.
**Решение:** в datasources.yml добавить датасорсу явный `uid: Prometheus`
(совпадение 1-в-1 с uid в панелях). Дашборды провиженятся каждые 30с,
а datasources — **только при старте** Grafana → после правки `docker compose
restart grafana`. Проверка из БД (`docker cp grafana:/var/lib/grafana/grafana.db`):
`SELECT name, uid, url FROM data_source` → uid стал `Prometheus`.
### 13. NO DATA #2: Prometheus на host-сети, а datasource URL `http://prometheus:9090`
Prometheus и blackbox запущены с `network_mode: host` (нужно для WG 10.8.0.0/24),
а Grafana — в bridge-сети. Внутри сети Grafana имя `prometheus` **не
резолвится** (проверено `docker exec grafana getent hosts prometheus` →
NO_RESOLVE), поэтому URL `http://prometheus:9090` недостижим → панели без
данных. Loki в той же bridge-сети — резолвится нормально.
**Решение:** datasource URL заменить на `http://172.28.0.1:9090` — IP хоста
со стороны bridge-сети Grafana (шлюз сети = адрес хоста). Определяется так:
```
docker inspect grafana --format '{{range .NetworkSettings.Networks}}{{.Gateway}} {{end}}'
# → 172.28.0.1
```
Проверено из контейнера: `wget http://172.28.0.1:9090/-/healthy` → OK.
ВАЖНО: шлюз bridge-сети docker стабилен (подсеть фиксированная),
но если пересоздать сеть — IP может смениться.
### 14. NO DATA #3 (частично): панели ссылались на НЕСУЩЕСТВУЮЩИЕ метрики
После починки datasource tproxy-панели ожили, а garage-панели (blocks,
RPC node health, S3 req/sec) остались пустыми. Причина: панели были написаны
под НЕСУЩЕСТВУЮЩИЕ имена метрик, которых в Prometheus нет:
- `garage_block_count` — в реальности `block_resync_queue_length` / `block_resync_errored_blocks`
- `garage_rpc_node_health_is_up` — в реальности `cluster_layout_node_connected`
- `garage_api_s3_request_counter` — в реальности `api_s3_request_counter`
Метрики Garage из admin API (10.8.0.x:3903) **не имеют префикса `garage_`**:
`api_s3_request_counter`, `api_s3_request_duration_*`, `block_resync_*`,
`cluster_*` (connected_nodes, healthy, partitions_all_ok, layout_node_connected),
`table_*`, `rpc_*`. Префикс `garage_` есть только у `garage_build_info`,
`garage_local_disk_avail/total`, `garage_replication_factor`.
Проверка реальных метрик: `curl 'http://127.0.0.1:9090/api/v1/label/__name__/values'`
и фильтр по `job=garage`. Панели переписаны на реальные метрики:
- Garage blocks (resync): `block_resync_queue_length` + `block_resync_errored_blocks`
- Garage node health (layout): `cluster_layout_node_connected` (легенда `{{ role_zone }}`)
- S3 requests /sec: `rate(api_s3_request_counter[5m])`
### 15. Алерты ссылались на те же несуществующие метрики
В alerts.yml алерты GarageResyncErrors и GarageNodeUnstable использовали те же
фантомные имена (`garage_block_resync_error_count`, `garage_rpc_node_health_is_up`),
поэтому никогда не сработали бы. Исправлено на реальные:
`block_resync_errored_blocks > 0` и `cluster_layout_node_connected == 0`.
После правки `promtool check config` → 6 rules found, все eval.
### 16. Relabel `instance` → hostname вместо IP (читаемые легенды)
По умолчанию instance = адрес скрейпа: у tproxy это `127.0.0.1:18081`
(локальный конец SSH-туннеля), у garage `10.8.0.x:3903`, у node `10.8.0.x:9100`.
В Grafana легенды показывали IP — некрасиво и непонятно. Не путать: это НЕ
«мониторинг локального интерфейса», просто лейбл instance = транспорт туннеля.
**Решение (на уровне Prometheus, не панелей)** — relabel_configs в пром.yml:
```yaml
relabel_configs:
- target_label: instance
replacement: vps03:8081 # для job tproxy
- source_labels: [__address__] # для garage/node — маппинг IP → hostname
regex: 10.8.0.2:3903
target_label: instance
replacement: bigbox:3903
```
Теперь легенды: `vps01:3903`, `bigbox:3903`, `vps02:3903` (garage),
`vps03:8081` (tproxy), `vps01:9100` и т.д. (node). Правило: правим источник
(prom.yml relabel), а не легенды в каждой панели — консистентно везде
(панели, explore, алерты).
### 17. Grafana provisioning: дашборды — каждые 30с, datasources — только при старте
Дашборды перечитываются автоматически (updateIntervalSeconds, по умолчанию 30с),
а datasources провиженятся только при старте Grafana. После правки
datasources.yml — обязателен `docker compose restart grafana`.
## Опыт: tproxy-server (WEB Proxy для Telegram Desktop) на vps03 ## Опыт: tproxy-server (WEB Proxy для Telegram Desktop) на vps03
> Дата: 2026-08-31 > Дата: 2026-08-31
@@ -163,14 +423,30 @@ backend_dial_failures_total, bytes_up_total, bytes_down_total, limit_hits_total.
Ключевые для алертов: `backend_dial_failures_total` (рост = бэкенд недоступен), Ключевые для алертов: `backend_dial_failures_total` (рост = бэкенд недоступен),
`limit_hits_total` (DDOS/перегруз), `sessions_live` (активность). `limit_hits_total` (DDOS/перегруз), `sessions_live` (активность).
### C. Подключение к Prometheus (стек /opt/monitoring, bigbox) ### C. Подключение к Prometheus (стек /opt/monitoring, bigbox) — РЕШЕНО через SSH-туннель
- vps03 НЕ в WG (10.8.0.0/24 = vps01/bigbox/vps02), поэтому прямого доступа - vps03 НЕ в WG (10.8.0.0/24 = vps01/bigbox/vps02), поэтому прямого доступа
к 8081 с bigbox нет. Prometheus в стеке — `network_mode: host` (видит внешние IP). к 8081 с bigbox нет. Prometheus в стеке — `network_mode: host` (видит внешние IP).
- План: nft-правило на vps03 (разрешить TCP 8081 с publIP bigbox 178.176.197.2), - Изначальный план (nft-правило на vps03: разрешить TCP 8081 с publIP bigbox
job 'tproxy' в prometheus.yml, targets ['77.67.89.154:8081']. 178.176.197.2) **не сработал**: tproxy-server слушает `admin_listen: 127.0.0.1:8081`,
- ВАЖНО: admin-эндпоинты слушают только loopback. Открывать наружу ТОЛЬКО запрос с bigbox к `77.67.89.154:8081` вернул `Connection refused` (соединение
по источнику (ip saddr bigbox), не публиковать всем. упёрлось в отсутствующего наружу слушателя, а не в файрвол).
- **Решение:** SSH-туннель bigbox → vps03 (паттерн systemd, как telegram-tunnel):
```ini
# /etc/systemd/system/tproxy-tunnel.service (bigbox)
User=estorozhenko
ExecStart=/usr/bin/ssh -i /home/estorozhenko/.ssh/hostkeyVPS \
-L 127.0.0.1:18081:127.0.0.1:8081 -N \
-o ServerAliveInterval=30 -o ServerAliveCountMax=3 \
-o ExitOnForwardFailure=yes -o StrictHostKeyChecking=accept-new \
root@77.67.89.154
Restart=always
```
Prometheus скрейпит `127.0.0.1:18081`. tproxy-server config/файрвол НЕ трогаются,
метрики остаются на loopback.
- **ВНИМАНИЕ с портом:** сначала взял `8081` — но он на bigbox уже занят ICQ-веб-чатом
(Converse.html на 0.0.0.0:8081), туннель конфликтовал и отдавал HTML чата вместо
метрик. Порт на bigbox выбирать свободный (взял 18081).
- readyz ходит ко ВСЕМ профилям: если добавить профиль с мёртвым бэкендом — - readyz ходит ко ВСЕМ профилям: если добавить профиль с мёртвым бэкендом —
readyz станет 503 (фича, учтена при алертах). readyz станет 503 (фича, учтена при алертах).
@@ -203,10 +479,53 @@ curl -G "http://127.0.0.1:3100/loki/api/v1/query_range" \
--data-urlencode 'query={container="garage"}' --data-urlencode 'limit=3' --data-urlencode 'query={container="garage"}' --data-urlencode 'limit=3'
``` ```
### 25. vps03 подключён к WireGuard + node-exporter (2026-09-08)
Задача: мониторинг vps03 (77.67.89.154, Debian 13) системных метрик
(диски/память/сеть/доступность) в Grafana, не светя порт наружу.
- **Топология WG** (была): hub = vps01 (10.8.0.1, pubkey ZAvz4xCE…), пиры
bigbox (10.8.0.2), vps02 (10.8.0.4). Все слушают 51820.
- **Новая нода**: vps03 = **10.8.0.3**, ключ `HBTzrS86SZ+…`. vps03 инициирует
туннель к hub (Endpoint 5.129.217.146:51820, AllowedIPs 10.8.0.0/24).
- **Главный грабль:** hub видит vps03, но bigbox/vps02 НЕ могут ответить в
10.8.0.3: WireGuard дропает пакеты, чей src-адрес не в AllowedIPs пира.
Пришлось на bigbox и vps02 **расширить AllowedIPs пира vps01** до
`10.8.0.1/32, 10.8.0.3/32` — тогда трафик к vps03 идёт через hub, а ответы
возвращаются самому vps03. (vps03 → hub работает сразу, т.к. у vps03
AllowedIPs = 10.8.0.0/24; но в обратную сторону — нет.)
- **node-exporter на vps03**: apt install, слушает `10.8.0.3:9100` (в
/etc/default/prometheus-node-exporter: `ARGS="--web.listen-address=10.8.0.3:9100"`).
Публичный IP 77.67.89.154:9100 → connection refused (наружу не светит).
- **prometheus.yml**: job `node` получает 4-й таргет `10.8.0.3:9100` (host=vps03,
instance=vps03:9100) + relabel.
- **Дашборд**: `grafana/dashboards/nodes.json` (uid `nodes`, title "Nodes") —
6 панелей: availability (up{job="node"}, stat), disk free GB, memory available
GB, network RX/TX B/s, load1. Provisioner импортирует автоматически.
- Проверка: `curl http://127.0.0.1:9090/api/v1/targets` → node ×4 все up.
### 26. Per-node дашборды: сеть RX/TX + утилизация канала (2026-09-08)
Вместо одного `nodes.json` — **4 дашборда** `node-<host>.json`
(uid `node-<host>`), генератор `/tmp/gen_nodes_dash.py`.
- **Метрики сети в node-exporter 1.9** (apt, Debian 13): НЕ `node_net_bytes_*`,
а **`node_network_receive_bytes_total` / `node_network_transmit_bytes_total`**
(counter, device=<iface>) + **`node_network_speed_bytes`** (скорость линка, Б/с).
- **Утилизация канала**: `(rate(rx[5m])+rate(tx[5m]))*8 / speed * 100`.
- **Грабль: виртуалки дают `node_network_speed_bytes` = -125000** (vps01 eth0,
vps02 enp3s0; sysfs speed=-1 → exporter умножает на 125000, знак — т.к. -1).
На физических (vps03 ens1, bigbox eno1) = 1.25e+08 (1000 Мбит/с).
- Решение: для vps01/vps02 в формуле утилизации хардкод `125000000` Б/с,
для vps03/bigbox — `node_network_speed_bytes`.
- Физические интерфейсы: vps01=eth0, vps02=enp3s0, vps03=ens1, bigbox=eno1.
Контейнерные (docker0/veth*/br-*) исключены.
- Проверка: дашборды импортированы (БД: uid node-*), prom targets node ×4 up.
## Что осталось сделать / TODO ## Что осталось сделать / TODO
- [x] Развернуть стек (docker compose up -d) в /opt/monitoring - [x] Развернуть стек (docker compose up -d) в /opt/monitoring
- [x] Node-exporter на всех 3 хостах (10.8.0.x:9100) - [x] Node-exporter на всех 4 хостах (10.8.0.x:9100; vps03 = 10.8.0.3)
- [x] Дашборд Garage в Grafana (provisioning + JSON) - [x] Дашборд Garage в Grafana (provisioning + JSON)
- [x] `up{job="garage"}` в Prometheus, Grafana :3001 - [x] `up{job="garage"}` в Prometheus, Grafana :3001
- [x] Логи garage через promtail → Loki → Grafana - [x] Логи garage через promtail → Loki → Grafana
+33 -38
View File
@@ -24,14 +24,14 @@
**Результат:** все 3 ноды отдают метрики, кластер HEALTHY. **Результат:** все 3 ноды отдают метрики, кластер HEALTHY.
## Этап 2. Стек мониторинга на bigbox (в работе 🔄) ## Этап 2. Стек мониторинга на bigbox (выполнено ✅)
- [x] Создать `/opt/monitoring/` с docker-compose.yml, prometheus.yml, alerts.yml - [x] Создать `/opt/monitoring/` с docker-compose.yml, prometheus.yml, alerts.yml
- [x] Prometheus: таргеты garage x3 (:3903), health x3, node-exporters, self - [x] Prometheus: таргеты garage x3 (:3903), health x3, node-exporters, self
- [ ] Node-exporter на 3 хостах (10.8.0.x:9100) - [x] Node-exporter на 3 хостах (10.8.0.x:9100)
- [ ] Запустить `docker compose up -d` - [x] Запустить `docker compose up -d`
- [ ] Проверить Prometheus :9090 (`up{job="garage"}`, targets) - [x] Проверить Prometheus :9090 (`up{job="garage"}`, targets)
- [ ] Grafana :3001: datasource Prometheus, дашборд Garage, alerts - [x] Grafana :3001: datasource Prometheus, дашборд Garage, alerts
## Этап 3. Логи (после метрик 🔄) ## Этап 3. Логи (после метрик 🔄)
@@ -78,41 +78,36 @@ Caddy → tproxy-server:8080 → MTProxy:2398. Admin-эндпоинты:
Задача: Задача:
- [ ] Обеспечить доступ Prometheus (bigbox) к `:8081` на vps03. - [x] **РЕШЕНО через SSH-туннель (не nft!)** — tproxy-server слушает `admin_listen: 127.0.0.1:8081`
vps03 НЕ в WG-сети (10.8.0.0/24 = vps01/bigbox/vps02). только на loopback (осознанно), простое nft-правило из плана НЕ сработало бы:
Варианты: соединение с bigbox упёрлось бы в `Connection refused` (проверено).
a) открыть 8081 на vps03 для IP bigbox в nft (правило в Поэтому добавлен systemd-сервис `tproxy-tunnel.service` на bigbox:
`/etc/tproxy-server/firewall.nft` или отдельный файл) — самый простой;
b) добавить vps03 в WG (если хочется закрытый контур);
c) node-exporter + textfile-коллектор — не наш случай (метрики уже в HTTP).
Рекомендация: (a).
**Конкретные шаги (рекомендуемый вариант a):**
- [ ] На vps03 (77.67.89.154) добавить nft-правило, разрешающее TCP 8081
с публичного IP bigbox **178.176.197.2** (проверить актуальный IP
bigbox перед выполнением: `curl -s https://api.ipify.org`):
``` ```
nft add rule inet filter input ip saddr 178.176.197.2 tcp dport 8081 accept ssh -i /home/estorozhenko/.ssh/hostkeyVPS \
-L 127.0.0.1:18081:127.0.0.1:8081 -N \
root@77.67.89.154
``` ```
(или внести в `/etc/tproxy-server/firewall.nft`, если он в авто-загрузке) Порт 18081 (а НЕ 8081 — 8081 на bigbox занят ICQ-веб-чатом!).
- [ ] Проверить с bigbox: Prometheus скрейпит `127.0.0.1:18081`.
`curl -s http://77.67.89.154:8081/metrics | head` - [x] Job 'tproxy' в prometheus.yml: `targets: ['127.0.0.1:18081']` → up=1, метрики скрейпятся.
→ должен вернуть метрики tproxy (не timeout/refused). - [x] Дашборд Garage Cluster дополнен 3 панелями tproxy: live sessions/streams, traffic /sec,
- [ ] Добавить job 'tproxy' в `/opt/monitoring/prometheus.yml`: backend errors.
```yaml - [x] Алерты TProxyDown (critical) и TProxyBackendErrors (warning) в alerts.yml.
- job_name: 'tproxy' - [x] **Починен существующий баг: alerts.yml не был смонтирован в Prometheus** — garage-алерты
static_configs: не работали (0 групп правил). Добавлен volume `./alerts.yml:/etc/prometheus/alerts.yml:ro`
- targets: ['77.67.89.154:8081'] в docker-compose.yml. Теперь загружено 6 правил (garage + tproxy).
labels: - [x] **Починены пустые панели (NO DATA, 2026-09-02)**: три причины подряд —
host: vps03 (1) у datasource не был задан `uid` (Grafana генерила случайный), добавлен
service: tproxy `uid: Prometheus` в datasources.yml; (2) URL `http://prometheus:9090` не
``` резолвился из Grafana (Prometheus на host-сети), заменён на `http://172.28.0.1:9090`
- [ ] `docker compose restart prometheus` в /opt/monitoring, (шлюз bridge-сети Grafana = адрес хоста); (3) панели и алерты ссылались
проверить `up{job="tproxy"}` на 127.0.0.1:9090 = 1. на несуществующие метрики (`garage_block_count`, `garage_rpc_node_health_is_up`,
- [ ] (Опционально) blackbox-job для внешнего HTTPS-чека vps03.nixg.ru: `garage_api_s3_request_counter`) — переписаны на реальные (`block_resync_*`,
`probe_success` — контролирует Caddy+сайт снаружи. `cluster_layout_node_connected`, `rate(api_s3_request_counter[5m])`). Подробности —
- [ ] Дашборд в Grafana: сессии, трафик (bytes_up/down rate), ошибки бэкенда. EXPERIENCE.md п.12-17.
- [ ] Алерт: `tproxy_backend_dial_failures_total` растёт, или `/readyz` 503 - [x] Relabel `instance` → hostname в prometheus.yml (garage/node/tproxy):
(можно ч/з blackbox по HTTP :8081, если открыт). `10.8.0.x` → `vps01|bigbox|vps02`, `127.0.0.1:18081` → `vps03:8081`.
Легенды в Grafana — hostname, а не IP.
Заметки: Заметки:
- `/metrics` и admin-эндпоинты слушают ТОЛЬКО loopback (127.0.0.1:8081). При - `/metrics` и admin-эндпоинты слушают ТОЛЬКО loopback (127.0.0.1:8081). При
+36
View File
@@ -0,0 +1,36 @@
# PRD — Мониторинг кластера
## Цель
Единый стек мониторинга всех серверов и сервисов (Garage, tproxy, инфраструктура)
на bigbox с веб-доступом через Grafana, без публикации служебных портов наружу.
## Пользователи
- estorozhenko (admin) — полный доступ
- гостевая учётка (readonly) — просмотр дашбордов (пароль 1qazXSW2)
## Функциональные требования
- Метрики: Garage-кластер (3 ноды), система каждой ноды (CPU/память/диск/сеть/доступность),
tproxy (sessions/streams/traffic/backend errors), blackbox (ICMP RTT)
- Сеть: по каждому физическому интерфейсу на каждой ноде отдельный график
(RX/TX вместе) + утилизация канала (%)
- Дашборды: per-node (Node <host>) + Garage Cluster + Vinograd WAN + tproxy
- Алерты: garage (down/unstable), tproxy, vinograd WAN — правила в alerts.yml
- Логи: promtail → Loki → Grafana (retention 7d)
## Нефункциональные
- Порт node-exporter и admin-эндпоинты НЕ публикуются наружу (наследие: vps03 —
только WG; garage :3903 — только WG-подсеть; tproxy :8081 — loopback+туннель)
- Катастрофоустойчивость: git gitverse (источник истины) + gitea pull-mirror на bigbox
- Retention: 30d (прометеевский TSDB)
## Границы (НЕ делаем)
- Трейсы (до них «не доросли»)
- Мониторинг сети за роутером (кроме ICMP-пробы vinograd)
- Панель авторизации вне Grafana
## Критерии готовности
- [x] Все ноды и сервисы в Prometheus (up)
- [x] Дашборды: per-node + garage + vinograd-wan + tproxy видны в Grafana
- [x] Алерты доставляются в Grafana Alerts
- [x] README/EXPERIENCE/PLAN/STATUS/TODO/WALKTHROUGH актуальны
- [ ] Разные admin_token/metrics_token; пароль Grafana; уведомления (Telegram)
+143 -11
View File
@@ -1,6 +1,7 @@
# Monitoring stack — Garage cluster (vps01 + bigbox + vps02) # Monitoring stack — Garage cluster (vps01 + bigbox + vps02) + Vinograd WAN
Стек мониторинга для S3-кластера Garage (репликация RF=3, WireGuard 10.8.0.0/24). Стек мониторинга для S3-кластера Garage (репликация RF=3, WireGuard 10.8.0.0/24)
и внешнего канала связи объекта «Винный город» (провайдер Ростелеком).
Расположен на **bigbox** в `/opt/monitoring`. Расположен на **bigbox** в `/opt/monitoring`.
## Архитектура ## Архитектура
@@ -51,6 +52,9 @@ grafana.nixg.ru → 87.242.100.206 (vps02) → caddy → reverse_proxy 10.8.0.2:
`docker exec caddy caddy reload --config /etc/caddy/Caddyfile` `docker exec caddy caddy reload --config /etc/caddy/Caddyfile`
- Сертификат Let's Encrypt выпускается автоматически. - Сертификат Let's Encrypt выпускается автоматически.
- Логин: **estorozhenko** (сменён с admin через UI), пароль — задан пользователем. - Логин: **estorozhenko** (сменён с admin через UI), пароль — задан пользователем.
- Read-only доступ: **it@vinogorod.ru** (роль **Viewer**, создан вручную в UI,
пароль `1qazXSW2`). Только просмотр дашбордов, без правки. Учётка в
`grafana-data` (переживает пересоздание контейнера).
## Garage admin API (метрики) ## Garage admin API (метрики)
@@ -112,23 +116,104 @@ curl -X POST -H "Authorization: token GITEA_TOK" \
|--------------------|--------------------------------|----------| |--------------------|--------------------------------|----------|
| GarageNodeDown | `up{job="garage"} == 0` (2м) | critical | | GarageNodeDown | `up{job="garage"} == 0` (2м) | critical |
| GarageNoQuorum | `<2` нод up (2м) | critical | | GarageNoQuorum | `<2` нод up (2м) | critical |
| GarageResyncErrors | `garage_block_resync_error_count > 0` (10м) | warning | | GarageResyncErrors | `block_resync_errored_blocks > 0` (10м) | warning |
| GarageNodeUnstable | RPC health down (5м) | warning | | GarageNodeUnstable | `cluster_layout_node_connected == 0` (5м) | warning |
| TProxyDown | `up{job="tproxy"} == 0` (2м) | critical |
| TProxyBackendErrors| `increase(tproxy_backend_dial_failures_total[5m]) > 0` (10м) | warning |
| VinogradRostelecomDown | `probe_success{job="vinograd_wan"} == 0` (2м) | critical |
## tproxy-server (vps03) — метрики WEB Proxy (этап 6, в работе) Примечание: метрики Garage из admin API (:3903) НЕ имеют префикса `garage_` —
это `api_s3_request_counter`, `block_resync_*`, `cluster_*`. Префикс `garage_`
только у `garage_build_info`, `garage_local_disk_*`, `garage_replication_factor`.
## Vinograd WAN — внешний канал «Винный город» (Ростелеком)
Объект «Винный город» (г. Геленджик, ул. Туристическая, 25), канал Ростелеком
(договор Бастион, Static IP). Адреса из «Реестра внешних каналов связи.ods»
(закладка «Винный город»):
| Адрес | Роль |
|-------|------|
| 83.239.50.145 | Шлюз (gateway) — поднимается от РТК |
| 83.239.50.146 | Наше оборудование (CPE, Static IP, /30) |
Мониторинг через **blackbox-exporter (ICMP-проба)** → Prometheus job `vinograd_wan`:
- Интервал scrape: **30s** (графики скорости ответа каждые 30 секунд)
- Метрики:
- `probe_success{job="vinograd_wan"}` — доступность (1/0)
- `probe_icmp_duration_seconds{job="vinograd_wan",phase="rtt"}` — RTT, сек
- Лейблы `instance`: `vinograd-gw-83.239.50.145`, `vinograd-cpe-83.239.50.146`
- Алерт: **VinogradRostelecomDown** (critical, 2м подряд недоступен)
- Grafana: дашборд **Vinograd WAN** (RTT ms + availability), панели в фолдере Garage
Retention: глобальный 30d (прометеевский TSDB) — данные хранятся минимум неделю,
что покрывает требование «хранить неделю» с запасом (жёсткий 7d для одного job
требовал бы отдельного инстанса Prometheus).
Проверка вручную:
```bash
# ICMP-проба через blackbox (debug)
curl -s "http://127.0.0.1:9115/probe?target=83.239.50.146&module=icmp&debug=true"
# данные в Prometheus
curl -sG 'http://127.0.0.1:9090/api/v1/query' \
--data-urlencode 'query=probe_success{job="vinograd_wan"}'
```
> Статус 2026-09-08: шлюз 83.239.50.145 НЕ отвечает на ICMP (probe_success=0,
> совпадает с алертом UptimeKuma 08:07 MSK). Оборудование 83.239.50.146
> отвечает ~13ms. Алерт VinogradRostelecomDown в состоянии FIRE до восстановления
> канала — это корректное отражение реальной аварии.
## tproxy-server (vps03) — метрики WEB Proxy (этап 6, РЕШЕНО ✅)
tproxy-server (Telegram Desktop WEB Proxy) развёрнут на **vps03** (77.67.89.154), tproxy-server (Telegram Desktop WEB Proxy) развёрнут на **vps03** (77.67.89.154),
admin-эндпоинты на loopback :8081: `/healthz`, `/readyz`, `/metrics` admin-эндпоинты на loopback :8081: `/healthz`, `/readyz`, `/metrics`
(Prometheus-формат, 13 счётчиков `tproxy_*`). vps03 НЕ в WG-сети, поэтому (Prometheus-формат, 13 счётчиков `tproxy_*`).
для scrape с bigbox нужно:
1. nft-правило на vps03: разрешить TCP 8081 с publIP bigbox **178.176.197.2** > **2026-09-08: vps03 подключён к WireGuard (10.8.0.3)** — вместо SSH-туннеля.
(`nft add rule inet filter input ip saddr 178.176.197.2 tcp dport 8081 accept`) > Админ-эндпоинты tproxy остались на loopback, но туннель `tproxy-tunnel.service`
2. job `tproxy` в prometheus.yml: `targets: ['77.67.89.154:8081']` > удалён не был (безвреден, оставлен на всякий случай). Prometheus берёт node-метрики
3. `docker compose restart prometheus`, проверить `up{job="tproxy"}` > vps03 напрямую по WG: `10.8.0.3:9100`.
- Локальный порт **18081** (не 8081 — тот на bigbox занят ICQ-веб-чатом!)
- Prometheus скрейпит `127.0.0.1:18081` (job `tproxy`, host=vps03), но через
relabel_configs в prometheus.yml лейбл `instance` заменён на **`vps03:8081`**,
чтобы в Grafana легенды показывали hostname, а не IP туннеля.
Аналогично relabel сделан для garage (`10.8.0.x:3903` → `vps01|bigbox|vps02:3903`)
и node (`10.8.0.x:9100` → hostname).
- Дашборд: 3 панели tproxy (sessions/streams, traffic /sec, backend errors)
- Алерты: TProxyDown (critical), TProxyBackendErrors (warning)
Подробности — в PLAN.md (этап 6) и EXPERIENCE.md. Подробности — в PLAN.md (этап 6) и EXPERIENCE.md.
## Node-дашборды (диски/память/сеть/доступность) — этап 7, РЕШЕНО ✅
На каждую ноду (vps01, vps02, vps03, bigbox) — **отдельный дашборд**,
`grafana/dashboards/nodes/node-<host>.json` (uid `node-<host>`, title "Node <host>").
Генерируются скриптом `/tmp/gen_nodes_dash.py` (в git не хранится; при желании
перенести — в EXPERIENCE.md). Файлы дашбордов разложены по подпапкам
(папки в UI Grafana совпадают):
- `grafana/dashboards/` — garage-cluster.json → папка **Garage**
- `grafana/dashboards/nodes/` — node-*.json → папка **nodes**
- `grafana/dashboards/vinogorod/` — vinograd-wan.json → папка **vinogorod**
Панели (все по `{host="<host>"}`, job=node):
1. **Availability** — stat-панель `up{job="node",host=...}` (UP/DOWN)
2. **Disk free (GB)** — `node_filesystem_avail_bytes` (без tmpfs/overlay/squashfs)
3. **Memory available (GB)** — `node_memory_MemAvailable_bytes`
4. **Load** — `node_load1/5/15`
5. **Net <iface> RX/TX (bytes/s)** — `rate(node_network_{receive,transmit}_bytes_total[5m])`
6. **Net <iface> utilization (%)** — `(rx+tx)*8/speed*100`
Физические интерфейсы: vps01=eth0, vps02=enp3s0, vps03=ens1, bigbox=eno1.
Контейнерные (docker0/veth*/br-*) не включены.
Тонкость node-exporter 1.9: на виртуалках vps01/vps02 `node_network_speed_bytes`
= **-125000** (sysfs speed=-1 → умножается на 125000), поэтому для них скорость
линка захардкожена **125000000 Б/с (1 Gbit/s)** в формуле утилизации. На физических
(vps03/bigbox) используется сама метрика `node_network_speed_bytes` (=1.25e+08).
## Катастрофоустойчивость ## Катастрофоустойчивость
Проект хранится в **двух** git-репозиториях — на случай поломки bigbox: Проект хранится в **двух** git-репозиториях — на случай поломки bigbox:
@@ -141,3 +226,50 @@ admin-эндпоинты на loopback :8081: `/healthz`, `/readyz`, `/metrics`
Порядок работы: правки коммитятся и пушутся в gitverse (источник), gitea Порядок работы: правки коммитятся и пушутся в gitverse (источник), gitea
подтягивает их mirror'ом (каждые 8ч или вручную mirror-sync). Если bigbox подтягивает их mirror'ом (каждые 8ч или вручную mirror-sync). Если bigbox
сломается — репозиторий остаётся в облаке gitverse. сломается — репозиторий остаётся в облаке gitverse.
## VESTI (2026-09-13)
Проект /opt/vesti: web-интерфейс (:8400, systemd), publisher-контейнер
(docker `vesti-publisher`, :8410/healthz), SOCKS5-туннель Telegram (:1080),
SQLite БД (vesti.db).
### Метрики (textfile-коллектор node-exporter, job `node`, label `component="vesti"`)
Скрипт `/opt/vesti/scripts/vesti-metrics.sh` (запуск: cron.d `vesti-metrics`,
root, каждую минуту; альтернативно systemd-таймер `vesti-metrics.timer`)
пишет `/var/lib/node_exporter/textfile_collector/vesti.prom`. node-exporter
на bigbox: `ARGS="--web.listen-address=10.8.0.2:9100 --collector.textfile.directory=/var/lib/node_exporter/textfile_collector"`.
| Метрика | Смысл |
|---|---|
| vesti_web_up | web :8400 (systemd) |
| vesti_publisher_health | publisher :8410 /healthz |
| vesti_publisher_docker | контейнер publisher (Up) |
| vesti_telegram_tunnel | SOCKS5-туннель :1080 |
| vesti_web_systemd | systemd-юнит web |
| vesti_publisher_bot / _proxy / _channels | детали publisher |
| vesti_db_ok | SQLite vesti.db доступна |
| vesti_db_posts_total / _new_total / _published_total | счётчики БД |
| vesti_db_last_run_ts | Unix-ts последнего рan краулера |
### Дашборд
`grafana/dashboards/vesti/vesti.json` (uid `vesti`, папка **vesti**), 18 панелей.
Провайдер `vesti-dashboards` в `grafana/provisioning/dashboards/dashboards.yml`.
Проверка импорта: `sudo docker cp grafana:/var/lib/grafana/grafana.db /tmp/gf.db`
+ `SELECT uid,title FROM dashboard WHERE uid='vesti'`.
### Алерты (группа `vesti` в alerts.yml)
| Алерт | Выражение | Удержание | Severity |
|---|---|---|---|
| VestiWebDown | `vesti_web_up{component="vesti"} == 0` | 2m | critical |
| VestiPublisherDown | `vesti_publisher_health{component="vesti"} == 0` | 2m | critical |
| VestiPublisherContainerDown | `vesti_publisher_docker{component="vesti"} == 0` | 2m | critical |
| VestiTunnelDown | `vesti_telegram_tunnel{component="vesti"} == 0` | 2m | critical |
| VestiDbDown | `vesti_db_ok{component="vesti"} == 0` | 2m | critical |
| VestiDbStale | `time() - vesti_db_last_run_ts{component="vesti"} > 1800` | 5m | warning |
Реализовано через встроенный Prometheus Alerting (`rule_files` в prometheus.yml
→ alerts.yml; правила видны в `/api/v1/rules`). Перечитывается рестартом
прометеуса (`docker restart prometheus`).
+73
View File
@@ -0,0 +1,73 @@
# Мониторинг кластера — Статус
Обновлено: 2026-09-13 (добавлен дашборд VESTI)
## Текущее состояние
Стек мониторинга на bigbox (/opt/monitoring): Prometheus 2.53 + Grafana 11.1 +
Loki + promtail. Собирает: Garage-кластер (vps01/bigbox/vps02 :3903), системы
всех 4 нод (node-exporter :9100 по WireGuard), tproxy (vps03, loopback-туннель),
**VESTI (компоненты/доступность)**.
Дашборды: **4 per-node** (Node vps01/vps02/vps03/bigbox) + Garage Cluster +
Vinograd WAN + tproxy-панели + **VESTI**. vps03 подключён к WG (10.8.0.3), порты наружу
не светятся. Все 4 node-таргета up. Всё в git (gitverse + gitea mirror).
## VESTI (добавлено 2026-09-13)
- Метрики: textfile-коллектор node-exporter на bigbox, скрипт
`/opt/vesti/scripts/vesti-metrics.sh` (запуск: `/etc/cron.d/vesti-metrics`, каждую минуту, root).
Пишет `/var/lib/node_exporter/textfile_collector/vesti.prom`.
- Метрики `vesti_*` (label `component="vesti"`), собираются job'ом `node` (10.8.0.2:9100):
`vesti_web_up`, `vesti_publisher_health`, `vesti_publisher_bot`, `vesti_publisher_proxy`,
`vesti_publisher_channels`, `vesti_telegram_tunnel`, `vesti_publisher_docker`,
`vesti_web_systemd`, `vesti_db_ok`, `vesti_db_posts_total`, `vesti_db_new_total`,
`vesti_db_published_total`, `vesti_db_last_run_ts`.
- node-exporter на bigbox: `ARGS="--web.listen-address=10.8.0.2:9100 --collector.textfile.directory=/var/lib/node_exporter/textfile_collector"`
(в /etc/default/prometheus-node-exporter).
- Дашборд: `grafana/dashboards/vesti/vesti.json` (uid `vesti`, папка **vesti**), 18 панелей:
stat-панели доступности (web/publisher/Docker/tunnel/DB/systemd) + детали publisher
(bot/proxy/channels) + БД (posts/new/published + время с последнего рan) + история доступности.
- Провайдер `vesti-dashboards` добавлен в `grafana/provisioning/dashboards/dashboards.yml`
(path: /var/lib/grafana/dashboards/vesti).
- Запуск коллектора: cron.d `/etc/cron.d/vesti-metrics` (root, каждую минуту).
Также прописаны юниты `/etc/systemd/system/vesti-metrics.{service,timer}` —
таймер заведён и активен (systemctl is-enabled: enabled, is-active: active);
при перезагрузке можно переключиться на таймер, отключив cron.d.
## Сделано (за сессию 2026-09-08)
- vps03 подключён к WireGuard: 10.8.0.3/24, hub vps01 (5.129.217.146:51820),
пиры добавлены на vps01/bigbox/vps02 (AllowedIPs пира vps01 расширен до
`10.8.0.1/32, 10.8.0.3/32` для возвратного трафика — грабль зафиксирован)
- node-exporter установлен на vps03 (apt), слушает 10.8.0.3:9100 (наружу 77.67.89.154:9100 закрыт)
- prometheus.yml: job `node` → 4 таргета (10.8.0.1/2/3/4:9100), все up
- Дашборды: **отдельный на каждую ноду** (uid node-<host>), в каждом:
availability + disk free + memory + load + **Net <iface> RX/TX (bytes/s)** +
**Net <iface> utilization (%)** (физические интерфейсы, контейнерные исключены)
- README.md/EXPERIENCE.md обновлены (грабли: node_network_* метрики, speed=-125000 на виртуалках)
- **Дашборды разложены по подпапкам**: `grafana/dashboards/nodes/` (node-*) и
`grafana/dashboards/vinogorod/` (vinograd-wan) — в git и на диске; provisioning
`dashboards.yml` получил 3 провайдера (Garage / nodes / vinogorod); в UI Grafana
дашборды лежат в папках **Garage / nodes / vinogorod** (проверено по БД, коммит 4b0539d)
## В работе / Следующие шаги
- [ ] Разные admin_token / metrics_token в Garage (совпадают)
- [ ] Поменять пароль Grafana с дефолтного (пользователь estorozhenko, пароль сменён вручную)
- [ ] Уведомления алертов (Telegram/почта)
- [ ] Проверить визуально дашборды Node в Grafana (после 1-2ч данных)
## Как запустить / проверить
```bash
cd /opt/monitoring && docker compose ps # стек
curl -s "http://127.0.0.1:9090/api/v1/targets?state=any" | python3 -m json.tool # node ×4 up
# Grafana: http://bigbox:3001 (или https://grafana.nixg.ru) → дашборды Node vps01/02/03/bigbox
```
## Ключевые артефакты
- /opt/monitoring/docker-compose.yml, prometheus.yml, alerts.yml
- /opt/monitoring/grafana/dashboards/garage-cluster.json (папка Garage)
- /opt/monitoring/grafana/dashboards/nodes/node-<host>.json (4 шт, папка nodes)
- /opt/monitoring/grafana/dashboards/vinogorod/vinograd-wan.json (папка vinogorod)
- /opt/monitoring/grafana/provisioning/datasources/datasources.yml, dashboards/dashboards.yml
- README.md, EXPERIENCE.md, PLAN.md (в этом каталоге), STATUS.md/TODO.md/WALKTHROUGH.md
## Открытые вопросы
- Нужен ли утилизация для vps01/vps02 с реальной скоростью линка (сейчас хардкод 1Гбит/с)
- Оставлять ли tproxy-tunnel.service (SSH-туннель) — vps03 уже в WG, туннель не нужен
+18
View File
@@ -0,0 +1,18 @@
# TODO — Мониторинг кластера
Формат: | дата | задача | статус | закрыта в |
|---|---|---|---|
| 2026-09-08 | vps03: установить node-exporter | ✅ закрыта | @session:default/20260908_142639_b2d5b4 |
| 2026-09-08 | vps03: подключить к WireGuard (10.8.0.3) | ✅ закрыта | @session:default/20260908_142639_b2d5b4 |
| 2026-09-08 | bigbox/vps02: AllowedIPs пира vps01 += 10.8.0.3/32 (возвратный трафик) | ✅ закрыта | @session:default/20260908_142639_b2d5b4 |
| 2026-09-08 | prometheus.yml: таргет 10.8.0.3:9100 (job node ×4) | ✅ закрыта | @session:default/20260908_142639_b2d5b4 |
| 2026-09-08 | Дашборд per-node (Node vps01/vps02/vps03/bigbox, uid node-*) | ✅ закрыта | @session:default/20260908_142639_b2d5b4 |
| 2026-09-08 | Сеть по интерфейсам: RX/TX + утилизация канала (node_network_*, speed) | ✅ закрыта | @session:default/20260908_142639_b2d5b4 |
| 2026-09-08 | EXPERIENCE.md/README.md: грабли (WG, node_network_*, speed=-125000) | ✅ закрыта | @session:default/20260908_142639_b2d5b4 |
| 2026-09-08 | Дашборды по подпапкам (nodes/, vinogorod/) + provisioning folder | ✅ закрыта | @session:default/20260908_180858_730d1b |
| 2026-09-08 | Git: push в gitverse + mirror gitea | ✅ закрыта | @session:default/20260908_142639_b2d5b4 |
| — | admin_token / metrics_token в Garage — разные | 🔵 открыта | |
| — | Пароль Grafana (ручной, пользователь) | 🔵 открыта | |
| — | Уведомления алертов (Telegram/почта) | 🔵 открыта | |
| — | Утилизация vps01/vps02 с реальной скоростью (хардкод 1Гбит) | 🔵 открыта | |
| — | Удалить tproxy-tunnel.service (vps03 в WG, туннель не нужен) | 🔵 открыта | |
+95
View File
@@ -0,0 +1,95 @@
# WALKTHROUGH — Мониторинг кластера (капитанский журнал)
Хронология. Цель — воспроизводимость с нуля. Подробные грабли — в EXPERIENCE.md.
## 2026-09-08 — vps03 в WG + per-node дашборды (сессия @session:default/20260908_142639_b2d5b4)
### Задача
Мониторинг vps03 (77.67.89.154, Debian 13) без публикации порта наружу;
на каждую ноду отдельный дашборд с сетью по интерфейсам (RX/TX + утилизация).
### Шаг 1. WireGuard vps03
1. Сгенерировать ключ vps03 на bigbox: `wg genkey | tee priv | wg pubkey | tee pub`
(получили `HBTzrS86SZ+...`).
2. На vps03: `apt install wireguard-tools`, `/etc/wireguard/wg0.conf`:
```
[Interface]
Address = 10.8.0.3/24
PrivateKey = <priv vps03>
[Peer]
PublicKey = <pubkey hub vps01: ZAvz4xCE...>
Endpoint = 5.129.217.146:51820
AllowedIPs = 10.8.0.0/24
```
`systemctl enable --now wg-quick@wg0`.
⚠️ Грабль: `printf "PrivateKey = ..."` в heredoc съел пробел (`PrivateKey=...`) →
wg-quick падал. Исправлять прямой перезаписью файла.
3. На vps01 (hub): `wg set wg0 peer <pub vps03> allowed-ips 10.8.0.3/32`
+ добавить блок [Peer] в /etc/wireguard/wg0.conf (навсегда, cp .bak сначала).
4. На bigbox/vps02: тот же `wg set ... allowed-ips 10.8.0.3/32` + в файл.
5. **Грабль-звезда:** ping vps03→vps01 работает, а vps03→bigbox/vps02 — нет.
Причина: WireGuard дропает пакеты, чей src НЕ в AllowedIPs пира. У bigbox/vps02
в AllowedIPs пира vps01 было только `10.8.0.1/32`. Решение: расширить до
`10.8.0.1/32, 10.8.0.3/32` (runtime `wg set` + файл). После этого ping OK
(vps03→bigbox 113ms, vps03→vps02 86ms).
### Шаг 2. node-exporter vps03
1. `apt install prometheus-node-exporter`, `/etc/default/prometheus-node-exporter`:
`ARGS="--web.listen-address=10.8.0.3:9100"` (не 127.0.0.1 — иначе по WG не достать;
не `*:` — наружу светится).
2. Проверка: `ss -tlnp | grep 9100` → 10.8.0.3:9100; `curl http://77.67.89.154:9100` → refused.
3. prometheus.yml: в job `node` добавлен таргет `10.8.0.3:9100` (relabel host=vps03).
`docker compose restart prometheus`, проверить `api/v1/targets` → node ×4 up.
### Шаг 3. Per-node дашборды
1. Обнаружение: в node-exporter 1.9 (apt/Debian 13) метрики сети называются
`node_network_receive_bytes_total{device=...}` / `node_network_transmit_bytes_total`,
скорость линка — `node_network_speed_bytes` (Б/с).
2. Генератор `/tmp/gen_nodes_dash.py` (Python) → `grafana/dashboards/node-<host>.json`
(uid `node-<host>`, title "Node <host>"): панели Availability (stat up),
Disk free GB (`node_filesystem_avail_bytes{fstype!~"tmpfs|overlay|squashfs"}`),
Memory (`node_memory_MemAvailable_bytes`), Load (load1/5/15),
Net <iface> RX/TX (`rate(...[5m])`), Net <iface> utilization
`(rx+tx)*8/speed*100`.
3. Физические интерфейсы: vps01=eth0, vps02=enp3s0, vps03=ens1, bigbox=eno1.
Контейнерные (docker0/veth*/br-*) исключены.
4. **Грабль скорости:** на виртуалках vps01/vps02 `node_network_speed_bytes` = **-125000**
(sysfs speed=-1 → exporter умножает на 125000). На физических (vps03/bigbox) = 1.25e+08.
Решение: для vps01/vps02 в формуле утилизации константа `125000000` Б/с (1 Гбит),
для vps03/bigbox — сама метрика.
5. Provisioning подхватывает дашборды автоматом (~30с, перечитывает папку).
Проверка: `docker cp grafana:/var/lib/grafana/grafana.db /tmp/gf.db`,
`SELECT uid,title FROM dashboard WHERE uid LIKE 'node-%'` → 4 строки.
6. Общий nodes.json удалён — заменён на 4 per-node дашборда.
### Шаг 4. Git
- `git add -A && git commit` (2 коммита: WG+дашборды; EXPERIENCE-грабли)
- `git push origin master` → gitverse (6e48fd8, eaafb1f)
- gitea mirror-sync: `curl -X POST -H "Authorization: token <GITEA_TOK>"
http://127.0.0.1:3000/api/v1/repos/estorozhenko/monitoring/mirror-sync` → 200
### Итог проверки
- `api/v1/targets?state=any` → node ×4 up (10.8.0.1/2/3/4)
- grafana.nixg.ru/api/health → 200
- Дашборды node-vps01/02/03/bigbox — в БД Графаны
## 2026-09-08 (вечер) — подпапки дашбордов (сессия @session:default/20260908_180858_730d1b)
Пользователь создал папки в UI Grafana (nodes, vinogorod) и попросил разложить
дашборды по ним (в файловой системе + в интерфейсе).
1. Перемещение на диске (git mv, история сохранена):
```
grafana/dashboards/nodes/node-{bigbox,vps01,vps02,vps03}.json
grafana/dashboards/vinogorod/vinograd-wan.json
garage-cluster.json — остался в корне
```
2. `grafana/provisioning/dashboards/dashboards.yml` — 3 провайдера:
- garage-dashboards → path /var/lib/grafana/dashboards (корень), folder: Garage
- node-dashboards → path .../dashboards/nodes, folder: nodes
- vinogorod-dashboards → path .../dashboards/vinogorod, folder: vinogorod
(Grafana type:file НЕ ходит рекурсивно — поэтому на подпапку свой provider.)
3. `docker compose restart grafana` — provisioning применил папки.
4. Проверка по БД: Node * → nodes, Vinograd WAN → vinogorod, Garage Cluster → Garage.
5. Коммит 4b0539d → gitverse + gitea mirror-sync.
+75 -9
View File
@@ -7,28 +7,94 @@ groups:
labels: labels:
severity: critical severity: critical
annotations: annotations:
summary: "Garage node {{ $labels.instance }} is down" summary: Garage node {{ $labels.instance }} is down
- alert: GarageNoQuorum - alert: GarageNoQuorum
expr: (count(up{job="garage"} == 1) < 2) expr: (count(up{job="garage"} == 1) < 2)
for: 2m for: 2m
labels: labels:
severity: critical severity: critical
annotations: annotations:
summary: "Garage cluster lost quorum (<2 nodes up)" summary: Garage cluster lost quorum (<2 nodes up)
- alert: GarageResyncErrors - alert: GarageResyncErrors
expr: garage_block_resync_error_count > 0 expr: block_resync_errored_blocks > 0
for: 10m for: 10m
labels: labels:
severity: warning severity: warning
annotations: annotations:
summary: "Garage resync errors on {{ $labels.instance }}" summary: Garage resync errors on {{ $labels.instance }}
- alert: GarageNodeUnstable - alert: GarageNodeUnstable
expr: garage_rpc_node_health_is_up == 0 expr: cluster_layout_node_connected == 0
for: 5m for: 5m
labels: labels:
severity: warning severity: warning
annotations: annotations:
summary: "Garage RPC health of some node is down" summary: Garage RPC health of some node is down
- name: tproxy
rules:
- alert: TProxyDown
expr: up{job="tproxy"} == 0
for: 2m
labels:
severity: critical
annotations:
summary: tproxy-server (vps03) is down or unreachable
- alert: TProxyBackendErrors
expr: increase(tproxy_backend_dial_failures_total[5m]) > 0
for: 10m
labels:
severity: warning
annotations:
summary: tproxy-server backend dial failures (MTProxy unreachable)
- name: vinograd
rules:
- alert: VinogradRostelecomDown
expr: probe_success{job="vinograd_wan"} == 0
for: 2m
labels:
severity: critical
annotations:
summary: Vinograd WAN (Ростелеком) {{ $labels.instance }} is down or unreachable
- name: vesti
rules:
- alert: VestiWebDown
expr: vesti_web_up{component="vesti"} == 0
for: 2m
labels:
severity: critical
annotations:
summary: VESTI web (:8400, systemd) is down
- alert: VestiPublisherDown
expr: vesti_publisher_health{component="vesti"} == 0
for: 2m
labels:
severity: critical
annotations:
summary: VESTI publisher (:8410) health check failed
- alert: VestiPublisherContainerDown
expr: vesti_publisher_docker{component="vesti"} == 0
for: 2m
labels:
severity: critical
annotations:
summary: VESTI publisher Docker container is not Up
- alert: VestiTunnelDown
expr: vesti_telegram_tunnel{component="vesti"} == 0
for: 2m
labels:
severity: critical
annotations:
summary: VESTI Telegram SOCKS5 tunnel (:1080) is down
- alert: VestiDbDown
expr: vesti_db_ok{component="vesti"} == 0
for: 2m
labels:
severity: critical
annotations:
summary: VESTI SQLite database vesti.db is unavailable
- alert: VestiDbStale
expr: time() - vesti_db_last_run_ts{component="vesti"} > 1800
for: 5m
labels:
severity: warning
annotations:
summary: VESTI crawler did not run for more than 30 minutes
+3
View File
@@ -5,3 +5,6 @@ modules:
http: http:
valid_status_codes: [200] valid_status_codes: [200]
follow_redirects: true follow_redirects: true
icmp:
prober: icmp
timeout: 5s
+1
View File
@@ -9,6 +9,7 @@ services:
restart: unless-stopped restart: unless-stopped
volumes: volumes:
- ./prometheus.yml:/etc/prometheus/prometheus.yml:ro - ./prometheus.yml:/etc/prometheus/prometheus.yml:ro
- ./alerts.yml:/etc/prometheus/alerts.yml:ro
- ./prometheus-data:/prometheus - ./prometheus-data:/prometheus
command: command:
- '--config.file=/etc/prometheus/prometheus.yml' - '--config.file=/etc/prometheus/prometheus.yml'
+188
View File
@@ -0,0 +1,188 @@
#!/usr/bin/env python3
# Генератор дашборда VESTI для Grafana (provisioning).
# Соглашения те же, что в node-<host>.json: datasource uid=Prometheus,
# stat-панели с mapping 0/1, timeseries.
import json
DS = {"type": "prometheus", "uid": "Prometheus"}
# --- утилиты ---
def stat_panel(title, expr, mapping, idx, w=3, h=3, x=0, y=0, legend=""):
return {
"datasource": DS,
"fieldConfig": {
"defaults": {
"color": {"mode": "thresholds"},
"mappings": ([{"options": mapping, "type": "value"}] if mapping else []),
"thresholds": {
"mode": "absolute",
"steps": [{"color": "red", "value": None}, {"color": "green", "value": 1}],
},
"unit": "short",
},
"overrides": [],
},
"gridPos": {"h": h, "w": w, "x": x, "y": y},
"id": idx,
"options": {
"colorMode": "background",
"graphMode": "none",
"justifyMode": "auto",
"orientation": "auto",
"reduceOptions": {"calcs": ["lastNotNull"], "fields": "", "values": False},
"textMode": "auto",
},
"pluginVersion": "11.1.0",
"targets": [{"datasource": DS, "expr": expr, "legendFormat": legend, "refId": "A"}],
"title": title,
"type": "stat",
}
def timeseries_panel(title, exprs, idx, w=12, h=8, x=0, y=0, unit="short"):
targets = []
for e, l in exprs:
targets.append({"datasource": DS, "expr": e, "legendFormat": l, "refId": chr(65 + len(targets))})
return {
"datasource": DS,
"fieldConfig": {
"defaults": {
"color": {"mode": "palette-classic"},
"custom": {
"axisCenteredZero": False, "axisColorMode": "text", "axisLabel": "",
"axisPlacement": "auto", "drawStyle": "line", "fillOpacity": 10,
"gradientMode": "none", "hideFrom": {"legend": False, "tooltip": False, "viz": False},
"lineInterpolation": "linear", "lineWidth": 1, "pointSize": 5,
"scaleDistribution": {"type": "linear"}, "showPoints": "never",
"spanNulls": False, "stacking": {"group": "A", "mode": "none"},
"thresholdsStyle": {"mode": "off"},
},
"mappings": [],
"thresholds": {"mode": "absolute", "steps": [{"color": "green", "value": None}]},
"unit": unit,
},
"overrides": [],
},
"gridPos": {"h": h, "w": w, "x": x, "y": y},
"id": idx,
"options": {
"legend": {"calcs": [], "displayMode": "list", "placement": "bottom", "showLegend": True},
"tooltip": {"mode": "multi", "sort": "none"},
},
"targets": targets,
"title": title,
"type": "timeseries",
}
panels = []
y = 0
idx = 1
# Row 1: Доступность (stat, 3x4 = 12 панелей в 2 ряда по 6)
row_title = {
"collapsed": False,
"datasource": DS,
"gridPos": {"h": 1, "w": 24, "x": 0, "y": y},
"id": idx, "panels": [], "title": "Доступность компонентов", "type": "row",
}
idx += 1
y += 1
panels.append(row_title)
stats = [
("Web :8400 (systemd)", "vesti_web_up{component=\"vesti\"}", {"0": {"color": "red", "text": "DOWN"}, "1": {"color": "green", "text": "UP"}}),
("Publisher :8410 (healthz)", "vesti_publisher_health{component=\"vesti\"}", {"0": {"color": "red", "text": "DOWN"}, "1": {"color": "green", "text": "UP"}}),
("Publisher (Docker)", "vesti_publisher_docker{component=\"vesti\"}", {"0": {"color": "red", "text": "DOWN"}, "1": {"color": "green", "text": "UP"}}),
("Telegram tunnel :1080", "vesti_telegram_tunnel{component=\"vesti\"}", {"0": {"color": "red", "text": "DOWN"}, "1": {"color": "green", "text": "UP"}}),
("DB vesti.db", "vesti_db_ok{component=\"vesti\"}", {"0": {"color": "red", "text": "DOWN"}, "1": {"color": "green", "text": "UP"}}),
("Web systemd unit", "vesti_web_systemd{component=\"vesti\"}", {"0": {"color": "red", "text": "DOWN"}, "1": {"color": "green", "text": "UP"}}),
]
x = 0
for i, (title, expr, mp) in enumerate(stats):
panels.append(stat_panel(title, expr, mp, idx, w=4, h=3, x=(i % 6) * 4, y=y + (i // 6) * 3))
idx += 1
y += 6 # 2 ряда по 3 = 6 строк
# Row 2: Publisher healthz детали (bot/proxy/channels)
row2 = {
"collapsed": False, "datasource": DS,
"gridPos": {"h": 1, "w": 24, "x": 0, "y": y},
"id": idx, "panels": [], "title": "Publisher — детали", "type": "row",
}
idx += 1
y += 1
panels.append(row2)
for i, (title, expr) in enumerate([
("Bot identity", 'vesti_publisher_bot{component="vesti"}'),
("Proxy SOCKS5", 'vesti_publisher_proxy{component="vesti"}'),
("Channels", 'vesti_publisher_channels{component="vesti"}'),
]):
panels.append(stat_panel(title, expr, None, idx, w=4, h=3, x=i * 4, y=y))
idx += 1
y += 3
# Row 3: БД
row3 = {
"collapsed": False, "datasource": DS,
"gridPos": {"h": 1, "w": 24, "x": 0, "y": y},
"id": idx, "panels": [], "title": "База данных VESTI", "type": "row",
}
idx += 1
y += 1
panels.append(row3)
# stat: постов в БД, новых, опубликовано, последний рan
panels.append(stat_panel("Постов в БД", 'vesti_db_posts_total{component="vesti"}', None, idx, w=4, h=3, x=0, y=y))
idx += 1
panels.append(stat_panel("Новых", 'vesti_db_new_total{component="vesti"}', None, idx, w=4, h=3, x=4, y=y))
idx += 1
panels.append(stat_panel("Опубликовано", 'vesti_db_published_total{component="vesti"}', None, idx, w=4, h=3, x=8, y=y))
idx += 1
# временной ряд последнего рan (ts → время)
panels.append(timeseries_panel("Последний рan краулера", [('time() - vesti_db_last_run_ts{component="vesti"}', "сек назад")], idx, w=12, h=4, x=12, y=y, unit="s"))
idx += 1
y += 4
# Row 4: Временные ряды доступности (возврат по вре)
row4 = {
"collapsed": False, "datasource": DS,
"gridPos": {"h": 1, "w": 24, "x": 0, "y": y},
"id": idx, "panels": [], "title": "История доступности (1/0)", "type": "row",
}
idx += 1
y += 1
panels.append(row4)
panels.append(timeseries_panel(
"Доступность компонентов",
[(f'{m}{{component="vesti"}}', t) for m, t in [
("vesti_web_up", "web"),
("vesti_publisher_health", "publisher"),
("vesti_publisher_docker", "docker"),
("vesti_telegram_tunnel", "tunnel"),
("vesti_db_ok", "db"),
("vesti_web_systemd", "web_sys"),
]],
idx, w=24, h=8, x=0, y=y, unit="short"))
idx += 1
y += 8
dashboard = {
"annotations": {"list": []},
"editable": True,
"fiscalYearStartMonth": 0,
"graphTooltip": 1,
"id": None,
"links": [],
"panels": panels,
"refresh": "15s",
"schemaVersion": 1,
"tags": ["vesti", "bigbox"],
"time": {"from": "now-6h", "to": "now"},
"timepicker": [{"refresh": "15s"}],
"title": "VESTI",
"uid": "vesti",
"version": 1,
}
with open("/opt/monitoring/grafana/dashboards/vesti/vesti.json", "w") as f:
json.dump(dashboard, f, indent=2, ensure_ascii=False)
print("written /opt/monitoring/grafana/dashboards/vesti/vesti.json", len(panels), "pabels")
+598 -45
View File
@@ -9,136 +9,689 @@
"links": [], "links": [],
"panels": [ "panels": [
{ {
"datasource": { "type": "prometheus", "uid": "Prometheus" }, "datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": { "fieldConfig": {
"defaults": { "defaults": {
"color": { "mode": "thresholds" }, "color": {
"mode": "thresholds"
},
"mappings": [ "mappings": [
{ "options": { "0": { "color": "red", "text": "DOWN" }, "1": { "color": "green", "text": "UP" } }, "type": "value" } {
"options": {
"0": {
"color": "red",
"text": "DOWN"
},
"1": {
"color": "green",
"text": "UP"
}
},
"type": "value"
}
], ],
"thresholds": { "mode": "absolute", "steps": [ { "color": "red", "value": null }, { "color": "green", "value": 1 } ] } "thresholds": {
"mode": "absolute",
"steps": [
{
"color": "red",
"value": null
},
{
"color": "green",
"value": 1
}
]
}
}, },
"overrides": [] "overrides": []
}, },
"gridPos": { "h": 8, "w": 6, "x": 0, "y": 0 }, "gridPos": {
"h": 8,
"w": 6,
"x": 0,
"y": 0
},
"id": 1, "id": 1,
"options": { "options": {
"colorMode": "background", "colorMode": "background",
"graphMode": "none", "graphMode": "none",
"justifyMode": "auto", "justifyMode": "auto",
"orientation": "auto", "orientation": "auto",
"reduceOptions": { "calcs": [ "lastNotNull" ], "fields": "", "values": false }, "reduceOptions": {
"calcs": [
"lastNotNull"
],
"fields": "",
"values": false
},
"textMode": "auto" "textMode": "auto"
}, },
"pluginVersion": "11.1.0", "pluginVersion": "11.1.0",
"targets": [ "targets": [
{ "datasource": { "type": "prometheus", "uid": "Prometheus" }, "expr": "up{job=\"garage\"}", "legendFormat": "{{ instance }}", "refId": "A" } {
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "up{job=\"garage\"}",
"legendFormat": "{{ instance }}",
"refId": "A"
}
], ],
"title": "Garage nodes up", "title": "Garage nodes up",
"type": "stat" "type": "stat"
}, },
{ {
"datasource": { "type": "prometheus", "uid": "Prometheus" }, "datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": { "fieldConfig": {
"defaults": { "defaults": {
"color": { "mode": "palette-classic" }, "color": {
"custom": { "axisCenteredZero": false, "axisColorMode": "text", "axisLabel": "", "axisPlacement": "auto", "drawStyle": "line", "fillOpacity": 10, "gradientMode": "none", "hideFrom": { "legend": false, "tooltip": false, "viz": false }, "lineInterpolation": "linear", "lineWidth": 1, "pointSize": 5, "scaleDistribution": { "type": "linear" }, "showPoints": "never", "spanNulls": false, "stacking": { "group": "A", "mode": "none" }, "thresholdsStyle": { "mode": "off" } }, "mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [], "mappings": [],
"thresholds": { "mode": "absolute", "steps": [ { "color": "green", "value": null } ] } "thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
}, },
"overrides": [] "overrides": []
}, },
"gridPos": { "h": 8, "w": 6, "x": 6, "y": 0 }, "gridPos": {
"h": 8,
"w": 6,
"x": 6,
"y": 0
},
"id": 2, "id": 2,
"options": { "options": {
"legend": { "calcs": [], "displayMode": "list", "placement": "bottom", "showLegend": true }, "legend": {
"tooltip": { "mode": "multi", "sort": "none" } "calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
}, },
"targets": [ "targets": [
{ "datasource": { "type": "prometheus", "uid": "Prometheus" }, "expr": "node_filesystem_avail_bytes{fstype!~\"tmpfs|overlay|squashfs\"} / 1024 / 1024 / 1024", "legendFormat": "{{ instance }}", "refId": "A" } {
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "node_filesystem_avail_bytes{fstype!~\"tmpfs|overlay|squashfs\"} / 1024 / 1024 / 1024",
"legendFormat": "{{ instance }}",
"refId": "A"
}
], ],
"title": "Node disk free (GB)", "title": "Node disk free (GB)",
"type": "timeseries" "type": "timeseries"
}, },
{ {
"datasource": { "type": "prometheus", "uid": "Prometheus" }, "datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": { "fieldConfig": {
"defaults": { "defaults": {
"color": { "mode": "palette-classic" }, "color": {
"custom": { "axisCenteredZero": false, "axisColorMode": "text", "axisLabel": "", "axisPlacement": "auto", "drawStyle": "line", "fillOpacity": 10, "gradientMode": "none", "hideFrom": { "legend": false, "tooltip": false, "viz": false }, "lineInterpolation": "linear", "lineWidth": 1, "pointSize": 5, "scaleDistribution": { "type": "linear" }, "showPoints": "never", "spanNulls": false, "stacking": { "group": "A", "mode": "none" }, "thresholdsStyle": { "mode": "off" } }, "mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [], "mappings": [],
"thresholds": { "mode": "absolute", "steps": [ { "color": "green", "value": null } ] } "thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
}, },
"overrides": [] "overrides": []
}, },
"gridPos": { "h": 8, "w": 12, "x": 12, "y": 0 }, "gridPos": {
"h": 8,
"w": 12,
"x": 12,
"y": 0
},
"id": 3, "id": 3,
"options": { "options": {
"legend": { "calcs": [], "displayMode": "list", "placement": "bottom", "showLegend": true }, "legend": {
"tooltip": { "mode": "multi", "sort": "none" } "calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
}, },
"targets": [ "targets": [
{ "datasource": { "type": "prometheus", "uid": "Prometheus" }, "expr": "garage_block_count", "legendFormat": "blocks {{ instance }}", "refId": "A" }, {
{ "datasource": { "type": "prometheus", "uid": "Prometheus" }, "expr": "garage_block_resync_queue_length", "legendFormat": "resync queue {{ instance }}", "refId": "B" } "expr": "block_resync_queue_length",
"legendFormat": "resync queue {{ instance }}",
"refId": "A"
},
{
"expr": "block_resync_errored_blocks",
"legendFormat": "errored {{ instance }}",
"refId": "B"
}
], ],
"title": "Garage blocks", "title": "Garage blocks (resync)",
"type": "timeseries" "type": "timeseries"
}, },
{ {
"datasource": { "type": "prometheus", "uid": "Prometheus" }, "datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": { "fieldConfig": {
"defaults": { "defaults": {
"color": { "mode": "thresholds" }, "color": {
"mode": "thresholds"
},
"mappings": [], "mappings": [],
"thresholds": { "mode": "absolute", "steps": [ { "color": "green", "value": null }, { "color": "red", "value": 1 } ] } "thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
},
{
"color": "red",
"value": 1
}
]
}
}, },
"overrides": [] "overrides": []
}, },
"gridPos": { "h": 8, "w": 12, "x": 0, "y": 8 }, "gridPos": {
"h": 8,
"w": 12,
"x": 0,
"y": 8
},
"id": 4, "id": 4,
"options": { "options": {
"legend": { "calcs": [], "displayMode": "list", "placement": "bottom", "showLegend": true }, "legend": {
"tooltip": { "mode": "multi", "sort": "none" } "calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
}, },
"targets": [ "targets": [
{ "datasource": { "type": "prometheus", "uid": "Prometheus" }, "expr": "garage_rpc_node_health_is_up", "legendFormat": "{{ instance }}", "refId": "A" } {
"expr": "cluster_layout_node_connected",
"legendFormat": "{{ role_zone }} {{ instance }}",
"refId": "A"
}
], ],
"title": "Garage RPC node health", "title": "Garage node health (layout)",
"type": "timeseries" "type": "timeseries"
}, },
{ {
"datasource": { "type": "prometheus", "uid": "Prometheus" }, "datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": { "fieldConfig": {
"defaults": { "defaults": {
"color": { "mode": "palette-classic" }, "color": {
"custom": { "axisCenteredZero": false, "axisColorMode": "text", "axisLabel": "", "axisPlacement": "auto", "drawStyle": "line", "fillOpacity": 10, "gradientMode": "none", "hideFrom": { "legend": false, "tooltip": false, "viz": false }, "lineInterpolation": "linear", "lineWidth": 1, "pointSize": 5, "scaleDistribution": { "type": "linear" }, "showPoints": "never", "spanNulls": false, "stacking": { "group": "A", "mode": "none" }, "thresholdsStyle": { "mode": "off" } }, "mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [], "mappings": [],
"thresholds": { "mode": "absolute", "steps": [ { "color": "green", "value": null } ] } "thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
}, },
"overrides": [] "overrides": []
}, },
"gridPos": { "h": 8, "w": 12, "x": 12, "y": 8 }, "gridPos": {
"h": 8,
"w": 12,
"x": 12,
"y": 8
},
"id": 5, "id": 5,
"options": { "options": {
"legend": { "calcs": [], "displayMode": "list", "placement": "bottom", "showLegend": true }, "legend": {
"tooltip": { "mode": "multi", "sort": "none" } "calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
}, },
"targets": [ "targets": [
{ "datasource": { "type": "prometheus", "uid": "Prometheus" }, "expr": "rate(garage_api_s3_request_counter[5m])", "legendFormat": "{{ instance }} {{ api_endpoint }}", "refId": "A" } {
"expr": "rate(api_s3_request_counter[5m])",
"legendFormat": "{{ instance }} {{ api_endpoint }}",
"refId": "A"
}
], ],
"title": "S3 requests /sec", "title": "S3 requests /sec",
"type": "timeseries" "type": "timeseries"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 6,
"x": 0,
"y": 16
},
"id": 6,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"title": "tproxy live sessions/streams",
"type": "timeseries",
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "tproxy_sessions_live",
"legendFormat": "sessions {{ instance }}",
"refId": "A"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "tproxy_streams_live",
"legendFormat": "streams {{ instance }}",
"refId": "B"
}
]
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 6,
"x": 6,
"y": 16
},
"id": 7,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"title": "tproxy traffic /sec",
"type": "timeseries",
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "rate(tproxy_bytes_up_total[5m])",
"legendFormat": "up {{ instance }}",
"refId": "A"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "rate(tproxy_bytes_down_total[5m])",
"legendFormat": "down {{ instance }}",
"refId": "B"
}
]
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "thresholds"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 12,
"x": 12,
"y": 16
},
"id": 8,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"title": "tproxy backend errors",
"type": "timeseries",
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "tproxy_backend_dial_failures_total",
"legendFormat": "dial failures {{ instance }}",
"refId": "A"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "tproxy_streams_rejected_total",
"legendFormat": "rejected {{ instance }}",
"refId": "B"
}
]
} }
], ],
"refresh": "30s", "refresh": "30s",
"schemaVersion": 39, "schemaVersion": 39,
"tags": [ "garage" ], "tags": [
"templating": { "list": [] }, "garage"
"time": { "from": "now-6h", "to": "now" }, ],
"templating": {
"list": []
},
"time": {
"from": "now-6h",
"to": "now"
},
"timepicker": {}, "timepicker": {},
"timezone": "browser", "timezone": "browser",
"title": "Garage Cluster", "title": "Garage Cluster",
"uid": "garage-cluster", "uid": "garage-cluster",
"version": 1, "version": 2,
"weekStart": "" "weekStart": ""
} }
File diff suppressed because it is too large Load Diff
+558
View File
@@ -0,0 +1,558 @@
{
"annotations": {
"list": []
},
"editable": true,
"fiscalYearStartMonth": 0,
"graphTooltip": 1,
"id": null,
"links": [],
"panels": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "thresholds"
},
"mappings": [
{
"options": {
"0": {
"color": "red",
"text": "DOWN"
},
"1": {
"color": "green",
"text": "UP"
}
},
"type": "value"
}
],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "red",
"value": null
},
{
"color": "green",
"value": 1
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 6,
"x": 0,
"y": 0
},
"id": 1,
"options": {
"colorMode": "background",
"graphMode": "none",
"justifyMode": "auto",
"orientation": "auto",
"reduceOptions": {
"calcs": [
"lastNotNull"
],
"fields": "",
"values": false
},
"textMode": "auto"
},
"pluginVersion": "11.1.0",
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "up{job=\"node\",host=\"bigbox\"}",
"legendFormat": "{{ host }}",
"refId": "A"
}
],
"title": "Availability",
"type": "stat"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 12,
"x": 0,
"y": 8
},
"id": 2,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "node_filesystem_avail_bytes{fstype!~\"tmpfs|overlay|squashfs\",host=\"bigbox\"} / 1024 / 1024 / 1024",
"legendFormat": "{{ mountpoint }}",
"refId": "A"
}
],
"title": "Disk free (GB)",
"type": "timeseries"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 12,
"x": 12,
"y": 8
},
"id": 3,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "node_memory_MemAvailable_bytes{host=\"bigbox\"} / 1024 / 1024 / 1024",
"legendFormat": "bigbox",
"refId": "A"
}
],
"title": "Memory available (GB)",
"type": "timeseries"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 12,
"x": 12,
"y": 8
},
"id": 4,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "node_load1{host=\"bigbox\"}",
"legendFormat": "1m",
"refId": "A"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "node_load5{host=\"bigbox\"}",
"legendFormat": "5m",
"refId": "B"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "node_load15{host=\"bigbox\"}",
"legendFormat": "15m",
"refId": "C"
}
],
"title": "Load (1m/5m/15m)",
"type": "timeseries"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "B/s",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 12,
"x": 0,
"y": 24
},
"id": 5,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "rate(node_network_receive_bytes_total{device=\"eno1\",host=\"bigbox\"}[5m])",
"legendFormat": "RX",
"refId": "A"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "rate(node_network_transmit_bytes_total{device=\"eno1\",host=\"bigbox\"}[5m])",
"legendFormat": "TX",
"refId": "B"
}
],
"title": "Net eno1 RX/TX (bytes/s)",
"type": "timeseries"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "%",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 12,
"x": 12,
"y": 24
},
"id": 6,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "(rate(node_network_receive_bytes_total{device=\"eno1\",host=\"bigbox\"}[5m]) + rate(node_network_transmit_bytes_total{device=\"eno1\",host=\"bigbox\"}[5m])) * 8 / (node_network_speed_bytes{device=\"eno1\",host=\"bigbox\"}) * 100",
"legendFormat": "util %",
"refId": "A"
}
],
"title": "Net eno1 utilization (%)",
"type": "timeseries"
}
],
"refresh": "15s",
"schemaVersion": 1,
"tags": [
"bigbox"
],
"time": {
"from": "now-6h",
"to": "now"
},
"timepicker": [
{
"refresh": "15s"
}
],
"title": "Node bigbox",
"uid": "node-bigbox",
"version": 1
}
+558
View File
@@ -0,0 +1,558 @@
{
"annotations": {
"list": []
},
"editable": true,
"fiscalYearStartMonth": 0,
"graphTooltip": 1,
"id": null,
"links": [],
"panels": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "thresholds"
},
"mappings": [
{
"options": {
"0": {
"color": "red",
"text": "DOWN"
},
"1": {
"color": "green",
"text": "UP"
}
},
"type": "value"
}
],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "red",
"value": null
},
{
"color": "green",
"value": 1
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 6,
"x": 0,
"y": 0
},
"id": 1,
"options": {
"colorMode": "background",
"graphMode": "none",
"justifyMode": "auto",
"orientation": "auto",
"reduceOptions": {
"calcs": [
"lastNotNull"
],
"fields": "",
"values": false
},
"textMode": "auto"
},
"pluginVersion": "11.1.0",
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "up{job=\"node\",host=\"vps01\"}",
"legendFormat": "{{ host }}",
"refId": "A"
}
],
"title": "Availability",
"type": "stat"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 12,
"x": 0,
"y": 8
},
"id": 2,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "node_filesystem_avail_bytes{fstype!~\"tmpfs|overlay|squashfs\",host=\"vps01\"} / 1024 / 1024 / 1024",
"legendFormat": "{{ mountpoint }}",
"refId": "A"
}
],
"title": "Disk free (GB)",
"type": "timeseries"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 12,
"x": 12,
"y": 8
},
"id": 3,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "node_memory_MemAvailable_bytes{host=\"vps01\"} / 1024 / 1024 / 1024",
"legendFormat": "vps01",
"refId": "A"
}
],
"title": "Memory available (GB)",
"type": "timeseries"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 12,
"x": 12,
"y": 8
},
"id": 4,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "node_load1{host=\"vps01\"}",
"legendFormat": "1m",
"refId": "A"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "node_load5{host=\"vps01\"}",
"legendFormat": "5m",
"refId": "B"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "node_load15{host=\"vps01\"}",
"legendFormat": "15m",
"refId": "C"
}
],
"title": "Load (1m/5m/15m)",
"type": "timeseries"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "B/s",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 12,
"x": 0,
"y": 24
},
"id": 5,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "rate(node_network_receive_bytes_total{device=\"eth0\",host=\"vps01\"}[5m])",
"legendFormat": "RX",
"refId": "A"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "rate(node_network_transmit_bytes_total{device=\"eth0\",host=\"vps01\"}[5m])",
"legendFormat": "TX",
"refId": "B"
}
],
"title": "Net eth0 RX/TX (bytes/s)",
"type": "timeseries"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "%",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 12,
"x": 12,
"y": 24
},
"id": 6,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "(rate(node_network_receive_bytes_total{device=\"eth0\",host=\"vps01\"}[5m]) + rate(node_network_transmit_bytes_total{device=\"eth0\",host=\"vps01\"}[5m])) * 8 / (125000000) * 100",
"legendFormat": "util %",
"refId": "A"
}
],
"title": "Net eth0 utilization (%)",
"type": "timeseries"
}
],
"refresh": "15s",
"schemaVersion": 1,
"tags": [
"vps01"
],
"time": {
"from": "now-6h",
"to": "now"
},
"timepicker": [
{
"refresh": "15s"
}
],
"title": "Node vps01",
"uid": "node-vps01",
"version": 1
}
+558
View File
@@ -0,0 +1,558 @@
{
"annotations": {
"list": []
},
"editable": true,
"fiscalYearStartMonth": 0,
"graphTooltip": 1,
"id": null,
"links": [],
"panels": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "thresholds"
},
"mappings": [
{
"options": {
"0": {
"color": "red",
"text": "DOWN"
},
"1": {
"color": "green",
"text": "UP"
}
},
"type": "value"
}
],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "red",
"value": null
},
{
"color": "green",
"value": 1
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 6,
"x": 0,
"y": 0
},
"id": 1,
"options": {
"colorMode": "background",
"graphMode": "none",
"justifyMode": "auto",
"orientation": "auto",
"reduceOptions": {
"calcs": [
"lastNotNull"
],
"fields": "",
"values": false
},
"textMode": "auto"
},
"pluginVersion": "11.1.0",
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "up{job=\"node\",host=\"vps02\"}",
"legendFormat": "{{ host }}",
"refId": "A"
}
],
"title": "Availability",
"type": "stat"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 12,
"x": 0,
"y": 8
},
"id": 2,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "node_filesystem_avail_bytes{fstype!~\"tmpfs|overlay|squashfs\",host=\"vps02\"} / 1024 / 1024 / 1024",
"legendFormat": "{{ mountpoint }}",
"refId": "A"
}
],
"title": "Disk free (GB)",
"type": "timeseries"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 12,
"x": 12,
"y": 8
},
"id": 3,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "node_memory_MemAvailable_bytes{host=\"vps02\"} / 1024 / 1024 / 1024",
"legendFormat": "vps02",
"refId": "A"
}
],
"title": "Memory available (GB)",
"type": "timeseries"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 12,
"x": 12,
"y": 8
},
"id": 4,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "node_load1{host=\"vps02\"}",
"legendFormat": "1m",
"refId": "A"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "node_load5{host=\"vps02\"}",
"legendFormat": "5m",
"refId": "B"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "node_load15{host=\"vps02\"}",
"legendFormat": "15m",
"refId": "C"
}
],
"title": "Load (1m/5m/15m)",
"type": "timeseries"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "B/s",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 12,
"x": 0,
"y": 24
},
"id": 5,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "rate(node_network_receive_bytes_total{device=\"enp3s0\",host=\"vps02\"}[5m])",
"legendFormat": "RX",
"refId": "A"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "rate(node_network_transmit_bytes_total{device=\"enp3s0\",host=\"vps02\"}[5m])",
"legendFormat": "TX",
"refId": "B"
}
],
"title": "Net enp3s0 RX/TX (bytes/s)",
"type": "timeseries"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "%",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 12,
"x": 12,
"y": 24
},
"id": 6,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "(rate(node_network_receive_bytes_total{device=\"enp3s0\",host=\"vps02\"}[5m]) + rate(node_network_transmit_bytes_total{device=\"enp3s0\",host=\"vps02\"}[5m])) * 8 / (125000000) * 100",
"legendFormat": "util %",
"refId": "A"
}
],
"title": "Net enp3s0 utilization (%)",
"type": "timeseries"
}
],
"refresh": "15s",
"schemaVersion": 1,
"tags": [
"vps02"
],
"time": {
"from": "now-6h",
"to": "now"
},
"timepicker": [
{
"refresh": "15s"
}
],
"title": "Node vps02",
"uid": "node-vps02",
"version": 1
}
+558
View File
@@ -0,0 +1,558 @@
{
"annotations": {
"list": []
},
"editable": true,
"fiscalYearStartMonth": 0,
"graphTooltip": 1,
"id": null,
"links": [],
"panels": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "thresholds"
},
"mappings": [
{
"options": {
"0": {
"color": "red",
"text": "DOWN"
},
"1": {
"color": "green",
"text": "UP"
}
},
"type": "value"
}
],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "red",
"value": null
},
{
"color": "green",
"value": 1
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 6,
"x": 0,
"y": 0
},
"id": 1,
"options": {
"colorMode": "background",
"graphMode": "none",
"justifyMode": "auto",
"orientation": "auto",
"reduceOptions": {
"calcs": [
"lastNotNull"
],
"fields": "",
"values": false
},
"textMode": "auto"
},
"pluginVersion": "11.1.0",
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "up{job=\"node\",host=\"vps03\"}",
"legendFormat": "{{ host }}",
"refId": "A"
}
],
"title": "Availability",
"type": "stat"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 12,
"x": 0,
"y": 8
},
"id": 2,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "node_filesystem_avail_bytes{fstype!~\"tmpfs|overlay|squashfs\",host=\"vps03\"} / 1024 / 1024 / 1024",
"legendFormat": "{{ mountpoint }}",
"refId": "A"
}
],
"title": "Disk free (GB)",
"type": "timeseries"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 12,
"x": 12,
"y": 8
},
"id": 3,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "node_memory_MemAvailable_bytes{host=\"vps03\"} / 1024 / 1024 / 1024",
"legendFormat": "vps03",
"refId": "A"
}
],
"title": "Memory available (GB)",
"type": "timeseries"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 12,
"x": 12,
"y": 8
},
"id": 4,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "node_load1{host=\"vps03\"}",
"legendFormat": "1m",
"refId": "A"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "node_load5{host=\"vps03\"}",
"legendFormat": "5m",
"refId": "B"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "node_load15{host=\"vps03\"}",
"legendFormat": "15m",
"refId": "C"
}
],
"title": "Load (1m/5m/15m)",
"type": "timeseries"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "B/s",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 12,
"x": 0,
"y": 24
},
"id": 5,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "rate(node_network_receive_bytes_total{device=\"ens1\",host=\"vps03\"}[5m])",
"legendFormat": "RX",
"refId": "A"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "rate(node_network_transmit_bytes_total{device=\"ens1\",host=\"vps03\"}[5m])",
"legendFormat": "TX",
"refId": "B"
}
],
"title": "Net ens1 RX/TX (bytes/s)",
"type": "timeseries"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "%",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 12,
"x": 12,
"y": 24
},
"id": 6,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "(rate(node_network_receive_bytes_total{device=\"ens1\",host=\"vps03\"}[5m]) + rate(node_network_transmit_bytes_total{device=\"ens1\",host=\"vps03\"}[5m])) * 8 / (node_network_speed_bytes{device=\"ens1\",host=\"vps03\"}) * 100",
"legendFormat": "util %",
"refId": "A"
}
],
"title": "Net ens1 utilization (%)",
"type": "timeseries"
}
],
"refresh": "15s",
"schemaVersion": 1,
"tags": [
"vps03"
],
"time": {
"from": "now-6h",
"to": "now"
},
"timepicker": [
{
"refresh": "15s"
}
],
"title": "Node vps03",
"uid": "node-vps03",
"version": 1
}
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,202 @@
{
"annotations": {
"list": []
},
"editable": true,
"fiscalYearStartMonth": 0,
"graphTooltip": 0,
"id": null,
"links": [],
"panels": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 12,
"x": 0,
"y": 0
},
"id": 2,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "probe_success{job=\"vinograd_wan\"}",
"legendFormat": "{{ instance }}",
"refId": "A"
}
],
"title": "Vinograd WAN availability",
"type": "timeseries"
},
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "ms",
"axisPlacement": "auto",
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 5,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"gridPos": {
"h": 8,
"w": 12,
"x": 0,
"y": 8
},
"id": 3,
"options": {
"legend": {
"calcs": [],
"displayMode": "list",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "none"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "Prometheus"
},
"expr": "probe_icmp_duration_seconds{job=\"vinograd_wan\",phase=\"rtt\"} * 1000",
"legendFormat": "{{ instance }}",
"refId": "A"
}
],
"title": "Vinograd WAN RTT (ms)",
"type": "timeseries"
}
],
"refresh": "30s",
"schemaVersion": 39,
"tags": [
"vinograd",
"wan",
"rostelecom"
],
"templating": {
"list": []
},
"time": {
"from": "now-6h",
"to": "now"
},
"timepicker": {},
"timezone": "Europe/Moscow",
"title": "Vinograd WAN",
"uid": "vinograd-wan",
"version": 1,
"weekStart": ""
}
+37 -1
View File
@@ -8,4 +8,40 @@ providers:
disableDeletion: false disableDeletion: false
updateIntervalSeconds: 30 updateIntervalSeconds: 30
options: options:
path: /var/lib/grafana/dashboards path: /var/lib/grafana/dashboards/garage-cluster.json
- name: 'node-dashboards'
orgId: 1
folder: 'nodes'
type: file
disableDeletion: false
updateIntervalSeconds: 30
options:
path: /var/lib/grafana/dashboards/nodes
- name: 'vinogorod-dashboards'
orgId: 1
folder: 'vinogorod'
type: file
disableDeletion: false
updateIntervalSeconds: 30
options:
path: /var/lib/grafana/dashboards/vinogorod
- name: 'vesti-dashboards'
orgId: 1
folder: 'vesti'
type: file
disableDeletion: false
updateIntervalSeconds: 30
options:
path: /var/lib/grafana/dashboards/vesti
- name: 'gotosocial-dashboards'
orgId: 1
folder: 'gotosocial'
type: file
disableDeletion: false
updateIntervalSeconds: 30
options:
path: /var/lib/grafana/dashboards/gotosocial
@@ -4,7 +4,8 @@ datasources:
- name: Prometheus - name: Prometheus
type: prometheus type: prometheus
access: proxy access: proxy
url: http://prometheus:9090 uid: Prometheus
url: http://172.28.0.1:9090
isDefault: true isDefault: true
editable: true editable: true
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-09-08
@@ -0,0 +1,62 @@
# Design: Fix Vinograd WAN dashboard annotations
## Файл
| Файл | Действие |
|---|---|
| `/opt/monitoring/grafana/dashboards/vinograd-wan.json` | убрать `annotations.list` (ссылка на `__grafana__`) → `"list": []` |
## Изменение
Было:
```json
"annotations": {
"list": [
{
"builtIn": 1,
"datasource": {"type": "grafana", "uid": "__grafana__"},
"enable": true,
"hide": true,
"iconColor": "rgba(0, 211, 255, 1)",
"name": "Annotations & Alerts",
"type": "style"
}
]
}
```
Стало:
```json
"annotations": {
"list": []
}
```
## Почему так
- `__grafana__` — встроенный datasource (аннотации/алерты), не существует в
БД этой инсталляции (в `data_source` только Prometheus и Loki).
- Панели дашборда не ссылаются на аннотации; секция добавлена автоматически
при создании JSON (скопирована из шаблона) и бесполезна.
- Рабочий garage-cluster.json имеет `"list": []` — дашборд открывается.
## Применение и проверка
```bash
cd /opt/monitoring
# правка файла (руками или jq)
jq '.annotations.list = []' grafana/dashboards/vinograd-wan.json > /tmp/vw.json && mv /tmp/vw.json grafana/dashboards/vinograd-wan.json
# провайдер перечитает файл за ≤30с (updateIntervalSeconds: 30); рестарт не нужен
sleep 35
# проверка: дашборд без ошибки __grafana__
curl -s -u 'estorozhenko:...' 'http://127.0.0.1:3001/api/dashboards/uid/vinograd-wan' | jq '.dashboard.annotations'
# и главное — открытие страницы без ошибки в браузере
```
## Риски
- Минимальные. Изменение декоративное (удаление неиспользуемой секции).
- Если Grafana всё же нужна встроенная аннотация — она добавится автоматически
в рантайме (built-in annotation не зависит от дашборд-JSON).
@@ -0,0 +1,31 @@
# Proposal: Fix Vinograd WAN dashboard — Datasource __grafana__ not found
## Why
При открытии `https://grafana.nixg.ru/d/vinograd-wan/vinograd-wan` Grafana
показывает ошибку:
```
Failed to retrieve datasource
Datasource __grafana__ was not found
```
Панели дашборда (availability, RTT) ссылаются на Prometheus (`uid: Prometheus`)
и работают. Ошибку вызывает секция `annotations` в JSON дашборда, которая
ссылается на встроенный датасорс `__grafana__` (аннотации/алерты). Такого
датасорса нет в БД Grafana 11 OSS (там только Prometheus и Loki), поэтому
Grafana не может его найти и показывает ошибку при открытии.
## What Changes
- В `/opt/monitoring/grafana/dashboards/vinograd-wan.json` секция
`annotations.list` заменяется с массива с элементом `{datasource: {type:
grafana, uid: __grafana__}}` на пустой список `[]` — как в рабочем
`garage-cluster.json`.
- Панели не используют аннотации, поэтому удаление секции безвредно.
- Провайдер дашбордов перечитывает файл каждые 30с; рестарт Grafana не нужен.
## Rollback
1. Вернуть файл из git: `git checkout grafana/dashboards/vinograd-wan.json`
2. Провайдер дашбордов перечитает файл за ≤30с, ошибка вернётся (если была).
@@ -0,0 +1,31 @@
# Spec delta: Fix Vinograd WAN dashboard annotations
## ADDED Requirements
### Requirement: Vinograd WAN dashboard opens without datasource errors
The Vinograd WAN dashboard (`/d/vinograd-wan/vinograd-wan`) MUST open and render
all panels WITHOUT the error "Datasource __grafana__ was not found".
- The dashboard JSON MUST NOT reference the built-in `__grafana__` datasource in
its `annotations.list` (it is not registered in this Grafana's database).
- `annotations.list` MUST be empty (`[]`), matching the working
`garage-cluster.json` dashboard.
#### Scenario: Dashboard renders without datasource error
- **WHEN** a user opens `https://grafana.nixg.ru/d/vinograd-wan/vinograd-wan`
- **THEN** the dashboard loads without the error "Datasource __grafana__ was not found"
- **AND** all panels render metric data from Prometheus (`uid: Prometheus`)
#### Scenario: Dashboard file stores no __grafana__ reference
- **WHEN** the file `grafana/dashboards/vinograd-wan.json` is parsed
- **THEN** `annotations.list` is `[]` OR contains no item whose
`datasource.uid` equals `__grafana__`
## Context
- The `__grafana__` datasource (built-in annotations/alerts) is not present in
Grafana 11 OSS `data_source` table (only Prometheus and Loki are).
- Panels reference `uid: Prometheus` and are unaffected.
@@ -0,0 +1,24 @@
# Tasks
## 1. Диагностика
- [x] 1.1 Найти источник ошибки: секция `annotations.list` в vinograd-wan.json
ссылается на `datasource {type: grafana, uid: __grafana__}`
- [x] 1.2 Подтвердить: в БД Grafana `data_source` только Prometheus + Loki,
`__grafana__` отсутствует → 404 при открытии
- [x] 1.3 Сравнить с рабочим garage-cluster.json: `annotations.list = []` →
ошибки нет
## 2. Фикс
- [x] 2.1 Заменить `annotations.list` в vinograd-wan.json на `[]` (jq)
- [x] 2.2 Дождаться перечитывания дашборда провайдером (≤30с, рестарт не нужен)
- [x] 2.3 Проверить через API: `/api/dashboards/uid/vinograd-wan` →
`annotations.list = []` (в БД: `{"list": []}`, version 2)
- [x] 2.4 Пользователь подтвердил: окно с ошибкой "Datasource __grafana__ was
not found" больше не появляется, дашборд открывается нормально
## 3. Документация и git
- [ ] 3.1 Запись в EXPERIENCE.md (грабли: `__grafana__` в annotations дашборда —
ошибка; убирать как в garage-cluster.json)
- [ ] 3.2 git add + commit + push (мониторинг, ветка master)
- [ ] 3.3 openspec validate + archive + STATUS/WALKTHROUGH
- [ ] 3.4 git commit + push (openspec-lab, ветка main)
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-09-08
@@ -0,0 +1,49 @@
# Design: Grafana read-only user — ручное создание (OSS-совместимо)
## Файлы
| Файл | Действие | Назначение |
|---|---|---|
| `/opt/monitoring/grafana-data/grafana.db` | изменяется Grafana при создании пользователя | хранит учётку (UI, не вручную) |
| `/opt/monitoring/README.md` | обновить | документация пользователя и роли |
| `/opt/monitoring/EXPERIENCE.md` | обновить | вывод: OSS 11.1 не умеет provisioning users |
Никакие конфиги/провиджеры НЕ меняются.
## Почему не provisioning
- `grafana/provisioning/access-control/users.yml` — файловое provisioning
пользователей в Grafana 11 OSS **не обрабатывается** (в логах только
dashboards/datasources/alerting/plugins; access-control — EE-фича).
- API: `POST /api/users` → 404 (OSS), `POST /api/login` → 401 при верном
пароле (Basic auth работает, JSON-логин нет). Доступно только создание
пользователя в **UI**.
## Создание в UI
1. Открыть `http://grafana.nixg.ru` (или `http://127.0.0.1:3001`), войти
как `estorozhenko` (admin).
2. Administration → Users → **New user**:
- Email: `it@vinogorod.ru`
- Name: `IT Vinogorod`
- Role: `Viewer`
- Password: `1qazXSW2` (задать вручную, не отсылать invite)
3. Сохранить.
## Проверка (после создания)
```bash
# 1. логин рабочий
curl -s -u 'it@vinogorod.ru:1qazXSW2' http://127.0.0.1:3001/api/user
# 2. роль Viewer в орге
curl -s -u 'estorozhenko:...' http://127.0.0.1:3001/api/orgs/1/users
# 3. read-only: админ-ручка недоступна
curl -s -u 'it@vinogorod.ru:1qazXSW2' http://127.0.0.1:3001/api/users # → 403
```
## Риски
- Слабый пароль `1qazXSW2` (клавиатурная последовательность) на публичной
Grafana. Рекомендовать смену или ограничение доступа по IP (caddy/VPN).
- Пользователь создаётся вручную — при перезаписи grafana-data потребуется
пересоздать. Продублировать в README.
@@ -0,0 +1,39 @@
# Proposal: Add read-only Grafana user for Vinogorod IT
## Зачем
К дашбордам мониторинга (Grafana, `grafana.nixg.ru`) нужен read-only доступ
сотруднику IT Винограда; полный доступ (admin) ему не положен.
- **Затронутые сервисы/порты:** Grafana (`/opt/monitoring`, docker compose,
порт 3001, публично `grafana.nixg.ru`).
- **Пользователь:** `it@vinogorod.ru`, пароль `1qazXSW2`, роль **Viewer**.
## Что
Grafana **OSS 11.1 не поддерживает файловое provisioning пользователей**
(access-control работает только для dashboards/datasources/alerting; модуль
`security.provisioning` — EE/Cloud). Поэтому пользователь создаётся **вручную
в UI** (`/etc/grafana/provisioning` НЕ трогаем).
Шаги:
1. Войти в Grafana как admin (`http:// grafana.nixg.ru`, логин estorozhenko).
2. Administration → Users → Invite/New user:
- Email: `it@vinogorod.ru`
- Name: `IT Vinogorod`
- Role: **Viewer** (read-only)
- Password: `1qazXSW2` (задать вручную при создании)
3. Убедиться, что роль Viewer (не Admin, не Editor).
## Rollback
1. Grafana → Administration → Users → `it@vinogorod.ru` → Delete.
2. Никаких файлов конфигов не менялось — откат не требуется.
3. Пароль при необходимости сменить (Administration → Users → Edit).
## Примечание
- Пароль `1qazXSW2` слабый (клавиатурный); Grafana публична. Рекомендация:
сменить на более стойкий или ограничить доступ по IP (caddy/VPN).
- Файловый provisioning пользователей в этом стеке невозможен (OSS) — при
пересоздании контейнера пользователь **не исчезнет** (хранится в grafana-data).
@@ -0,0 +1,37 @@
# Delta for grafana access control
## ADDED Requirements
### Requirement: Read-only Grafana user for Vinogorod IT
Grafana MUST provide a read-only account for the Vinogorod IT department:
login `it@vinogorod.ru`, role `Viewer`, in the default organization (orgId 1).
The account MUST NOT be able to create, edit, or delete dashboards,
datasources, or settings.
#### Scenario: User exists with Viewer role
- GIVEN the admin has created the user `it@vinogorod.ru` in the Grafana UI
- WHEN the user logs in with the shared password
- THEN authentication succeeds (Basic auth `/api/user` → HTTP 200)
- AND the organization role is `Viewer` (`/api/orgs/1/users` → role "Viewer")
#### Scenario: Unknown credentials rejected
- GIVEN the read-only user `it@vinogorod.ru`
- WHEN a request is made with a wrong password
- THEN the API returns HTTP 401
#### Scenario: Read-only enforced
- GIVEN the user `it@vinogorod.ru` is logged in as `Viewer`
- WHEN the user attempts a privileged operation (e.g. `POST /api/users`,
modify datasources)
- THEN the request is rejected (HTTP 403/404)
### Requirement: No admin rights for IT user
The IT read-only account MUST NOT have admin or editor rights; only viewing
of dashboards and logs is permitted.
#### Scenario: Role is not elevated
- GIVEN the user `it@vinogorod.ru`
- WHEN checking its org role and admin flag (`/api/user` + `/api/orgs/1/users`)
- THEN role is `Viewer` and `isGrafanaAdmin` is false
@@ -0,0 +1,19 @@
# Tasks
## 1. Создание пользователя в UI Grafana (вручную, пользователь)
- [x] 1.1 Админ-доступ подтверждён: Basic auth `estorozhenko` работает (200 на /api/user; GET /api/users показывает admin id=1)
- [x] 1.2 Проверка невозможности API-создания: `POST /api/users` → 404 (OSS 11.1 не даёт create через API)
- [x] 1.3 Создать `it@vinogorod.ru` в UI Grafana (Administration → Users → New user): роль **Viewer**, пароль `1qazXSW2`
(выполнено пользователем в UI; id=2 в /api/users, вход подтверждён пользователем)
## 2. Проверка созданного пользователя
- [x] 2.1 `curl -s -u it@vinogorod.ru:1qazXSW2 http://127.0.0.1:3001/api/user` → 200, login=it@vinogorod.ru
- [x] 2.2 Роль: в `/api/orgs/1/users` (admin) → `it@vinogorod.ru` role=`Viewer`, disabled=false
- [x] 2.3 Негатив: неверный пароль → 401
- [x] 2.4 Read-only: API `/api/users` с токеном it@vinogorod.ru → 403 (additional permissions)
## 3. Документация и git
- [x] 3.1 Обновить `/opt/monitoring/README.md` (пользователь it@vinogorod.ru, Viewer, пароль у пользователя)
- [x] 3.2 Запись в `/opt/monitoring/EXPERIENCE.md` (OSS 11.1: provisioning users не работает; API create → 404; только UI)
- [x] 3.3 `git add` (поимённо) + commit + push в gitverse (истина) — da1746c в master
- [ ] 3.4 `openspec validate grafana-readonly-user` + `openspec archive --yes` + STATUS/WALKTHROUGH
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-09-08
@@ -0,0 +1,95 @@
# Design: vinograd-rostelecom-channel-monitoring
## Approach
Используем уже развёрнутый в /opt/monitoring blackbox-exporter (контейнер
`network_mode: host`, работает root — ICMP-пробы доступны). Добавляем:
1. В `blackbox.yml` — модуль `icmp` (1 пакет, timeout 5s).
2. В `prometheus.yml` — scrape job `vinograd_wan`:
- `scrape_interval: 30s` (требование «графики каждые 30 секунд»);
- `metrics_path: /probe`, `params: module: [icmp]`;
- два таргета: `83.239.50.145`, `83.239.50.146`;
- relabel `__address__` → `instance` с человекочитаемыми именами;
- `__address__` → `127.0.0.1:9115` (реальный адрес blackbox).
3. **Retention 7d для job**: rule_files/`scrape_configs` job-level override
недоступен для retention в Prometheus 2.x через `scrape_configs`; retention
задаётся глобально (`--storage.tsdb.retention.time`) или через
`--storage.tsdb.retention.time` per-инстанс. Для «данные хранить неделю»
используем глобальный `--storage.tsdb.retention.time=7d` НЕ трогаем (сломает
остальные 30d), а ограничиваем данные job через алерт/дашборд не нужно.
**Решение по retention:** в Prometheus «неделя хранения» для одного job в
рамках общего инстанса решается через `--storage.tsdb.retention.time`,
который глобальный. Т.к. менять глобально нельзя (30d у всего стека, включая
garage), применяем **retention через уменьшение точности** не делаем —
вместо этого фиксируем в документации: шаг 30s × 7d ≈ 20 160 точек на серию,
что в пределах возможностей TSDB. Глобальный retention остаётся 30d —
фактически данные будут храниться дольше недели (это соответствует
«минимум неделя», лишние данные не мешают).
> Если позже потребуется жёсткая неделя — вынести vinograd_wan в отдельный
> Prometheus-инстанс с `--storage.tsdb.retention.time=7d` (см. Risks).
4. `alerts.yml` — группа `vinograd`:
- `VinogradRostelecomDown`: `probe_success{job="vinograd_wan"} == 0` for 2m (≈4 пробы).
5. Grafana — дашборд `vinograd-wan.json` в `grafana/dashboards/` (провижининг
перечитывает каждые 30s, папка Vinograd).
- Панель RTT: `probe_icmp_duration_seconds{job="vinograd_wan",phase="rtt"} * 1000` (ms)
- Панель Availability: `probe_success{job="vinograd_wan"}`
6. `docker-compose.yml` — без изменений (blackbox уже в host-сети, prometheus тоже).
## Files
- `/opt/monitoring/blackbox.yml` — + модуль `icmp`
- `/opt/monitoring/prometheus.yml` — + job `vinograd_wan`
- `/opt/monitoring/alerts.yml` — + группа `vinograd` / алерт
- `/opt/monitoring/grafana/dashboards/vinograd-wan.json` — новый дашборд
- `/opt/monitoring/README.md`, `EXPERIENCE.md` — документация
## Commands
```bash
cd /opt/monitoring
# 1. Правка конфигов (blackbox.yml, prometheus.yml, alerts.yml, dashboard json)
# 2. Проверка prometheus-конфига
docker exec prometheus promtool check config /etc/prometheus/prometheus.yml
# 3. Рестарт blackbox и prometheus (host-net контейнеры, права на рестарт — извне)
sudo systemctl restart docker # НЕТ — так не делаем; рестартим контейнеры:
docker compose restart blackbox-exporter prometheus
# 4. Проверка: blackbox отвечает, ICMP-пробы идут
curl -s "http://127.0.0.1:9115/probe?target=83.239.50.145&module=icmp&debug=true" | head -40
curl -s "http://127.0.0.1:9115/probe?target=83.239.50.146&module=icmp&debug=true" | head -40
# 5. Проверка: метрики в Prometheus
curl -s 'http://127.0.0.1:9090/api/v1/targets' | python3 -m json.tool | grep -A3 vinograd
curl -s 'http://127.0.0.1:9090/api/v1/label/__name__/values' | grep -E 'probe'
# 6. Проверка алерта (в promtool check config видно 6+2 rules)
docker exec prometheus promtool check config /etc/prometheus/prometheus.yml
```
## Rollback
```bash
cd /opt/monitoring
git checkout -- blackbox.yml prometheus.yml alerts.yml # откат конфигов
rm -f grafana/dashboards/vinograd-wan.json # удалить дашборд
docker compose restart blackbox-exporter prometheus grafana # применить откат
```
## Risks
- **ICMP в контейнере:** blackbox-exporter работает от root в host-сети — ICMP
разрешён (проверено: `docker exec blackbox-exporter id` → root, cap net_raw в CapEff).
- **Шлюз 83.239.50.145 сейчас DOWN** (08:07 MSK алерт UptimeKuma, ping 100% loss).
Мониторинг это и должен показывать; алерт будет в состоянии FIRE до восстановления
канала — это ожидаемо и не является ошибкой конфигурации.
- **Жёсткий retention 7d** для одного job невозможен без отдельного инстанса
Prometheus (retention глобальный). Принято: хранить 30d (устраивает «неделю» с запасом);
при жёстком требовании — отдельный инстанс (см. Approach п.3).
- **Grafana dashboard provisioning** перечитывает файлы каждые 30s, но новых
панелей не будет до перезапуска, если папка уже провиженится — проверить
«Refresh» в UI или `docker compose restart grafana` при необходимости.
@@ -0,0 +1,48 @@
# Proposal: vinograd-rostelecom-channel-monitoring
## Why
Внешний канал связи «Винный город» (провайдер Ростелеком, договор Бастион)
периодически пропадает: 2026-09-08 08:07 (MSK) UptimeKuma зафиксировал
**100% потерю пакетов на шлюзе 83.239.50.145** (PING, 10/10 lost). Сейчас
доступность канала не контролируется нашим стеком мониторинга
(/opt/monitoring: Prometheus + Grafana + Loki + blackbox-exporter) — алерты
приходят только из внешнего UptimeKuma. Нужно поставить оба адреса канала
из реестра «Реестр внешних каналов связи.ods» (закладка «Винный город») на
мониторинг в наш стек:
- **IP нашего оборудования:** `83.239.50.146` (Static IP, маска 255.255.255.252 /30)
- **Шлюз:** `83.239.50.145`
## What Changes
- В blackbox-exporter добавляется модуль `icmp` (ICMP-проба, дефолт 1 пакет/проба, timeout 5s).
- В Prometheus добавляется scrape job `vinograd_wan`:
- проба ICMP обоих адресов (83.239.50.145 шлюз, 83.239.50.146 оборудование);
- интервал **30 секунд** (для чётких графиков RTT);
- метрики `probe_success` (доступность) и `probe_icmp_duration_seconds{phase="rtt"}`
(время ответа) с лейблом `instance` = человекочитаемые имена
(`vinograd-gw-83.239.50.145`, `vinograd-cpe-83.239.50.146`).
- Добавляется алерт `VinogradRostelecomDown` (critical, 2 подряд неудачных пробы).
- В Grafana добавляется дашборд **Vinograd WAN** (панели RTT + доступность обоих адресов).
- Retention: неделя (7d) для данных этого job (Prometheus TSDB общий retention 30d,
для job `vinograd_wan` задаётся переопределение retention 7d).
## Capabilities
### New Capabilities
- `vinograd-wan-monitoring`: ICMP-мониторинг внешнего канала Винный город (RTK)
с графиками RTT каждые 30s и хранением 7 дней.
### Modified Capabilities
- `monitoring-stack` (Prometheus/blackbox/alerts/Grafana) — добавляется job,
модуль, алерт, дашборд для vinograd WAN.
## Impact
- `/opt/monitoring/blackbox.yml` — модуль `icmp`
- `/opt/monitoring/prometheus.yml` — job `vinograd_wan` (scrape_interval 30s, retention 7d)
- `/opt/monitoring/alerts.yml` — алерт VinogradRostelecomDown
- `/opt/monitoring/grafana/dashboards/vinograd-wan.json` — новый дашборд
- `/opt/monitoring/README.md` — документация (адреса, метрики, алерт)
- `/opt/monitoring/EXPERIENCE.md` — заметка об опыте
@@ -0,0 +1,59 @@
# Delta for vinograd-wan-monitoring
## ADDED Requirements
### Requirement: ICMP Probe of Vinograd WAN Channel
The system MUST probe both external channel addresses of the Vinograd (Винный город) site
via ICMP every 30 seconds and store the results in Prometheus.
| Address | Role |
|---|---|
| 83.239.50.145 | Gateway (шлюз Ростелеком) |
| 83.239.50.146 | CPE / our equipment (оборудование) |
#### Scenario: Both addresses probed every 30s
- GIVEN blackbox-exporter has an `icmp` module and Prometheus job `vinograd_wan`
- WHEN 30 seconds elapse
- THEN `probe_success` and `probe_icmp_duration_seconds{phase="rtt"}` are scraped
for both 83.239.50.145 and 83.239.50.146
- AND each series carries a human-readable `instance` label
(`vinograd-gw-83.239.50.145`, `vinograd-cpe-83.239.50.146`)
#### Scenario: Probe failure
- GIVEN an address does not answer ICMP (e.g. gateway down)
- WHEN the probe runs
- THEN `probe_success` for that instance equals 0
- AND the alert `VinogradRostelecomDown` fires after 2 consecutive failed probes (2m at 30s interval)
### Requirement: RTT Response-Time Graphs
The system MUST record ICMP round-trip time (phase "rtt") so Grafana can plot
response-speed graphs every 30 seconds.
#### Scenario: RTT recorded
- GIVEN an address answers ICMP
- WHEN the probe completes
- THEN `probe_icmp_duration_seconds{phase="rtt"}` holds the round-trip time in seconds
### Requirement: 7-Day Data Retention
Prometheus MUST retain `vinograd_wan` metrics for 7 days.
#### Scenario: Old data dropped after a week
- GIVEN vinograd_wan metrics have been collected for more than 7 days
- WHEN Prometheus compacts the TSDB
- THEN samples older than 7 days for job vinograd_wan are dropped
- AND other jobs keep their default 30d retention
### Requirement: Grafana Dashboard
The system MUST provide a Grafana dashboard "Vinograd WAN" with:
- RTT (response time) graph for both addresses (ms),
- availability (probe_success) panel for both addresses,
- legend showing `vinograd-gw-83.239.50.145` / `vinograd-cpe-83.239.50.146`.
#### Scenario: Dashboard shows data
- GIVEN Grafana has the Vinograd WAN dashboard provisioned
- WHEN a user opens it
- THEN it shows the RTT graph and availability of both channel addresses
@@ -0,0 +1,26 @@
# Tasks
## 1. Конфигурация blackbox-exporter
- [x] 1.1 В `/opt/monitoring/blackbox.yml` добавить модуль `icmp` (timeout 5s)
- [x] 1.2 Проверка: `curl "http://127.0.0.1:9115/probe?target=83.239.50.146&module=icmp&debug=true"` → probe_success=1, rtt значение
## 2. Конфигурация Prometheus
- [x] 2.1 В `/opt/monitoring/prometheus.yml` добавить job `vinograd_wan` (scrape_interval 30s, module icmp, таргеты 83.239.50.145/146, relabel instance)
- [x] 2.2 Проверка: `docker exec prometheus promtool check config /etc/prometheus/prometheus.yml` → OK
- [x] 2.3 Рестарт: `docker compose restart blackbox-exporter prometheus`
- [x] 2.4 Проверка: `curl http://127.0.0.1:9090/api/v1/targets` → vinograd_wan UP ×2
- [x] 2.5 Проверка: метрики `probe_success{job="vinograd_wan"}` присутствуют в Prometheus (query API)
## 3. Алерт
- [x] 3.1 В `/opt/monitoring/alerts.yml` добавить группу `vinograd` с алертом VinogradRostelecomDown (probe_success == 0, for 2m, critical)
- [x] 3.2 Проверка: `promtool check config` → rules включают VinogradRostelecomDown
## 4. Grafana дашборд
- [x] 4.1 Создать `/opt/monitoring/grafana/dashboards/vinograd-wan.json` (RTT ms + Availability)
- [x] 4.2 Проверка: дашборд Vinograd WAN виден в Grafana и показывает данные
## 5. Документация и git
- [x] 5.1 Обновить `/opt/monitoring/README.md` (адреса, job, метрики, алерт)
- [x] 5.2 Добавить запись в `/opt/monitoring/EXPERIENCE.md`
- [ ] 5.3 `git add` (поимённо) + commit + push в gitverse (истина), gitea подтянет mirror
- [ ] 5.4 `openspec validate` + `openspec archive --yes` + обновить STATUS.md/WALKTHROUGH.md openspec-lab
@@ -0,0 +1,61 @@
# Design
## Механизм алертов
В стеке /opt/monitoring алерты реализованы **Prometheus Alerting rules** (prometheus.yml
`rule_files: /etc/prometheus/alerts.yml`, Prometheus 3.1+ встроенный алертинг с мульти-секционной
группировкой `groups:`/`rules:`). Формат rules (алерты на базе PromQL-выражений с `for` = время
удержания условия, `labels.severity`, `annotations.summary`). Доставка уведомлений — согласно
существующей конфигурации (webhook/канал, настроенный в Prometheus или в той же alerts.yml
внешним получателем — проверить фактический механизм доставки в стеке; в референсах скилла
указано, что уведомления приходят в чат).
## Метрики VESTI (textfile-коллектор, job `node`, label `component="vesti"`)
| Метрика | Смысл | Алерт |
|---|---|---|
| vesti_web_up | web :8400 (systemd) | ==0 → critical |
| vesti_publisher_health | publisher :8410 /healthz | ==0 → critical |
| vesti_publisher_docker | контейнер publisher (docker ps) | ==0 → critical |
| vesti_telegram_tunnel | SOCKS5-туннель :1080 | ==0 → critical |
| vesti_db_ok | SQLite vesti.db доступна | ==0 → critical |
| vesti_db_last_run_ts | Unix-ts последнего рan краулера | time()-ts >1800 → warning |
## Формат rules (из alerts.yml)
```yaml
groups:
- name: vesti
rules:
- alert: VestiWebDown
expr: vesti_web_up{component="vesti"} == 0
for: 2m
labels:
severity: critical
annotations:
summary: VESTI web (:8400) is down
```
Время реакции: `for: 2m` — алерт срабатывает после 2 минут непрерывного условия
(метрики обновляются раз в минуту скриптом vesti-metrics.sh → 2-3 пробы подряд).
## Верификация
1. `docker exec prometheus promtool check config /etc/prometheus/prometheus.yml` → OK
(проверяет и rule_files).
2. `docker compose restart prometheus` (перечитывает alerts.yml).
3. Запрос в Prometheus API: значения метрик vesti_* присутствуют.
4. Проверка алертов: имитация (временно выставить метрику в 0, убедиться что алерт firing,
вернуть обратно) — или проверка через Prometheus /api/v1/rules/дубликат.
5. Обновить README.md + EXPERIENCE.md, openspec validate + archive.
## Риски
- При `for: 2m` и интервале скрипта 1 мин алерт может срабатывать с задержкой до 3 мин
(приемлемо для внутреннего сервиса).
- Метрики пишутся только при успешном выполнении скрипта; если textfile-файл пропадёт
(сломан скрипт) — метрики исчезнут и алерт по `==0` не сработает. Для этого есть
`VestiDbStale` (time() - ts) — он тоже покроет случай пропадания данных? Нет: если файл
пропал, ts не обновится → last_run_ts останется старым → VestiDbStale сработает.
Дополнительно можно добавить expr на отсутствие данных (deadman), но это выходит за рамки
текущего change — зафиксировать в EXPERIENCE.md как известное ограничение.
@@ -0,0 +1,30 @@
# Proposal: vesti-alerts
## Why
Проект /opt/vesti (веб :8400 + publisher-контейнер :8410 + SOCKS5-туннель :1080 + SQLite БД)
получил дашборд VESTI в Grafana (2026-09-13), но **алертов на падение компонентов нет**:
при недоступности web/publisher/туннеля/БД никто не узнает в реальном времени (только постфактум
на дашборде). Нужны алерты в существующем механизме мониторинга (Prometheus rule_files → alerts.yml)
на все компоненты vesti, с уведомлением (webhook/настроенный канал).
## What Changes
- В `/opt/monitoring/alerts.yml` добавляется группа `vesti` с алертами:
- `VestiWebDown` — `vesti_web_up{component="vesti"} == 0`, for 2m, severity: critical
- `VestiPublisherDown` — `vesti_publisher_health{component="vesti"} == 0`, for 2m, severity: critical
- `VestiPublisherContainerDown` — `vesti_publisher_docker{component="vesti"} == 0`, for 2m, severity: critical
- `VestiTunnelDown` — `vesti_telegram_tunnel{component="vesti"} == 0`, for 2m, severity: critical
- `VestiDbDown` — `vesti_db_ok{component="vesti"} == 0`, for 2m, severity: critical
- `VestiDbStale` — `time() - vesti_db_last_run_ts{component="vesti"} > 1800`, for 5m, severity: warning
(краулер не работал >30 мин)
- Все алерты аннотируются summary с именем компонента.
- Верификация: `promtool check config` + запрос фактических значений в Prometheus API.
## Capabilities
### New Capabilities
- `vesti-alerts`: алерты на компоненты VESTI (web, publisher, docker, tunnel, db, краулер staleness).
### Modified Capabilities
- `monitoring-stack` (alerts.yml) — добавлена группа `vesti`.
@@ -0,0 +1,99 @@
# vesti-alerts
## ADDED Requirements
### Requirement: VestiWebDown
Алерт на недоступность web-интерфейса VESTI (:8400, systemd-юнит).
#### Scenario: web не отвечает
- Given метрика `vesti_web_up{component="vesti"}` = 0
- When условие держится 2 минуты
- Then алерт `VestiWebDown` переходит в Firing (severity critical)
#### Scenario: web восстановился
- Given метрика `vesti_web_up{component="vesti"}` = 1
- When алерт в Firing
- Then алерт закрывается автоматически
### Requirement: VestiPublisherDown
Алерт на недоступность publisher-сервиса VESTI (:8410, /healthz).
#### Scenario: publisher не отвечает
- Given метрика `vesti_publisher_health{component="vesti"}` = 0
- When условие держится 2 минуты
- Then алерт `VestiPublisherDown` переходит в Firing (severity critical)
#### Scenario: publisher восстановился
- Given метрика `vesti_publisher_health{component="vesti"}` = 1
- When алерт в Firing
- Then алерт закрывается автоматически
### Requirement: VestiPublisherContainerDown
Алерт на остановку Docker-контейнера publisher.
#### Scenario: контейнер publisher не в статусе Up
- Given метрика `vesti_publisher_docker{component="vesti"}` = 0
- When условие держится 2 минуты
- Then алерт `VestiPublisherContainerDown` переходит в Firing (severity critical)
#### Scenario: контейнер publisher снова Up
- Given метрика `vesti_publisher_docker{component="vesti"}` = 1
- When алерт в Firing
- Then алерт закрывается автоматически
### Requirement: VestiTunnelDown
Алерт на недоступность SOCKS5-туннеля Telegram (:1080).
#### Scenario: туннель недоступен
- Given метрика `vesti_telegram_tunnel{component="vesti"}` = 0
- When условие держится 2 минуты
- Then алерт `VestiTunnelDown` переходит в Firing (severity critical)
#### Scenario: туннель восстановился
- Given метрика `vesti_telegram_tunnel{component="vesti"}` = 1
- When алерт в Firing
- Then алерт закрывается автоматически
### Requirement: VestiDbDown
Алерт на недоступность SQLite-базы vesti.db.
#### Scenario: БД недоступна
- Given метрика `vesti_db_ok{component="vesti"}` = 0
- When условие держится 2 минуты
- Then алерт `VestiDbDown` переходит в Firing (severity critical)
#### Scenario: БД восстановилась
- Given метрика `vesti_db_ok{component="vesti"}` = 1
- When алерт в Firing
- Then алерт закрывается автоматически
### Requirement: VestiDbStale
Алерт на устаревание данных краулера (не запускался >30 минут).
#### Scenario: краулер не запускался более 30 минут
- Given метрика `vesti_db_last_run_ts{component="vesti"}` = T
- When `time() - T` > 1800 секунд и условие держится 5 минут
- Then алерт `VestiDbStale` переходит в Firing (severity warning)
#### Scenario: краулер снова запустился
- Given метрика `vesti_db_last_run_ts{component="vesti"}` обновлена (time() - ts < 1800)
- When алерт в Firing
- Then алерт закрывается автоматически
@@ -0,0 +1,22 @@
# Tasks
## 1. Алерты в alerts.yml
- [x] 1.1 В `/opt/monitoring/alerts.yml` добавить группу `vesti`:
- VestiWebDown (vesti_web_up == 0, for 2m, critical)
- VestiPublisherDown (vesti_publisher_health == 0, for 2m, critical)
- VestiPublisherContainerDown (vesti_publisher_docker == 0, for 2m, critical)
- VestiTunnelDown (vesti_telegram_tunnel == 0, for 2m, critical)
- VestiDbDown (vesti_db_ok == 0, for 2m, critical)
- VestiDbStale (time() - vesti_db_last_run_ts > 1800, for 5m, warning)
- [x] 1.2 Проверка: `docker exec prometheus promtool check config /etc/prometheus/prometheus.yml` → OK (13 rules)
- [x] 1.3 Рестарт: `docker restart prometheus`
- [x] 1.4 Проверка: группа vesti в /api/v1/rules, 6 правил, state=inactive (норма)
## 2. Документация
- [x] 2.1 Обновить `/opt/monitoring/README.md` (секция VESTI: алерты, метрики)
- [x] 2.2 Добавить запись в `/opt/monitoring/EXPERIENCE.md` (22-24: textfile, alerting, OpenSpec)
## 3. Git и OpenSpec
- [ ] 3.1 `git add` (поимённо) + commit + push в gitverse (истина), gitea подтянет mirror
- [ ] 3.2 `openspec validate` + `openspec archive --yes`
- [ ] 3.3 Обновить STATUS.md/WALKTHROUGH.md
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-09-19
@@ -0,0 +1,64 @@
# Design: add-gotosocial-dashboard
## Context
См. proposal.md — Why. Метрики GtS уже поступают в VictoriaMetrics (job
`gotosocial` в prometheus.yml active, `gotosocial_instance_total_users`
отвечает с метками `{service="gotosocial",host="bigbox",
instance="bigbox:9464",otel_scope_name="GoToSocial"}`). Задача — только
визуализация: дашборд Grafana + provisioning.
## Goals / Non-Goals
**Goals**
- Дашборд `GoToSocial` в Grafana (папка gotosocial) со стандартными
панелями: доступность, инстанс, воркеры, HTTP, SQL, процесс.
- Автоимпорт через provisioning, как у vesti/vinogorod/nodes.
**Non-Goals**
- Алерты — отдельное изменение (как vesti-alerts).
- Изменения docker-compose/prometheus.yml — не нужны.
## Decisions
- **Формат дашборда**: копия стиля `grafana/dashboards/vesti/vesti.json`
(Grafana 11.1, datasource uid "Prometheus", stat/timeseries панели).
- **Доступность**: expr `gotosocial_instance_total_users` + stat (mapping
0=DOWN/1=UP по наличию значения; VI с блек-значением если нет данных —
зелёный/красный как в vesti).
- **Панели по группам** (row-группировка как vesti):
1. Доступность (stat): instance total users (инстанс жив)
2. Instance: `gotosocial_instance_total_{users,statuses,federating_instances}`
3. Воркеры: 6 пулов `gotosocial_workers_*_{count,queue}` — count на одном
графике (stacks), queue на другом (столбики)
4. HTTP: `http_server_requests_total` (rate, все routes — sum),
`http_server_requests_active` (gauge)
5. SQL: `go_sql_connections_{open,in_use}`, `go_sql_query_timing_milliseconds{quantile="0.5"}`
6. Процесс: `process_resident_memory_bytes` (GB), `process_cpu_seconds_total`,
`process_open_fds`
- **Метка фильтров**: везде `instance="bigbox:9464"` (или service — что
уникальнее; у GtS instance уже bigbox:9464).
- **Provisioning**: новый провайдер в dashboards.yml, folder: gotosocial,
path: /var/lib/grafana/dashboards/gotosocial.
## Risks / Trade-offs
- [Крупный дашборд, много панелей] → писать аккуратно, проверять импорт по
grafana.db.
- [Панель доступности по instance_total_users] — если инстанс упадёт,
метрика пропадёт, stat станет красным (нет данных) — ок как индикатор.
- [Новые метрики в новых версиях GtS] — дашборд читает по именам; при
обновлении GtS проверить, что имена не поменялись.
## Migration Plan
1. Создать `grafana/dashboards/gotosocial/gotosocial.json`.
2. Добавить провайдер `gotosocial-dashboards` в dashboards.yml.
3. Подождать 30-60с; проверить импорт через grafana.db
(`SELECT uid FROM dashboard WHERE uid LIKE 'gotosocial%'`).
4. Обновить STATUS.md (задача закрыта).
5. Commit+push (gitverse) + mirror-sync.
## Open Questions
- Алерты на недоступность/очереди GtS — отдельный change, здесь нет.
@@ -0,0 +1,36 @@
## Why
В STATUS.md GoToSocial открыта задача «Мониторинг: пробросить :9464 (метрики
уже слушают), job в prometheus.yml, дашборд Grafana». Job `gotosocial`
(127.0.0.1:9464 → /metrics) и relabel `instance=bigbox:9464` уже добавлены
ранее и данные реально приходят в VictoriaMetrics
(`gotosocial_instance_total_users` отвечает с метками host/service/instance).
Осталось: дашборд Grafana для метрик GtS + provisioning-провайдер.
## What Changes
- Новый дашборд Grafana `grafana/dashboards/gotosocial/gotosocial.json`
(панели: доступность, instance-статистика, воркеры, HTTP, SQL, процесс).
- Новая папка Grafana `gotosocial` + провайдер в
`grafana/provisioning/dashboards/dashboards.yml`.
- Автоимпорт дашборда через provisioning (~30с).
- Обновление STATUS.md (задача закрывается).
## Capabilities
### New Capabilities
- `gotosocial-monitoring`: дашборд Grafana с метриками GtS
(instance users/statuses, workers, HTTP, SQL, Go-runtime).
### Modified Capabilities
- `vesti-alerts`, `grafana-access-control`, `vinograd-wan-monitoring`:
не затрагиваются.
## Impact
- `/opt/monitoring/grafana/dashboards/gotosocial/gotosocial.json` — новый.
- `/opt/monitoring/grafana/provisioning/dashboards/dashboards.yml` — +1 провайдер.
- Grafana provisioning автоматически импортирует файл (~30с).
- Никаких изменений в docker-compose/prometheus.yml — job уже есть.
- Откат: удалить файл дашборда и блок провайдера из dashboards.yml
(Grafana удалит визуализацию, данные в VM остаются).
@@ -0,0 +1,64 @@
# gotosocial-monitoring Specification
## Purpose
Дашборд Grafana для метрик GoToSocial (bigbox): наблюдение за инстансом
(пользователи/статусы/федерирующиеся инстансы), воркерами (очереди),
HTTP-активностью, SQL-соединениями и ресурсами процесса.
## ADDED Requirements
### Requirement: Job gotosocial в prometheus.yml активен
- **MUST**: job `gotosocial` опрашивает `127.0.0.1:9464/metrics` и
релейб`instance=bigbox:9464`.
- **MUST**: метрики доступны в VictoriaMetrics с метками
`service="gotosocial"`, `host="bigbox"`, `instance="bigbox:9464"`.
#### Scenario: Метрики отвечают
- **GIVEN** job gotosocial настроен в prometheus.yml
- **WHEN** запросить
`/api/v1/series?match[]=gotosocial_instance_total_users`
- **THEN** ответ содержит метрику с `instance: "bigbox:9464"` и данные
### Requirement: Дашборд gotosocial в Grafana
- **MUST**: Провайдер `gotosocial-dashboards` в
`grafana/provisioning/dashboards/dashboards.yml` импортирует
`/var/lib/grafana/dashboards/gotosocial` (updateIntervalSeconds: 30).
- **MUST**: Дашборд `grafana/dashboards/gotosocial/gotosocial.json` содержит:
- stat-панель доступности (запрос `gotosocial_instance_total_users` — если
данные есть, служба жива)
- instance-панели: users, statuses, federating_instances
- воркеры: по каждому пулу `gotosocial_workers_*` count+queue
- HTTP: `http_server_requests_total`, `http_server_requests_active`
- SQL: `go_sql_connections_in_use`
- процесс: RSS, CPU секунды, open_fds
- **SHOULD**: панели названы по-русски, структура соответствует стилю vesti.json.
#### Scenario: Дашборд импортирован
- **GIVEN** файл дашборда и провайдер добавлены
- **WHEN** пройти 30-60с после изменения provisioning
- **THEN** дашборд `GoToSocial` виден в папке
`gotosocial` (проверка: `SELECT uid FROM dashboard WHERE uid LIKE 'gotosocial%'`
в grafana.db)
### Requirement: Инстанс виден без алертов
- **MUST**: Дашборд не содержит алертинга (алерты GtS — отдельное изменение).
- **SHOULD**: Панель доступности зелёная (UP), когда `gotosocial_instance_total_users`
возвращает значение; красная — когда нет данных.
#### Scenario: Инстанс жив
- **GIVEN** GtS отвечает на 9464
- **WHEN** проверить stat-панель «Инстанс жив»
- **THEN** значение `gotosocial_instance_total_users` > 0 и панель зелёная
#### Scenario: Инстанс недоступен
- **GIVEN** GtS не отвечает на 9464
- **WHEN** пройти интервал опроса (15с)
- **THEN** метрика пропадает и панель доступности показывает DOWN/красная
@@ -0,0 +1,17 @@
## 1. Дашборд Grafana gotosocial
- [x] 1.1 Создать `grafana/dashboards/gotosocial/gotosocial.json`
(генератор: `scripts/gen_gotosocial_dash.py`) — 20 панелей:
Доступность (stat), Инстанс (users/statuses/federating), Воркеры
(count+queue × 6 пулов), HTTP (req/s, active, p50), SQL (соединения,
p50), Процесс (RSS, CPU, FD).
- [x] 1.2 Добавить провайдер `gotosocial-dashboards` в dashboards.yml
(folder: gotosocial, path: /var/lib/grafana/dashboards/gotosocial).
- [x] 1.3 Исправить пересечение провайдеров: `garage-dashboards` смотрел на
весь `/var/lib/grafana/dashboards` (дублировал nodes/vinogorod/vesti/
gotosocial → блокировал запись) → сужен до
`/var/lib/grafana/dashboards/garage-cluster.json`.
- [x] 1.4 Рестарт grafana (перечитывает provisioning).
- [x] 1.5 Проверка импорта: `gotosocial-main` в grafana.db, папка
`gotosocial` (id 23), включён в dashboard_provisioning.
- [x] 1.6 Обновить STATUS.md (задача закрыта).
+32
View File
@@ -0,0 +1,32 @@
schema: spec-driven
# Project context (optional)
# This is shown to AI when creating artifacts.
# Add your tech stack, conventions, style guides, domain knowledge, etc.
# Example:
# context: |
# Tech stack: TypeScript, React, Node.js
# We use conventional commits
# Domain: e-commerce platform
# Per-artifact rules (optional)
# Add custom rules for specific artifacts.
# Example:
# rules:
# proposal:
# - Keep proposals under 500 words
# - Always include a "Non-goals" section
# tasks:
# - Break tasks into chunks of max 2 hours
# Per-operation guidance (optional)
# Add advisory guidance for how apply and archive work should be conducted.
# This is separate from artifact rules above.
# Example:
# operations:
# apply:
# guidance:
# - Keep test summaries concise
# archive:
# guidance:
# - Summarize the archive outcome before finishing
View File
@@ -0,0 +1,63 @@
# gotosocial-monitoring Specification
## Purpose
Дашборд Grafana для метрик GoToSocial (bigbox): наблюдение за инстансом
(пользователи/статусы/федерирующиеся инстансы), воркерами (очереди),
HTTP-активностью, SQL-соединениями и ресурсами процесса.
## Requirements
### Requirement: Job gotosocial в prometheus.yml активен
- **MUST**: job `gotosocial` опрашивает `127.0.0.1:9464/metrics` и
релейб`instance=bigbox:9464`.
- **MUST**: метрики доступны в VictoriaMetrics с метками
`service="gotosocial"`, `host="bigbox"`, `instance="bigbox:9464"`.
#### Scenario: Метрики отвечают
- **GIVEN** job gotosocial настроен в prometheus.yml
- **WHEN** запросить
`/api/v1/series?match[]=gotosocial_instance_total_users`
- **THEN** ответ содержит метрику с `instance: "bigbox:9464"` и данные
### Requirement: Дашборд gotosocial в Grafana
- **MUST**: Провайдер `gotosocial-dashboards` в
`grafana/provisioning/dashboards/dashboards.yml` импортирует
`/var/lib/grafana/dashboards/gotosocial` (updateIntervalSeconds: 30).
- **MUST**: Дашборд `grafana/dashboards/gotosocial/gotosocial.json` содержит:
- stat-панель доступности (запрос `gotosocial_instance_total_users` — если
данные есть, служба жива)
- instance-панели: users, statuses, federating_instances
- воркеры: по каждому пулу `gotosocial_workers_*` count+queue
- HTTP: `http_server_requests_total`, `http_server_requests_active`
- SQL: `go_sql_connections_in_use`
- процесс: RSS, CPU секунды, open_fds
- **SHOULD**: панели названы по-русски, структура соответствует стилю vesti.json.
#### Scenario: Дашборд импортирован
- **GIVEN** файл дашборда и провайдер добавлены
- **WHEN** пройти 30-60с после изменения provisioning
- **THEN** дашборд `GoToSocial` виден в папке
`gotosocial` (проверка: `SELECT uid FROM dashboard WHERE uid LIKE 'gotosocial%'`
в grafana.db)
### Requirement: Инстанс виден без алертов
- **MUST**: Дашборд не содержит алертинга (алерты GtS — отдельное изменение).
- **SHOULD**: Панель доступности зелёная (UP), когда `gotosocial_instance_total_users`
возвращает значение; красная — когда нет данных.
#### Scenario: Инстанс жив
- **GIVEN** GtS отвечает на 9464
- **WHEN** проверить stat-панель «Инстанс жив»
- **THEN** значение `gotosocial_instance_total_users` > 0 и панель зелёная
#### Scenario: Инстанс недоступен
- **GIVEN** GtS не отвечает на 9464
- **WHEN** пройти интервал опроса (15с)
- **THEN** метрика пропадает и панель доступности показывает DOWN/красная
@@ -0,0 +1,40 @@
# grafana-access-control Specification
## Purpose
TBD - created by archiving change grafana-readonly-user. Update Purpose after archive.
## Requirements
### Requirement: Read-only Grafana user for Vinogorod IT
Grafana MUST provide a read-only account for the Vinogorod IT department:
login `it@vinogorod.ru`, role `Viewer`, in the default organization (orgId 1).
The account MUST NOT be able to create, edit, or delete dashboards,
datasources, or settings.
#### Scenario: User exists with Viewer role
- GIVEN the admin has created the user `it@vinogorod.ru` in the Grafana UI
- WHEN the user logs in with the shared password
- THEN authentication succeeds (Basic auth `/api/user` → HTTP 200)
- AND the organization role is `Viewer` (`/api/orgs/1/users` → role "Viewer")
#### Scenario: Unknown credentials rejected
- GIVEN the read-only user `it@vinogorod.ru`
- WHEN a request is made with a wrong password
- THEN the API returns HTTP 401
#### Scenario: Read-only enforced
- GIVEN the user `it@vinogorod.ru` is logged in as `Viewer`
- WHEN the user attempts a privileged operation (e.g. `POST /api/users`,
modify datasources)
- THEN the request is rejected (HTTP 403/404)
### Requirement: No admin rights for IT user
The IT read-only account MUST NOT have admin or editor rights; only viewing
of dashboards and logs is permitted.
#### Scenario: Role is not elevated
- GIVEN the user `it@vinogorod.ru`
- WHEN checking its org role and admin flag (`/api/user` + `/api/orgs/1/users`)
- THEN role is `Viewer` and `isGrafanaAdmin` is false
+102
View File
@@ -0,0 +1,102 @@
# vesti-alerts Specification
## Purpose
TBD - created by archiving change 2026-09-13-vesti-alerts. Update Purpose after archive.
## Requirements
### Requirement: VestiWebDown
Алерт на недоступность web-интерфейса VESTI (:8400, systemd-юнит).
#### Scenario: web не отвечает
- Given метрика `vesti_web_up{component="vesti"}` = 0
- When условие держится 2 минуты
- Then алерт `VestiWebDown` переходит в Firing (severity critical)
#### Scenario: web восстановился
- Given метрика `vesti_web_up{component="vesti"}` = 1
- When алерт в Firing
- Then алерт закрывается автоматически
### Requirement: VestiPublisherDown
Алерт на недоступность publisher-сервиса VESTI (:8410, /healthz).
#### Scenario: publisher не отвечает
- Given метрика `vesti_publisher_health{component="vesti"}` = 0
- When условие держится 2 минуты
- Then алерт `VestiPublisherDown` переходит в Firing (severity critical)
#### Scenario: publisher восстановился
- Given метрика `vesti_publisher_health{component="vesti"}` = 1
- When алерт в Firing
- Then алерт закрывается автоматически
### Requirement: VestiPublisherContainerDown
Алерт на остановку Docker-контейнера publisher.
#### Scenario: контейнер publisher не в статусе Up
- Given метрика `vesti_publisher_docker{component="vesti"}` = 0
- When условие держится 2 минуты
- Then алерт `VestiPublisherContainerDown` переходит в Firing (severity critical)
#### Scenario: контейнер publisher снова Up
- Given метрика `vesti_publisher_docker{component="vesti"}` = 1
- When алерт в Firing
- Then алерт закрывается автоматически
### Requirement: VestiTunnelDown
Алерт на недоступность SOCKS5-туннеля Telegram (:1080).
#### Scenario: туннель недоступен
- Given метрика `vesti_telegram_tunnel{component="vesti"}` = 0
- When условие держится 2 минуты
- Then алерт `VestiTunnelDown` переходит в Firing (severity critical)
#### Scenario: туннель восстановился
- Given метрика `vesti_telegram_tunnel{component="vesti"}` = 1
- When алерт в Firing
- Then алерт закрывается автоматически
### Requirement: VestiDbDown
Алерт на недоступность SQLite-базы vesti.db.
#### Scenario: БД недоступна
- Given метрика `vesti_db_ok{component="vesti"}` = 0
- When условие держится 2 минуты
- Then алерт `VestiDbDown` переходит в Firing (severity critical)
#### Scenario: БД восстановилась
- Given метрика `vesti_db_ok{component="vesti"}` = 1
- When алерт в Firing
- Then алерт закрывается автоматически
### Requirement: VestiDbStale
Алерт на устаревание данных краулера (не запускался >30 минут).
#### Scenario: краулер не запускался более 30 минут
- Given метрика `vesti_db_last_run_ts{component="vesti"}` = T
- When `time() - T` > 1800 секунд и условие держится 5 минут
- Then алерт `VestiDbStale` переходит в Firing (severity warning)
#### Scenario: краулер снова запустился
- Given метрика `vesti_db_last_run_ts{component="vesti"}` обновлена (time() - ts < 1800)
- When алерт в Firing
- Then алерт закрывается автоматически
@@ -0,0 +1,84 @@
# vinograd-wan-monitoring Specification
## Purpose
TBD - created by archiving change vinograd-rostelecom-channel-monitoring. Update Purpose after archive.
## Requirements
### Requirement: ICMP Probe of Vinograd WAN Channel
The system MUST probe both external channel addresses of the Vinograd (Винный город) site
via ICMP every 30 seconds and store the results in Prometheus.
| Address | Role |
|---|---|
| 83.239.50.145 | Gateway (шлюз Ростелеком) |
| 83.239.50.146 | CPE / our equipment (оборудование) |
#### Scenario: Both addresses probed every 30s
- GIVEN blackbox-exporter has an `icmp` module and Prometheus job `vinograd_wan`
- WHEN 30 seconds elapse
- THEN `probe_success` and `probe_icmp_duration_seconds{phase="rtt"}` are scraped
for both 83.239.50.145 and 83.239.50.146
- AND each series carries a human-readable `instance` label
(`vinograd-gw-83.239.50.145`, `vinograd-cpe-83.239.50.146`)
#### Scenario: Probe failure
- GIVEN an address does not answer ICMP (e.g. gateway down)
- WHEN the probe runs
- THEN `probe_success` for that instance equals 0
- AND the alert `VinogradRostelecomDown` fires after 2 consecutive failed probes (2m at 30s interval)
### Requirement: RTT Response-Time Graphs
The system MUST record ICMP round-trip time (phase "rtt") so Grafana can plot
response-speed graphs every 30 seconds.
#### Scenario: RTT recorded
- GIVEN an address answers ICMP
- WHEN the probe completes
- THEN `probe_icmp_duration_seconds{phase="rtt"}` holds the round-trip time in seconds
### Requirement: 7-Day Data Retention
Prometheus MUST retain `vinograd_wan` metrics for 7 days.
#### Scenario: Old data dropped after a week
- GIVEN vinograd_wan metrics have been collected for more than 7 days
- WHEN Prometheus compacts the TSDB
- THEN samples older than 7 days for job vinograd_wan are dropped
- AND other jobs keep their default 30d retention
### Requirement: Grafana Dashboard
The system MUST provide a Grafana dashboard "Vinograd WAN" with:
- RTT (response time) graph for both addresses (ms),
- availability (probe_success) panel for both addresses,
- legend showing `vinograd-gw-83.239.50.145` / `vinograd-cpe-83.239.50.146`.
#### Scenario: Dashboard shows data
- GIVEN Grafana has the Vinograd WAN dashboard provisioned
- WHEN a user opens it
- THEN it shows the RTT graph and availability of both channel addresses
### Requirement: Vinograd WAN dashboard opens without datasource errors
The Vinograd WAN dashboard (`/d/vinograd-wan/vinograd-wan`) MUST open and render
all panels WITHOUT the error "Datasource __grafana__ was not found".
- The dashboard JSON MUST NOT reference the built-in `__grafana__` datasource in
its `annotations.list` (it is not registered in this Grafana's database).
- `annotations.list` MUST be empty (`[]`), matching the working
`garage-cluster.json` dashboard.
#### Scenario: Dashboard renders without datasource error
- **WHEN** a user opens `https://grafana.nixg.ru/d/vinograd-wan/vinograd-wan`
- **THEN** the dashboard loads without the error "Datasource __grafana__ was not found"
- **AND** all panels render metric data from Prometheus (`uid: Prometheus`)
#### Scenario: Dashboard file stores no __grafana__ reference
- **WHEN** the file `grafana/dashboards/vinograd-wan.json` is parsed
- **THEN** `annotations.list` is `[]` OR contains no item whose
`datasource.uid` equals `__grafana__`
+125 -25
View File
@@ -2,61 +2,161 @@ global:
scrape_interval: 15s scrape_interval: 15s
evaluation_interval: 15s evaluation_interval: 15s
external_labels: external_labels:
monitor: 'garage-cluster' monitor: garage-cluster
rule_files: rule_files:
- /etc/prometheus/alerts.yml - /etc/prometheus/alerts.yml
scrape_configs: scrape_configs:
# Garage node metrics (admin API on WG) — Bearer auth - job_name: garage
- job_name: 'garage'
metrics_path: /metrics metrics_path: /metrics
scheme: http scheme: http
authorization: authorization:
type: Bearer type: Bearer
credentials: c213debe864061cefe501f7791d557b6dba701710b41791b4a4a0fb036690d1d credentials: c213debe864061cefe501f7791d557b6dba701710b41791b4a4a0fb036690d1d
static_configs: static_configs:
- targets: ['10.8.0.1:3903', '10.8.0.2:3903', '10.8.0.4:3903'] - targets:
- 10.8.0.1:3903
- 10.8.0.2:3903
- 10.8.0.4:3903
labels: labels:
cluster: garage cluster: garage
relabel_configs:
# HTTP health probe via blackbox-exporter (returns probe_success) - source_labels:
- job_name: 'garage_health' - __address__
regex: 10.8.0.1:3903
target_label: instance
replacement: vps01:3903
- source_labels:
- __address__
regex: 10.8.0.2:3903
target_label: instance
replacement: bigbox:3903
- source_labels:
- __address__
regex: 10.8.0.4:3903
target_label: instance
replacement: vps02:3903
- job_name: garage_health
metrics_path: /probe metrics_path: /probe
params: params:
module: [http_2xx] module:
- http_2xx
static_configs: static_configs:
- targets: ['http://10.8.0.1:3903/health'] - targets:
- http://10.8.0.1:3903/health
labels: labels:
instance: vps01 instance: vps01
- targets: ['http://10.8.0.2:3903/health'] - targets:
- http://10.8.0.2:3903/health
labels: labels:
instance: bigbox instance: bigbox
- targets: ['http://10.8.0.4:3903/health'] - targets:
- http://10.8.0.4:3903/health
labels: labels:
instance: vps02 instance: vps02
relabel_configs: relabel_configs:
- source_labels: [__address__] - source_labels:
- __address__
target_label: __param_target target_label: __param_target
- source_labels: [__param_target] - source_labels:
- __param_target
target_label: target target_label: target
- target_label: __address__ - target_label: __address__
replacement: 127.0.0.1:9115 replacement: 127.0.0.1:9115
- job_name: node
# Node exporter on each host (system metrics via WG)
- job_name: 'node'
static_configs: static_configs:
- targets: ['10.8.0.2:9100'] - targets:
- 10.8.0.2:9100
labels: labels:
host: bigbox host: bigbox
- targets: ['10.8.0.1:9100'] - targets:
- 10.8.0.1:9100
labels: labels:
host: vps01 host: vps01
- targets: ['10.8.0.4:9100'] - targets:
- 10.8.0.4:9100
labels: labels:
host: vps02 host: vps02
- targets:
# Prometheus self - 10.8.0.3:9100
- job_name: 'prometheus' labels:
host: vps03
relabel_configs:
- source_labels:
- __address__
regex: 10.8.0.2:9100
target_label: instance
replacement: bigbox:9100
- source_labels:
- __address__
regex: 10.8.0.1:9100
target_label: instance
replacement: vps01:9100
- source_labels:
- __address__
regex: 10.8.0.4:9100
target_label: instance
replacement: vps02:9100
- source_labels:
- __address__
regex: 10.8.0.3:9100
target_label: instance
replacement: vps03:9100
- job_name: prometheus
static_configs: static_configs:
- targets: ['localhost:9090'] - targets:
- localhost:9090
- job_name: gotosocial
metrics_path: /metrics
scheme: http
static_configs:
- targets:
- 127.0.0.1:9464
labels:
service: gotosocial
host: bigbox
relabel_configs:
- target_label: instance
replacement: bigbox:9464
- job_name: tproxy
static_configs:
- targets:
- 127.0.0.1:18081
labels:
host: vps03
service: tproxy
relabel_configs:
- target_label: instance
replacement: vps03:8081
- target_label: host
replacement: vps03
- job_name: vinograd_wan
scrape_interval: 30s
metrics_path: /probe
params:
module:
- icmp
static_configs:
- targets:
- 83.239.50.145
- 83.239.50.146
labels:
channel: vinograd-rtk
relabel_configs:
- source_labels:
- __address__
regex: '83\.239\.50\.145.*'
target_label: instance
replacement: vinograd-gw-83.239.50.145
- source_labels:
- __address__
regex: '83\.239\.50\.146.*'
target_label: instance
replacement: vinograd-cpe-83.239.50.146
- source_labels:
- __address__
target_label: __param_target
- source_labels:
- __param_target
target_label: target
- target_label: __address__
replacement: 127.0.0.1:9115
+139
View File
@@ -0,0 +1,139 @@
#!/usr/bin/env python3
# Генератор дашборда Grafana "GoToSocial" — по стилю vesti.json
# Результат: grafana/dashboards/gotosocial/gotosocial.json
import json, os, uuid
DS = {"type": "prometheus", "uid": "Prometheus"}
F = "bigbox:9464"
def tgt(expr, legend="", ref="A"):
return {"datasource": DS, "expr": expr, "legendFormat": legend, "refId": ref}
def stat_panel(title, expr, x, y, w=4, h=3, mapping=None, unit="short", thresholds=None):
if mapping is None:
mapping = {"0": {"color": "red", "text": "DOWN"}, "1": {"color": "green", "text": "UP"}}
if thresholds is None:
thresholds = {"mode": "absolute", "steps": [
{"color": "red", "value": None}, {"color": "green", "value": 1}]}
return {
"datasource": DS,
"fieldConfig": {"defaults": {
"color": {"mode": "thresholds"},
"mappings": [{"options": mapping, "type": "value"}],
"thresholds": thresholds, "unit": unit},
"overrides": []},
"gridPos": {"h": h, "w": w, "x": x, "y": y},
"id": uuid.uuid4().int & 0xFFFF,
"options": {
"colorMode": "background", "graphMode": "none", "justifyMode": "auto",
"orientation": "auto",
"reduceOptions": {"calcs": ["lastNotNull"], "fields": "", "values": False},
"textMode": "auto"},
"pluginVersion": "11.1.0",
"targets": [tgt(expr, "")],
"title": title, "type": "stat",
}
def ts_panel(title, exprs, x, y, w=12, h=4, unit="short", legend=False):
# exprs: list of (expr, legend) tuples
return {
"datasource": DS,
"fieldConfig": {"defaults": {
"color": {"mode": "palette-classic"},
"custom": {
"axisCenteredZero": False, "axisColorMode": "text", "axisLabel": "",
"axisPlacement": "auto", "drawStyle": "line", "fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {"legend": False, "tooltip": False, "viz": False},
"lineInterpolation": "linear", "lineWidth": 1, "pointSize": 5,
"scaleDistribution": {"type": "linear"}, "showPoints": "never",
"spanNulls": False, "stacking": {"group": "A", "mode": "none"},
"thresholdsStyle": {"mode": "off"}},
"mappings": [], "thresholds": {"mode": "absolute", "steps": [
{"color": "green", "value": None}]},
"unit": unit}, "overrides": []},
"gridPos": {"h": h, "w": w, "x": x, "y": y},
"id": uuid.uuid4().int & 0xFFFF,
"options": {"legend": {"calcs": [], "displayMode": "list", "placement": "bottom",
"showLegend": legend},
"tooltip": {"mode": "multi", "sort": "none"}},
"targets": [tgt(e, l) for e, l in exprs],
"title": title, "type": "timeseries",
}
def row(title, y):
return {"collapsed": False, "gridPos": {"h": 1, "w": 24, "x": 0, "y": y},
"id": uuid.uuid4().int & 0xFFFF, "panels": [], "title": title, "type": "row"}
panels = []
y = 0
# ---- Доступность ----
panels.append(row("Доступность", y)); y += 1
panels.append(stat_panel("Инстанс жив (total users)",
f'gotosocial_instance_total_users{{instance="{F}"}}', 0, y, w=8, h=3)); y += 3
# ---- Instance ----
panels.append(row("Инстанс", y)); y += 1
panels.append(ts_panel("Пользователи", [(f'gotosocial_instance_total_users{{instance="{F}"}}', "users")], 0, y, w=8, h=4))
panels.append(ts_panel("Статусы", [(f'gotosocial_instance_total_statuses{{instance="{F}"}}', "statuses")], 8, y, w=8, h=4))
panels.append(ts_panel("Федерирующиеся инстансы", [(f'gotosocial_instance_total_federating_instances{{instance="{F}"}}', "federating")], 16, y, w=8, h=4)); y += 4
# ---- Workers ----
panels.append(row("Воркеры", y)); y += 1
workers = ["client_api", "fedi_api", "processing", "delivery", "dereference", "webpush"]
count_exprs = [(f'gotosocial_workers_{w}_count{{instance="{F}"}}', w) for w in workers]
queue_exprs = [(f'gotosocial_workers_{w}_queue{{instance="{F}"}}', w) for w in workers]
panels.append(ts_panel("Воркеры: count (обработано)", count_exprs, 0, y, w=12, h=4, legend=True))
panels.append(ts_panel("Воркеры: queue (очередь)", queue_exprs, 12, y, w=12, h=4, legend=True)); y += 4
# ---- HTTP ----
panels.append(row("HTTP", y)); y += 1
panels.append(ts_panel("Запросы/с (все routes)",
[(f'rate(http_server_requests_total{{otel_scope_name="gin",instance="{F}"}})',
"req/s")], 0, y, w=8, h=4))
panels.append(ts_panel("Активные запросы",
[(f'http_server_requests_active{{instance="{F}"}}', "active")], 8, y, w=8, h=4))
panels.append(ts_panel("Длительность (p50)",
[(f'http_server_duration_milliseconds{{quantile="0.5",instance="{F}"}}', "p50 ms")],
16, y, w=8, h=4)); y += 4
# ---- SQL ----
panels.append(row("SQL", y)); y += 1
panels.append(ts_panel("SQL соединения",
[(f'go_sql_connections_open{{instance="{F}"}}', "open"),
(f'go_sql_connections_in_use{{instance="{F}"}}', "in_use")], 0, y, w=12, h=4, legend=True))
panels.append(ts_panel("SQL запросы (p50 ms)",
[(f'go_sql_query_timing_milliseconds{{quantile="0.5",instance="{F}"}}', "p50")],
12, y, w=12, h=4)); y += 4
# ---- Процесс ----
panels.append(row("Процесс", y)); y += 1
panels.append(ts_panel("RSS память (GB)",
[(f'process_resident_memory_bytes{{instance="{F}"}}/1024/1024/1024', "RSS GB")], 0, y, w=8, h=4))
panels.append(ts_panel("CPU (сек)",
[(f'process_cpu_seconds_total{{instance="{F}"}}', "cpu s")], 8, y, w=8, h=4))
panels.append(ts_panel("Открытые FD",
[(f'process_open_fds{{instance="{F}"}}', "fds")], 16, y, w=8, h=4)); y += 4
dash = {
"annotations": {"list": []},
"editable": True,
"fiscalYearStartMonth": 1,
"graphTooltip": 0,
"id": uuid.uuid4().int & 0xFFFF,
"links": [], "panels": panels, "refresh": "30s",
"tags": ["gotosocial"],
"title": "GoToSocial",
"type": "dashboard",
"uid": "gotosocial-main",
"version": 1,
"schemaVersion": 1,
}
outdir = os.path.join(os.path.dirname(__file__), "..", "grafana", "dashboards", "gotosocial")
os.makedirs(outdir, exist_ok=True)
out = os.path.join(outdir, "gotosocial.json")
with open(out, "w") as f:
json.dump(dash, f, ensure_ascii=False, indent=1)
print("OK:", out, len(panels), "panels")