Initial commit: Hermes skill project-doc-processing

This commit is contained in:
estorozhenko
2026-09-06 13:51:20 +00:00
commit d53e4b88f6
4 changed files with 214 additions and 0 deletions
+103
View File
@@ -0,0 +1,103 @@
---
name: project-doc-processing
description: "Process large engineering PDF/DWG sets one file at a time."
version: 1.0.0
author: Hermes Agent
license: MIT
platforms: [linux, macos, windows]
metadata:
hermes:
tags: [PDF, DWG, DrawIO, CAD, TokenEconomy, DocumentSet, Vision, Engineering]
related_skills: [ocr-and-documents, file-tree-catalog, pdf, docx]
---
# Project Document Set Processing (token-economy)
For processing a **large set of engineering/design documents** (СКС/ЛВС/КТСБ drawings, CAD-derived PDFs, architect sets) one file at a time, WITHOUT burning tokens by dumping whole docs into LLM context. Complements `ocr-and-documents` (which is single-PDF text extraction) and `file-tree-catalog` (which catalogs a tree). This skill is about the *processing loop* for a recurring, multi-session project.
## Core principle
The LLM should never read a full large PDF. Extract text to a file ONCE, then have code find structure and pull ONLY the slices worth human/LLM attention.
## Workflow
### 1. Find a clean list of files to process
Catalog the tree first (see `file-tree-catalog` skill). Filter junk (Thumbs.db, ~$*, .dwl) at ANY nesting depth. The user processes **one file at a time** — never batch-spend tokens across many files.
### 2. Get a PDF tool into an isolated venv
System Python is often PEP-668 externally-managed (`pip install` fails without `--break-system-packages`). Do NOT touch the Hermes venv or system site-packages.
```bash
python3 -m venv /path/to/project/.venv-pdf
/path/to/project/.venv-pdf/bin/pip install -q pymupdf
```
Use `pymupdf` first — instant, no models. Only reach for `marker-pdf` (OCR, ~3-5GB models) if there's NO text layer.
### 3. Detect the text layer before choosing extractor
Big engineering PDFs are often DWG renders / scans with a partial text layer. Probe cheaply first:
```python
import pymupdf
doc = pymupdf.open(path)
print(doc.page_count, "pages")
for i in range(min(6, doc.page_count)):
t = doc[i].get_text().strip()
print(f"page {i+1}: {len(t)} chars: {t[:80]!r}")
```
If pages have real text → pymupdf text extraction is enough. If empty → it's a scan: either render + vision (few pages, `doc[i].get_pixmap()`) or marker-pdf (bulk OCR). 20MB PDF with 90 pages / 270K chars of text layer still extracts fine — size alone does not mean "scanned".
### 4. Extract full text ONCE to a file, with page markers
```python
out = "name.txt"
with open(out, "w", encoding="utf-8") as f:
for i, page in enumerate(doc):
f.write(f"\n\n===== СТРАНИЦА {i+1} =====\n")
f.write(page.get_text())
```
This gives you a paginated, greppable corpus. Never paste the whole thing into the LLM.
### 5. Parse structure programmatically, don't read linearly
Use `execute_code`/`search_files` to find section markers (СПЕЦИФИКАЦИЯ, ВЕДОМОСТЬ, СХЕМА, ОБЩИЕ ДАННЫЕ, СОДЕРЖАНИЕ) and extract **only** the important ranges by line number via `read_file(offset=...)`. For spec tables, write a small parser that produces structured data (JSON) and diffs/condenses it — far cheaper than reading thousands of lines:
```python
# heuristic: обозначение = whitespace-free upper/digit/dash string >3 chars, name until next marker or a bare int (qty)
```
Then store the structured result (JSON) and write a **concise** summary .md for the record.
### 6. Record progress as editable status checkboxes
Keep a `_конспект.md` (summary) per processed file with:
- object / customer / GIP / revision
- doc composition (ведомость чертежей)
- the "what next / unread" list
- a `## Статус обработки` section with `[x]`/`[ ]` checkboxes (which files done, which pending)
This makes a multi-session project resumable: next session reads the status file and continues.
## Schematics → draw.io rebuild (design intent)
When the goal is to *reconstruct* a schematic (structural diagram, rack/ТШ layout, cable-run diagram), do NOT have the LLM write .drawio XML directly (it produces broken XML). Instead:
1. Render the schematic page to PNG: `page.get_pixmap(dpi=150).save("/tmp/page.png")`.
2. Send the PNG to a **vision model** describing nodes + edges.
3. Have vision output a **structured graph** (JSON: nodes with labels, edges), then serialize to `.drawio` (mxGraphModel) in code.
Honest capability table (set user expectations up front):
| Schematic type | draw.io rebuild |
|---|---|
| Structural diagram (nodes+links) | ✅ good — ideal candidate |
| Rack/ТШ layout (grid) | ✅ good (spec already parsed → grid) |
| Cable-run / fiber-splice diagram | ⚠️ medium — many tiny labels |
| Equipment/route layout plan | ⚠️ schematic only, NOT to scale |
| Isometric (axonometric) runs | ❌ poor — 3D iso doesn't fit draw.io |
Caveats to state:
- Local vision (qwen3-vl) reads diagrams but errs more on dense small-label sheets than cloud vision — local=privacy+free, cloud=accuracy. Hybrid allowed.
- Rendering DWG-render PDF → PNG loses geometry; output is a **semantic** diagram (right elements/links), not a pixel-exact copy.
- If pixel-exact matters, the real source is the `.dwg` CAD files (needs DXF/ODA conversion — heavier). For "understand & rebuild", vision is enough.
- Test on ONE simple representative sheet before scaling.
## Pitfalls
- PEP-668: use a dedicated venv, never pollute system/Hermes site-packages.
- A 20MB PDF is not necessarily a scan — ALWAYS probe text layer first.
- pymupdf text layer from DWG renders is partial and spuriously ordered (rack U-number grids, SFP/QSF labels leak into spec text) — the parser will catch noise; filter obvious non-spec tokens.
- Never dump an extracted .txt into the LLM whole. Parse to structure, then summarize.
- Preserve the original text file and JSON next to the summary .md — they're the auditable source for the summary.
## Verification
- After parsing a spec, spot-check 1-2 rows against the original PDF page (`read_file` at the known line) before trusting counts.
- After any vision→drawio sheet, open the .drawio (or re-parse) to confirm valid XML / expected node count.
+38
View File
@@ -0,0 +1,38 @@
# Легенда обозначений розеток — комплект СКС «Винный город» (блок A)
Источник: свежий лист шкафа /mnt/vinogorod/ИТ/trash/2026-09-05_08-44-17.png (нижняя часть листа, таблица «Условные графические обозначения»). Прочитано vision-моделью (polza deepseek-vision) по кропу нижней полосы.
ВАЖНО: легенда может отличаться между листами комплекта — перед импортом сверять с каждым листом, не копировать вслепую.
## Маркировка розеток/назначений (таблица «Маркировка патч-панелей»)
| Код | Назначение |
|---|---|
| (A) | Гнездо доступа Wi-Fi |
| (V) | IP-камера видеонаблюдения |
| (S) | IP-контроллер (КДУ/СОТС) |
| (I) | Гнездо доступа ИБП |
| (D) | Гнездо IP-телефонии |
| (T) | Телефон |
| (Д) | Инфракрасный (сетевой) датчик периметра |
| (СР) | Вызывная панель |
## Линии (условные графические обозначения)
- сплошная линия — кабель электрический (СВ)
- двойная сплошная — кабель электрический коммутационный с оборудованием
- пунктирная — Ethernet UTP cat 6a
- волнистая — Ethernet UTP cat 6a внешней прокладки
- тройная вертикальная — патч-корд
- двойная вертикальная — патч-корд оптический
- длинный штрих — DAC-кабель
## Маркировка оборудования в шкафу
- SW# — коммутатор доступа; SWA — коммутатор агрегации
- PR / PR1 / PR2 — устройство распределения / управления
- патч-панель, розетка RJ-45, сетевой шкаф
## Замечание по вчерашнему разбору (2026-09-04)
На первом скриншоте (PP.E.1/А.1) модель дала коды (v), (d), (a), (w). По этой легенде:
- (v) в легенде нет — есть (V) = IP-камера
- (a) == (A) Wi-Fi, (d) == (D) IP-телефония вероятно
- (w) в легенде отсутствует → перепроверить на исходном листе
Не утверждать расшифровку, пока не сверено с листом.
+50
View File
@@ -0,0 +1,50 @@
# Spec/структура parser — engineering PDF text layer
Pattern that worked on the wine-city СКС project (Спецификация оборудования ТШ шкафов из 90-page DWG-render PDF, 384KB extracted text). Reusable for any "оборудование по шкафам/узлам" spec table whose text layer is radially extracted.
## Goal
Turn the extracted .txt (with `===== СТРАНИЦА N =====` page markers) into structured items WITHOUT reading 2000 lines/spec into the LLM.
## Heuristic-обозначение parser
```python
import re
def parse_spec(lines, start, end):
"""items = [ {oboz, name, qty}, ... ]"""
items, i = [], start
while i < end:
s = lines[i-1].strip()
if not s:
i += 1; continue
# обозначение: one token, upper/alnum/dash/dot, >3 chars, not starting with digit
if re.fullmatch(r'[A-ZА-Я0-9][A-ZА-Я0-9\-.]+', s) and len(s) > 3 and not s[0].isdigit():
cur = {"oboz": s, "name": [], "qty": None}
j = i + 1; name_lines = []
while j < end and len(name_lines) < 8: # cap: наименование ≤8 строк
t = lines[j-1].strip()
if re.fullmatch(r'[A-ZА-Я0-9][A-ZА-Я0-9\-.]+', t) and len(t) > 3 and not t[0].isdigit():
break # next обозначение
if re.fullmatch(r'\d+', t):
cur["qty"] = int(t); j += 1; break # qty = bare int
if t and not re.fullmatch(r'\d+', t):
name_lines.append(t)
j += 1
cur["name"] = " ".join(name_lines)[:140]
items.append(cur); i = j
else:
i += 1
return items
```
## Known noise to filter (DWG-render text layer)
- Rack U-number grids: long runs of bare ints (1..47) — compiler may capture as qty=None items; ignore `qty is None` rows unless they carry a real name (they won't: name stays empty).
- SFP1..4 / QSFP1..6 / "резерв" / "220V Port 1..." label rows leak in from switch faceplate text. Drop tokens in {SFP*, QSFP*} and empty-name items.
- Real spec items are the ones with BOTH a multi-char обозначение AND a non-empty name AND a qty — keep those as the equipment table.
## Diff across nearly-identical sections (e.g. ТШ-A.1..A.4)
Since соседние spec-блоки почти одинаковы, don't re-read each: parse all to JSON, then compare counts per (oboz, name) to spot the deltas (A.1 vs A.2 vs A.3 vs A.4) and only narrate what differs.
## Finding section boundaries without full reads
Scan the .txt lines for header regexes once (cheap): `СПЕЦИФИКАЦИЯ`, `ВЕДОМОСТЬ`, `СХЕМА`, `ОБЩИЕ ДАННЫЕ`, `СОДЕРЖАНИЕ`; record their line numbers, then `read_file(offset=…)` ONLY the interesting ranges.
## Verification
Before trusting counts, `read_file` 1-2 rows at their known line and match против исходной таблицы on the actual PDF page (`page.get_text()` for the page). Spec numbers are consumable-accurate (закупочные количества) — worth the check.
+23
View File
@@ -0,0 +1,23 @@
# Vision-reading CAD screenshots — crop & zoom technique
When the source is a **PNG screenshot of a drawing sheet** (rack/ТШ schematic, port→socket table) instead of a PDF, the vision model reads it directly — but dense sheets need the crop-zoom technique. This reference captures the working pattern from the vinogorod СКС screenshot sessions.
## Workflow
1. Probe the sheet size first (`PIL.Image.open().size`). Large sheets (~2300px) read well whole; small ones (~500px) need crops.
2. Cut the sheet into bands (top/mid/bottom strips, left/right halves, quarters) with `img.crop()`, upscale ×2-4 with `Image.LANCZOS`, save to /tmp.
3. Send each crop with a **narrow, fact-only question**: "перечисли все подписи по одной на строку, ничего не пропускай". Explicit formatted-single-line output beats "опиши лист".
4. **Timeout retry pattern**: cloud vision (polza deepseek-vision) times out on dense crops — retry the SAME region as a smaller/narrower crop, or split the question. A timeout is not a dead end; the region is readable in pieces.
5. Cross-check across crops of the same sheet (top = device labels, bottom = legend/tables); merge answers. Vision text is imperfect — treat "(v) vs (V)", "Swd vs SWd" as the same until proven otherwise.
6. On engineer sheet sets, the **legend/«Условные графические обозначения» sits at the BOTTOM of each rack sheet** — crop the bottom strip for socket-code decoding. Legends can differ BETWEEN sheets of one set — verify per sheet, never copy from a previous one.
## Proven example (2026-09-05)
Sheet `/mnt/vinogorod/ИТ/trash/2026-09-05_08-44-17.png` (2299×1789): bottom strip crop (55-100% height, ×1.6 LANCZOS) exposed the legend tables «Условные графические обозначения», «Маркировка патч-панелей», «Маркировка оборудования в шкафу». Full decode in `sks-rack-legend.md`.
## Model quirks
- polza deepseek-vision (~0.1₽/request): fine on large sheets; times out on dense crops → narrow the crop, not the timeout.
- Small sheets (~500px wide) read worse than large ones; always upscale ×3+ before sending.
- Split one big "read everything" question into several per-region questions — the model returns cleaner single-line lists.
## Cross-links
- Legend decode: `sks-rack-legend.md`
- Spec/table parsing (PDFs): `spec-parser-pattern.md`