mirror of
https://gitverse.ru/kpa39l/project-doc-processing.git
synced 2026-09-29 09:15:04 +00:00
Initial commit: Hermes skill project-doc-processing
This commit is contained in:
@@ -0,0 +1,103 @@
|
|||||||
|
---
|
||||||
|
name: project-doc-processing
|
||||||
|
description: "Process large engineering PDF/DWG sets one file at a time."
|
||||||
|
version: 1.0.0
|
||||||
|
author: Hermes Agent
|
||||||
|
license: MIT
|
||||||
|
platforms: [linux, macos, windows]
|
||||||
|
metadata:
|
||||||
|
hermes:
|
||||||
|
tags: [PDF, DWG, DrawIO, CAD, TokenEconomy, DocumentSet, Vision, Engineering]
|
||||||
|
related_skills: [ocr-and-documents, file-tree-catalog, pdf, docx]
|
||||||
|
---
|
||||||
|
|
||||||
|
# Project Document Set Processing (token-economy)
|
||||||
|
|
||||||
|
For processing a **large set of engineering/design documents** (СКС/ЛВС/КТСБ drawings, CAD-derived PDFs, architect sets) one file at a time, WITHOUT burning tokens by dumping whole docs into LLM context. Complements `ocr-and-documents` (which is single-PDF text extraction) and `file-tree-catalog` (which catalogs a tree). This skill is about the *processing loop* for a recurring, multi-session project.
|
||||||
|
|
||||||
|
## Core principle
|
||||||
|
The LLM should never read a full large PDF. Extract text to a file ONCE, then have code find structure and pull ONLY the slices worth human/LLM attention.
|
||||||
|
|
||||||
|
## Workflow
|
||||||
|
|
||||||
|
### 1. Find a clean list of files to process
|
||||||
|
Catalog the tree first (see `file-tree-catalog` skill). Filter junk (Thumbs.db, ~$*, .dwl) at ANY nesting depth. The user processes **one file at a time** — never batch-spend tokens across many files.
|
||||||
|
|
||||||
|
### 2. Get a PDF tool into an isolated venv
|
||||||
|
System Python is often PEP-668 externally-managed (`pip install` fails without `--break-system-packages`). Do NOT touch the Hermes venv or system site-packages.
|
||||||
|
```bash
|
||||||
|
python3 -m venv /path/to/project/.venv-pdf
|
||||||
|
/path/to/project/.venv-pdf/bin/pip install -q pymupdf
|
||||||
|
```
|
||||||
|
Use `pymupdf` first — instant, no models. Only reach for `marker-pdf` (OCR, ~3-5GB models) if there's NO text layer.
|
||||||
|
|
||||||
|
### 3. Detect the text layer before choosing extractor
|
||||||
|
Big engineering PDFs are often DWG renders / scans with a partial text layer. Probe cheaply first:
|
||||||
|
```python
|
||||||
|
import pymupdf
|
||||||
|
doc = pymupdf.open(path)
|
||||||
|
print(doc.page_count, "pages")
|
||||||
|
for i in range(min(6, doc.page_count)):
|
||||||
|
t = doc[i].get_text().strip()
|
||||||
|
print(f"page {i+1}: {len(t)} chars: {t[:80]!r}")
|
||||||
|
```
|
||||||
|
If pages have real text → pymupdf text extraction is enough. If empty → it's a scan: either render + vision (few pages, `doc[i].get_pixmap()`) or marker-pdf (bulk OCR). 20MB PDF with 90 pages / 270K chars of text layer still extracts fine — size alone does not mean "scanned".
|
||||||
|
|
||||||
|
### 4. Extract full text ONCE to a file, with page markers
|
||||||
|
```python
|
||||||
|
out = "name.txt"
|
||||||
|
with open(out, "w", encoding="utf-8") as f:
|
||||||
|
for i, page in enumerate(doc):
|
||||||
|
f.write(f"\n\n===== СТРАНИЦА {i+1} =====\n")
|
||||||
|
f.write(page.get_text())
|
||||||
|
```
|
||||||
|
This gives you a paginated, greppable corpus. Never paste the whole thing into the LLM.
|
||||||
|
|
||||||
|
### 5. Parse structure programmatically, don't read linearly
|
||||||
|
Use `execute_code`/`search_files` to find section markers (СПЕЦИФИКАЦИЯ, ВЕДОМОСТЬ, СХЕМА, ОБЩИЕ ДАННЫЕ, СОДЕРЖАНИЕ) and extract **only** the important ranges by line number via `read_file(offset=...)`. For spec tables, write a small parser that produces structured data (JSON) and diffs/condenses it — far cheaper than reading thousands of lines:
|
||||||
|
```python
|
||||||
|
# heuristic: обозначение = whitespace-free upper/digit/dash string >3 chars, name until next marker or a bare int (qty)
|
||||||
|
```
|
||||||
|
Then store the structured result (JSON) and write a **concise** summary .md for the record.
|
||||||
|
|
||||||
|
### 6. Record progress as editable status checkboxes
|
||||||
|
Keep a `_конспект.md` (summary) per processed file with:
|
||||||
|
- object / customer / GIP / revision
|
||||||
|
- doc composition (ведомость чертежей)
|
||||||
|
- the "what next / unread" list
|
||||||
|
- a `## Статус обработки` section with `[x]`/`[ ]` checkboxes (which files done, which pending)
|
||||||
|
This makes a multi-session project resumable: next session reads the status file and continues.
|
||||||
|
|
||||||
|
## Schematics → draw.io rebuild (design intent)
|
||||||
|
|
||||||
|
When the goal is to *reconstruct* a schematic (structural diagram, rack/ТШ layout, cable-run diagram), do NOT have the LLM write .drawio XML directly (it produces broken XML). Instead:
|
||||||
|
|
||||||
|
1. Render the schematic page to PNG: `page.get_pixmap(dpi=150).save("/tmp/page.png")`.
|
||||||
|
2. Send the PNG to a **vision model** describing nodes + edges.
|
||||||
|
3. Have vision output a **structured graph** (JSON: nodes with labels, edges), then serialize to `.drawio` (mxGraphModel) in code.
|
||||||
|
|
||||||
|
Honest capability table (set user expectations up front):
|
||||||
|
| Schematic type | draw.io rebuild |
|
||||||
|
|---|---|
|
||||||
|
| Structural diagram (nodes+links) | ✅ good — ideal candidate |
|
||||||
|
| Rack/ТШ layout (grid) | ✅ good (spec already parsed → grid) |
|
||||||
|
| Cable-run / fiber-splice diagram | ⚠️ medium — many tiny labels |
|
||||||
|
| Equipment/route layout plan | ⚠️ schematic only, NOT to scale |
|
||||||
|
| Isometric (axonometric) runs | ❌ poor — 3D iso doesn't fit draw.io |
|
||||||
|
|
||||||
|
Caveats to state:
|
||||||
|
- Local vision (qwen3-vl) reads diagrams but errs more on dense small-label sheets than cloud vision — local=privacy+free, cloud=accuracy. Hybrid allowed.
|
||||||
|
- Rendering DWG-render PDF → PNG loses geometry; output is a **semantic** diagram (right elements/links), not a pixel-exact copy.
|
||||||
|
- If pixel-exact matters, the real source is the `.dwg` CAD files (needs DXF/ODA conversion — heavier). For "understand & rebuild", vision is enough.
|
||||||
|
- Test on ONE simple representative sheet before scaling.
|
||||||
|
|
||||||
|
## Pitfalls
|
||||||
|
- PEP-668: use a dedicated venv, never pollute system/Hermes site-packages.
|
||||||
|
- A 20MB PDF is not necessarily a scan — ALWAYS probe text layer first.
|
||||||
|
- pymupdf text layer from DWG renders is partial and spuriously ordered (rack U-number grids, SFP/QSF labels leak into spec text) — the parser will catch noise; filter obvious non-spec tokens.
|
||||||
|
- Never dump an extracted .txt into the LLM whole. Parse to structure, then summarize.
|
||||||
|
- Preserve the original text file and JSON next to the summary .md — they're the auditable source for the summary.
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
- After parsing a spec, spot-check 1-2 rows against the original PDF page (`read_file` at the known line) before trusting counts.
|
||||||
|
- After any vision→drawio sheet, open the .drawio (or re-parse) to confirm valid XML / expected node count.
|
||||||
@@ -0,0 +1,38 @@
|
|||||||
|
# Легенда обозначений розеток — комплект СКС «Винный город» (блок A)
|
||||||
|
|
||||||
|
Источник: свежий лист шкафа /mnt/vinogorod/ИТ/trash/2026-09-05_08-44-17.png (нижняя часть листа, таблица «Условные графические обозначения»). Прочитано vision-моделью (polza deepseek-vision) по кропу нижней полосы.
|
||||||
|
|
||||||
|
ВАЖНО: легенда может отличаться между листами комплекта — перед импортом сверять с каждым листом, не копировать вслепую.
|
||||||
|
|
||||||
|
## Маркировка розеток/назначений (таблица «Маркировка патч-панелей»)
|
||||||
|
| Код | Назначение |
|
||||||
|
|---|---|
|
||||||
|
| (A) | Гнездо доступа Wi-Fi |
|
||||||
|
| (V) | IP-камера видеонаблюдения |
|
||||||
|
| (S) | IP-контроллер (КДУ/СОТС) |
|
||||||
|
| (I) | Гнездо доступа ИБП |
|
||||||
|
| (D) | Гнездо IP-телефонии |
|
||||||
|
| (T) | Телефон |
|
||||||
|
| (Д) | Инфракрасный (сетевой) датчик периметра |
|
||||||
|
| (СР) | Вызывная панель |
|
||||||
|
|
||||||
|
## Линии (условные графические обозначения)
|
||||||
|
- сплошная линия — кабель электрический (СВ)
|
||||||
|
- двойная сплошная — кабель электрический коммутационный с оборудованием
|
||||||
|
- пунктирная — Ethernet UTP cat 6a
|
||||||
|
- волнистая — Ethernet UTP cat 6a внешней прокладки
|
||||||
|
- тройная вертикальная — патч-корд
|
||||||
|
- двойная вертикальная — патч-корд оптический
|
||||||
|
- длинный штрих — DAC-кабель
|
||||||
|
|
||||||
|
## Маркировка оборудования в шкафу
|
||||||
|
- SW# — коммутатор доступа; SWA — коммутатор агрегации
|
||||||
|
- PR / PR1 / PR2 — устройство распределения / управления
|
||||||
|
- патч-панель, розетка RJ-45, сетевой шкаф
|
||||||
|
|
||||||
|
## Замечание по вчерашнему разбору (2026-09-04)
|
||||||
|
На первом скриншоте (PP.E.1/А.1) модель дала коды (v), (d), (a), (w). По этой легенде:
|
||||||
|
- (v) в легенде нет — есть (V) = IP-камера
|
||||||
|
- (a) == (A) Wi-Fi, (d) == (D) IP-телефония вероятно
|
||||||
|
- (w) в легенде отсутствует → перепроверить на исходном листе
|
||||||
|
Не утверждать расшифровку, пока не сверено с листом.
|
||||||
@@ -0,0 +1,50 @@
|
|||||||
|
# Spec/структура parser — engineering PDF text layer
|
||||||
|
|
||||||
|
Pattern that worked on the wine-city СКС project (Спецификация оборудования ТШ шкафов из 90-page DWG-render PDF, 384KB extracted text). Reusable for any "оборудование по шкафам/узлам" spec table whose text layer is radially extracted.
|
||||||
|
|
||||||
|
## Goal
|
||||||
|
Turn the extracted .txt (with `===== СТРАНИЦА N =====` page markers) into structured items WITHOUT reading 2000 lines/spec into the LLM.
|
||||||
|
|
||||||
|
## Heuristic-обозначение parser
|
||||||
|
```python
|
||||||
|
import re
|
||||||
|
def parse_spec(lines, start, end):
|
||||||
|
"""items = [ {oboz, name, qty}, ... ]"""
|
||||||
|
items, i = [], start
|
||||||
|
while i < end:
|
||||||
|
s = lines[i-1].strip()
|
||||||
|
if not s:
|
||||||
|
i += 1; continue
|
||||||
|
# обозначение: one token, upper/alnum/dash/dot, >3 chars, not starting with digit
|
||||||
|
if re.fullmatch(r'[A-ZА-Я0-9][A-ZА-Я0-9\-.]+', s) and len(s) > 3 and not s[0].isdigit():
|
||||||
|
cur = {"oboz": s, "name": [], "qty": None}
|
||||||
|
j = i + 1; name_lines = []
|
||||||
|
while j < end and len(name_lines) < 8: # cap: наименование ≤8 строк
|
||||||
|
t = lines[j-1].strip()
|
||||||
|
if re.fullmatch(r'[A-ZА-Я0-9][A-ZА-Я0-9\-.]+', t) and len(t) > 3 and not t[0].isdigit():
|
||||||
|
break # next обозначение
|
||||||
|
if re.fullmatch(r'\d+', t):
|
||||||
|
cur["qty"] = int(t); j += 1; break # qty = bare int
|
||||||
|
if t and not re.fullmatch(r'\d+', t):
|
||||||
|
name_lines.append(t)
|
||||||
|
j += 1
|
||||||
|
cur["name"] = " ".join(name_lines)[:140]
|
||||||
|
items.append(cur); i = j
|
||||||
|
else:
|
||||||
|
i += 1
|
||||||
|
return items
|
||||||
|
```
|
||||||
|
|
||||||
|
## Known noise to filter (DWG-render text layer)
|
||||||
|
- Rack U-number grids: long runs of bare ints (1..47) — compiler may capture as qty=None items; ignore `qty is None` rows unless they carry a real name (they won't: name stays empty).
|
||||||
|
- SFP1..4 / QSFP1..6 / "резерв" / "220V Port 1..." label rows leak in from switch faceplate text. Drop tokens in {SFP*, QSFP*} and empty-name items.
|
||||||
|
- Real spec items are the ones with BOTH a multi-char обозначение AND a non-empty name AND a qty — keep those as the equipment table.
|
||||||
|
|
||||||
|
## Diff across nearly-identical sections (e.g. ТШ-A.1..A.4)
|
||||||
|
Since соседние spec-блоки почти одинаковы, don't re-read each: parse all to JSON, then compare counts per (oboz, name) to spot the deltas (A.1 vs A.2 vs A.3 vs A.4) and only narrate what differs.
|
||||||
|
|
||||||
|
## Finding section boundaries without full reads
|
||||||
|
Scan the .txt lines for header regexes once (cheap): `СПЕЦИФИКАЦИЯ`, `ВЕДОМОСТЬ`, `СХЕМА`, `ОБЩИЕ ДАННЫЕ`, `СОДЕРЖАНИЕ`; record their line numbers, then `read_file(offset=…)` ONLY the interesting ranges.
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
Before trusting counts, `read_file` 1-2 rows at their known line and match против исходной таблицы on the actual PDF page (`page.get_text()` for the page). Spec numbers are consumable-accurate (закупочные количества) — worth the check.
|
||||||
@@ -0,0 +1,23 @@
|
|||||||
|
# Vision-reading CAD screenshots — crop & zoom technique
|
||||||
|
|
||||||
|
When the source is a **PNG screenshot of a drawing sheet** (rack/ТШ schematic, port→socket table) instead of a PDF, the vision model reads it directly — but dense sheets need the crop-zoom technique. This reference captures the working pattern from the vinogorod СКС screenshot sessions.
|
||||||
|
|
||||||
|
## Workflow
|
||||||
|
1. Probe the sheet size first (`PIL.Image.open().size`). Large sheets (~2300px) read well whole; small ones (~500px) need crops.
|
||||||
|
2. Cut the sheet into bands (top/mid/bottom strips, left/right halves, quarters) with `img.crop()`, upscale ×2-4 with `Image.LANCZOS`, save to /tmp.
|
||||||
|
3. Send each crop with a **narrow, fact-only question**: "перечисли все подписи по одной на строку, ничего не пропускай". Explicit formatted-single-line output beats "опиши лист".
|
||||||
|
4. **Timeout retry pattern**: cloud vision (polza deepseek-vision) times out on dense crops — retry the SAME region as a smaller/narrower crop, or split the question. A timeout is not a dead end; the region is readable in pieces.
|
||||||
|
5. Cross-check across crops of the same sheet (top = device labels, bottom = legend/tables); merge answers. Vision text is imperfect — treat "(v) vs (V)", "Swd vs SWd" as the same until proven otherwise.
|
||||||
|
6. On engineer sheet sets, the **legend/«Условные графические обозначения» sits at the BOTTOM of each rack sheet** — crop the bottom strip for socket-code decoding. Legends can differ BETWEEN sheets of one set — verify per sheet, never copy from a previous one.
|
||||||
|
|
||||||
|
## Proven example (2026-09-05)
|
||||||
|
Sheet `/mnt/vinogorod/ИТ/trash/2026-09-05_08-44-17.png` (2299×1789): bottom strip crop (55-100% height, ×1.6 LANCZOS) exposed the legend tables «Условные графические обозначения», «Маркировка патч-панелей», «Маркировка оборудования в шкафу». Full decode in `sks-rack-legend.md`.
|
||||||
|
|
||||||
|
## Model quirks
|
||||||
|
- polza deepseek-vision (~0.1₽/request): fine on large sheets; times out on dense crops → narrow the crop, not the timeout.
|
||||||
|
- Small sheets (~500px wide) read worse than large ones; always upscale ×3+ before sending.
|
||||||
|
- Split one big "read everything" question into several per-region questions — the model returns cleaner single-line lists.
|
||||||
|
|
||||||
|
## Cross-links
|
||||||
|
- Legend decode: `sks-rack-legend.md`
|
||||||
|
- Spec/table parsing (PDFs): `spec-parser-pattern.md`
|
||||||
Reference in New Issue
Block a user