Files
project-doc-processing/SKILL.md
T
2026-09-06 13:51:20 +00:00

103 lines
6.0 KiB
Markdown

---
name: project-doc-processing
description: "Process large engineering PDF/DWG sets one file at a time."
version: 1.0.0
author: Hermes Agent
license: MIT
platforms: [linux, macos, windows]
metadata:
hermes:
tags: [PDF, DWG, DrawIO, CAD, TokenEconomy, DocumentSet, Vision, Engineering]
related_skills: [ocr-and-documents, file-tree-catalog, pdf, docx]
---
# Project Document Set Processing (token-economy)
For processing a **large set of engineering/design documents** (СКС/ЛВС/КТСБ drawings, CAD-derived PDFs, architect sets) one file at a time, WITHOUT burning tokens by dumping whole docs into LLM context. Complements `ocr-and-documents` (which is single-PDF text extraction) and `file-tree-catalog` (which catalogs a tree). This skill is about the *processing loop* for a recurring, multi-session project.
## Core principle
The LLM should never read a full large PDF. Extract text to a file ONCE, then have code find structure and pull ONLY the slices worth human/LLM attention.
## Workflow
### 1. Find a clean list of files to process
Catalog the tree first (see `file-tree-catalog` skill). Filter junk (Thumbs.db, ~$*, .dwl) at ANY nesting depth. The user processes **one file at a time** — never batch-spend tokens across many files.
### 2. Get a PDF tool into an isolated venv
System Python is often PEP-668 externally-managed (`pip install` fails without `--break-system-packages`). Do NOT touch the Hermes venv or system site-packages.
```bash
python3 -m venv /path/to/project/.venv-pdf
/path/to/project/.venv-pdf/bin/pip install -q pymupdf
```
Use `pymupdf` first — instant, no models. Only reach for `marker-pdf` (OCR, ~3-5GB models) if there's NO text layer.
### 3. Detect the text layer before choosing extractor
Big engineering PDFs are often DWG renders / scans with a partial text layer. Probe cheaply first:
```python
import pymupdf
doc = pymupdf.open(path)
print(doc.page_count, "pages")
for i in range(min(6, doc.page_count)):
t = doc[i].get_text().strip()
print(f"page {i+1}: {len(t)} chars: {t[:80]!r}")
```
If pages have real text → pymupdf text extraction is enough. If empty → it's a scan: either render + vision (few pages, `doc[i].get_pixmap()`) or marker-pdf (bulk OCR). 20MB PDF with 90 pages / 270K chars of text layer still extracts fine — size alone does not mean "scanned".
### 4. Extract full text ONCE to a file, with page markers
```python
out = "name.txt"
with open(out, "w", encoding="utf-8") as f:
for i, page in enumerate(doc):
f.write(f"\n\n===== СТРАНИЦА {i+1} =====\n")
f.write(page.get_text())
```
This gives you a paginated, greppable corpus. Never paste the whole thing into the LLM.
### 5. Parse structure programmatically, don't read linearly
Use `execute_code`/`search_files` to find section markers (СПЕЦИФИКАЦИЯ, ВЕДОМОСТЬ, СХЕМА, ОБЩИЕ ДАННЫЕ, СОДЕРЖАНИЕ) and extract **only** the important ranges by line number via `read_file(offset=...)`. For spec tables, write a small parser that produces structured data (JSON) and diffs/condenses it — far cheaper than reading thousands of lines:
```python
# heuristic: обозначение = whitespace-free upper/digit/dash string >3 chars, name until next marker or a bare int (qty)
```
Then store the structured result (JSON) and write a **concise** summary .md for the record.
### 6. Record progress as editable status checkboxes
Keep a `_конспект.md` (summary) per processed file with:
- object / customer / GIP / revision
- doc composition (ведомость чертежей)
- the "what next / unread" list
- a `## Статус обработки` section with `[x]`/`[ ]` checkboxes (which files done, which pending)
This makes a multi-session project resumable: next session reads the status file and continues.
## Schematics → draw.io rebuild (design intent)
When the goal is to *reconstruct* a schematic (structural diagram, rack/ТШ layout, cable-run diagram), do NOT have the LLM write .drawio XML directly (it produces broken XML). Instead:
1. Render the schematic page to PNG: `page.get_pixmap(dpi=150).save("/tmp/page.png")`.
2. Send the PNG to a **vision model** describing nodes + edges.
3. Have vision output a **structured graph** (JSON: nodes with labels, edges), then serialize to `.drawio` (mxGraphModel) in code.
Honest capability table (set user expectations up front):
| Schematic type | draw.io rebuild |
|---|---|
| Structural diagram (nodes+links) | ✅ good — ideal candidate |
| Rack/ТШ layout (grid) | ✅ good (spec already parsed → grid) |
| Cable-run / fiber-splice diagram | ⚠️ medium — many tiny labels |
| Equipment/route layout plan | ⚠️ schematic only, NOT to scale |
| Isometric (axonometric) runs | ❌ poor — 3D iso doesn't fit draw.io |
Caveats to state:
- Local vision (qwen3-vl) reads diagrams but errs more on dense small-label sheets than cloud vision — local=privacy+free, cloud=accuracy. Hybrid allowed.
- Rendering DWG-render PDF → PNG loses geometry; output is a **semantic** diagram (right elements/links), not a pixel-exact copy.
- If pixel-exact matters, the real source is the `.dwg` CAD files (needs DXF/ODA conversion — heavier). For "understand & rebuild", vision is enough.
- Test on ONE simple representative sheet before scaling.
## Pitfalls
- PEP-668: use a dedicated venv, never pollute system/Hermes site-packages.
- A 20MB PDF is not necessarily a scan — ALWAYS probe text layer first.
- pymupdf text layer from DWG renders is partial and spuriously ordered (rack U-number grids, SFP/QSF labels leak into spec text) — the parser will catch noise; filter obvious non-spec tokens.
- Never dump an extracted .txt into the LLM whole. Parse to structure, then summarize.
- Preserve the original text file and JSON next to the summary .md — they're the auditable source for the summary.
## Verification
- After parsing a spec, spot-check 1-2 rows against the original PDF page (`read_file` at the known line) before trusting counts.
- After any vision→drawio sheet, open the .drawio (or re-parse) to confirm valid XML / expected node count.