--- name: project-doc-processing description: "Process large engineering PDF/DWG sets one file at a time." version: 1.0.0 author: Hermes Agent license: MIT platforms: [linux, macos, windows] metadata: hermes: tags: [PDF, DWG, DrawIO, CAD, TokenEconomy, DocumentSet, Vision, Engineering] related_skills: [ocr-and-documents, file-tree-catalog, pdf, docx] --- # Project Document Set Processing (token-economy) For processing a **large set of engineering/design documents** (СКС/ЛВС/КТСБ drawings, CAD-derived PDFs, architect sets) one file at a time, WITHOUT burning tokens by dumping whole docs into LLM context. Complements `ocr-and-documents` (which is single-PDF text extraction) and `file-tree-catalog` (which catalogs a tree). This skill is about the *processing loop* for a recurring, multi-session project. ## Core principle The LLM should never read a full large PDF. Extract text to a file ONCE, then have code find structure and pull ONLY the slices worth human/LLM attention. ## Workflow ### 1. Find a clean list of files to process Catalog the tree first (see `file-tree-catalog` skill). Filter junk (Thumbs.db, ~$*, .dwl) at ANY nesting depth. The user processes **one file at a time** — never batch-spend tokens across many files. ### 2. Get a PDF tool into an isolated venv System Python is often PEP-668 externally-managed (`pip install` fails without `--break-system-packages`). Do NOT touch the Hermes venv or system site-packages. ```bash python3 -m venv /path/to/project/.venv-pdf /path/to/project/.venv-pdf/bin/pip install -q pymupdf ``` Use `pymupdf` first — instant, no models. Only reach for `marker-pdf` (OCR, ~3-5GB models) if there's NO text layer. ### 3. Detect the text layer before choosing extractor Big engineering PDFs are often DWG renders / scans with a partial text layer. Probe cheaply first: ```python import pymupdf doc = pymupdf.open(path) print(doc.page_count, "pages") for i in range(min(6, doc.page_count)): t = doc[i].get_text().strip() print(f"page {i+1}: {len(t)} chars: {t[:80]!r}") ``` If pages have real text → pymupdf text extraction is enough. If empty → it's a scan: either render + vision (few pages, `doc[i].get_pixmap()`) or marker-pdf (bulk OCR). 20MB PDF with 90 pages / 270K chars of text layer still extracts fine — size alone does not mean "scanned". ### 4. Extract full text ONCE to a file, with page markers ```python out = "name.txt" with open(out, "w", encoding="utf-8") as f: for i, page in enumerate(doc): f.write(f"\n\n===== СТРАНИЦА {i+1} =====\n") f.write(page.get_text()) ``` This gives you a paginated, greppable corpus. Never paste the whole thing into the LLM. ### 5. Parse structure programmatically, don't read linearly Use `execute_code`/`search_files` to find section markers (СПЕЦИФИКАЦИЯ, ВЕДОМОСТЬ, СХЕМА, ОБЩИЕ ДАННЫЕ, СОДЕРЖАНИЕ) and extract **only** the important ranges by line number via `read_file(offset=...)`. For spec tables, write a small parser that produces structured data (JSON) and diffs/condenses it — far cheaper than reading thousands of lines: ```python # heuristic: обозначение = whitespace-free upper/digit/dash string >3 chars, name until next marker or a bare int (qty) ``` Then store the structured result (JSON) and write a **concise** summary .md for the record. ### 6. Record progress as editable status checkboxes Keep a `_конспект.md` (summary) per processed file with: - object / customer / GIP / revision - doc composition (ведомость чертежей) - the "what next / unread" list - a `## Статус обработки` section with `[x]`/`[ ]` checkboxes (which files done, which pending) This makes a multi-session project resumable: next session reads the status file and continues. ## Schematics → draw.io rebuild (design intent) When the goal is to *reconstruct* a schematic (structural diagram, rack/ТШ layout, cable-run diagram), do NOT have the LLM write .drawio XML directly (it produces broken XML). Instead: 1. Render the schematic page to PNG: `page.get_pixmap(dpi=150).save("/tmp/page.png")`. 2. Send the PNG to a **vision model** describing nodes + edges. 3. Have vision output a **structured graph** (JSON: nodes with labels, edges), then serialize to `.drawio` (mxGraphModel) in code. Honest capability table (set user expectations up front): | Schematic type | draw.io rebuild | |---|---| | Structural diagram (nodes+links) | ✅ good — ideal candidate | | Rack/ТШ layout (grid) | ✅ good (spec already parsed → grid) | | Cable-run / fiber-splice diagram | ⚠️ medium — many tiny labels | | Equipment/route layout plan | ⚠️ schematic only, NOT to scale | | Isometric (axonometric) runs | ❌ poor — 3D iso doesn't fit draw.io | Caveats to state: - Local vision (qwen3-vl) reads diagrams but errs more on dense small-label sheets than cloud vision — local=privacy+free, cloud=accuracy. Hybrid allowed. - Rendering DWG-render PDF → PNG loses geometry; output is a **semantic** diagram (right elements/links), not a pixel-exact copy. - If pixel-exact matters, the real source is the `.dwg` CAD files (needs DXF/ODA conversion — heavier). For "understand & rebuild", vision is enough. - Test on ONE simple representative sheet before scaling. ## Pitfalls - PEP-668: use a dedicated venv, never pollute system/Hermes site-packages. - A 20MB PDF is not necessarily a scan — ALWAYS probe text layer first. - pymupdf text layer from DWG renders is partial and spuriously ordered (rack U-number grids, SFP/QSF labels leak into spec text) — the parser will catch noise; filter obvious non-spec tokens. - Never dump an extracted .txt into the LLM whole. Parse to structure, then summarize. - Preserve the original text file and JSON next to the summary .md — they're the auditable source for the summary. ## Verification - After parsing a spec, spot-check 1-2 rows against the original PDF page (`read_file` at the known line) before trusting counts. - After any vision→drawio sheet, open the .drawio (or re-parse) to confirm valid XML / expected node count.