Files
2026-09-06 13:51:20 +00:00

6.0 KiB

name, description, version, author, license, platforms, metadata
name description version author license platforms metadata
project-doc-processing Process large engineering PDF/DWG sets one file at a time. 1.0.0 Hermes Agent MIT
linux
macos
windows
hermes
tags related_skills
PDF
DWG
DrawIO
CAD
TokenEconomy
DocumentSet
Vision
Engineering
ocr-and-documents
file-tree-catalog
pdf
docx

Project Document Set Processing (token-economy)

For processing a large set of engineering/design documents (СКС/ЛВС/КТСБ drawings, CAD-derived PDFs, architect sets) one file at a time, WITHOUT burning tokens by dumping whole docs into LLM context. Complements ocr-and-documents (which is single-PDF text extraction) and file-tree-catalog (which catalogs a tree). This skill is about the processing loop for a recurring, multi-session project.

Core principle

The LLM should never read a full large PDF. Extract text to a file ONCE, then have code find structure and pull ONLY the slices worth human/LLM attention.

Workflow

1. Find a clean list of files to process

Catalog the tree first (see file-tree-catalog skill). Filter junk (Thumbs.db, ~$*, .dwl) at ANY nesting depth. The user processes one file at a time — never batch-spend tokens across many files.

2. Get a PDF tool into an isolated venv

System Python is often PEP-668 externally-managed (pip install fails without --break-system-packages). Do NOT touch the Hermes venv or system site-packages.

python3 -m venv /path/to/project/.venv-pdf
/path/to/project/.venv-pdf/bin/pip install -q pymupdf

Use pymupdf first — instant, no models. Only reach for marker-pdf (OCR, ~3-5GB models) if there's NO text layer.

3. Detect the text layer before choosing extractor

Big engineering PDFs are often DWG renders / scans with a partial text layer. Probe cheaply first:

import pymupdf
doc = pymupdf.open(path)
print(doc.page_count, "pages")
for i in range(min(6, doc.page_count)):
    t = doc[i].get_text().strip()
    print(f"page {i+1}: {len(t)} chars: {t[:80]!r}")

If pages have real text → pymupdf text extraction is enough. If empty → it's a scan: either render + vision (few pages, doc[i].get_pixmap()) or marker-pdf (bulk OCR). 20MB PDF with 90 pages / 270K chars of text layer still extracts fine — size alone does not mean "scanned".

4. Extract full text ONCE to a file, with page markers

out = "name.txt"
with open(out, "w", encoding="utf-8") as f:
    for i, page in enumerate(doc):
        f.write(f"\n\n===== СТРАНИЦА {i+1} =====\n")
        f.write(page.get_text())

This gives you a paginated, greppable corpus. Never paste the whole thing into the LLM.

5. Parse structure programmatically, don't read linearly

Use execute_code/search_files to find section markers (СПЕЦИФИКАЦИЯ, ВЕДОМОСТЬ, СХЕМА, ОБЩИЕ ДАННЫЕ, СОДЕРЖАНИЕ) and extract only the important ranges by line number via read_file(offset=...). For spec tables, write a small parser that produces structured data (JSON) and diffs/condenses it — far cheaper than reading thousands of lines:

# heuristic: обозначение = whitespace-free upper/digit/dash string >3 chars, name until next marker or a bare int (qty)

Then store the structured result (JSON) and write a concise summary .md for the record.

6. Record progress as editable status checkboxes

Keep a _конспект.md (summary) per processed file with:

  • object / customer / GIP / revision
  • doc composition (ведомость чертежей)
  • the "what next / unread" list
  • a ## Статус обработки section with [x]/[ ] checkboxes (which files done, which pending) This makes a multi-session project resumable: next session reads the status file and continues.

Schematics → draw.io rebuild (design intent)

When the goal is to reconstruct a schematic (structural diagram, rack/ТШ layout, cable-run diagram), do NOT have the LLM write .drawio XML directly (it produces broken XML). Instead:

  1. Render the schematic page to PNG: page.get_pixmap(dpi=150).save("/tmp/page.png").
  2. Send the PNG to a vision model describing nodes + edges.
  3. Have vision output a structured graph (JSON: nodes with labels, edges), then serialize to .drawio (mxGraphModel) in code.

Honest capability table (set user expectations up front):

Schematic type draw.io rebuild
Structural diagram (nodes+links) ✅ good — ideal candidate
Rack/ТШ layout (grid) ✅ good (spec already parsed → grid)
Cable-run / fiber-splice diagram ⚠️ medium — many tiny labels
Equipment/route layout plan ⚠️ schematic only, NOT to scale
Isometric (axonometric) runs ❌ poor — 3D iso doesn't fit draw.io

Caveats to state:

  • Local vision (qwen3-vl) reads diagrams but errs more on dense small-label sheets than cloud vision — local=privacy+free, cloud=accuracy. Hybrid allowed.
  • Rendering DWG-render PDF → PNG loses geometry; output is a semantic diagram (right elements/links), not a pixel-exact copy.
  • If pixel-exact matters, the real source is the .dwg CAD files (needs DXF/ODA conversion — heavier). For "understand & rebuild", vision is enough.
  • Test on ONE simple representative sheet before scaling.

Pitfalls

  • PEP-668: use a dedicated venv, never pollute system/Hermes site-packages.
  • A 20MB PDF is not necessarily a scan — ALWAYS probe text layer first.
  • pymupdf text layer from DWG renders is partial and spuriously ordered (rack U-number grids, SFP/QSF labels leak into spec text) — the parser will catch noise; filter obvious non-spec tokens.
  • Never dump an extracted .txt into the LLM whole. Parse to structure, then summarize.
  • Preserve the original text file and JSON next to the summary .md — they're the auditable source for the summary.

Verification

  • After parsing a spec, spot-check 1-2 rows against the original PDF page (read_file at the known line) before trusting counts.
  • After any vision→drawio sheet, open the .drawio (or re-parse) to confirm valid XML / expected node count.