6.0 KiB
name, description, version, author, license, platforms, metadata
| name | description | version | author | license | platforms | metadata | |||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| project-doc-processing | Process large engineering PDF/DWG sets one file at a time. | 1.0.0 | Hermes Agent | MIT |
|
|
Project Document Set Processing (token-economy)
For processing a large set of engineering/design documents (СКС/ЛВС/КТСБ drawings, CAD-derived PDFs, architect sets) one file at a time, WITHOUT burning tokens by dumping whole docs into LLM context. Complements ocr-and-documents (which is single-PDF text extraction) and file-tree-catalog (which catalogs a tree). This skill is about the processing loop for a recurring, multi-session project.
Core principle
The LLM should never read a full large PDF. Extract text to a file ONCE, then have code find structure and pull ONLY the slices worth human/LLM attention.
Workflow
1. Find a clean list of files to process
Catalog the tree first (see file-tree-catalog skill). Filter junk (Thumbs.db, ~$*, .dwl) at ANY nesting depth. The user processes one file at a time — never batch-spend tokens across many files.
2. Get a PDF tool into an isolated venv
System Python is often PEP-668 externally-managed (pip install fails without --break-system-packages). Do NOT touch the Hermes venv or system site-packages.
python3 -m venv /path/to/project/.venv-pdf
/path/to/project/.venv-pdf/bin/pip install -q pymupdf
Use pymupdf first — instant, no models. Only reach for marker-pdf (OCR, ~3-5GB models) if there's NO text layer.
3. Detect the text layer before choosing extractor
Big engineering PDFs are often DWG renders / scans with a partial text layer. Probe cheaply first:
import pymupdf
doc = pymupdf.open(path)
print(doc.page_count, "pages")
for i in range(min(6, doc.page_count)):
t = doc[i].get_text().strip()
print(f"page {i+1}: {len(t)} chars: {t[:80]!r}")
If pages have real text → pymupdf text extraction is enough. If empty → it's a scan: either render + vision (few pages, doc[i].get_pixmap()) or marker-pdf (bulk OCR). 20MB PDF with 90 pages / 270K chars of text layer still extracts fine — size alone does not mean "scanned".
4. Extract full text ONCE to a file, with page markers
out = "name.txt"
with open(out, "w", encoding="utf-8") as f:
for i, page in enumerate(doc):
f.write(f"\n\n===== СТРАНИЦА {i+1} =====\n")
f.write(page.get_text())
This gives you a paginated, greppable corpus. Never paste the whole thing into the LLM.
5. Parse structure programmatically, don't read linearly
Use execute_code/search_files to find section markers (СПЕЦИФИКАЦИЯ, ВЕДОМОСТЬ, СХЕМА, ОБЩИЕ ДАННЫЕ, СОДЕРЖАНИЕ) and extract only the important ranges by line number via read_file(offset=...). For spec tables, write a small parser that produces structured data (JSON) and diffs/condenses it — far cheaper than reading thousands of lines:
# heuristic: обозначение = whitespace-free upper/digit/dash string >3 chars, name until next marker or a bare int (qty)
Then store the structured result (JSON) and write a concise summary .md for the record.
6. Record progress as editable status checkboxes
Keep a _конспект.md (summary) per processed file with:
- object / customer / GIP / revision
- doc composition (ведомость чертежей)
- the "what next / unread" list
- a
## Статус обработкиsection with[x]/[ ]checkboxes (which files done, which pending) This makes a multi-session project resumable: next session reads the status file and continues.
Schematics → draw.io rebuild (design intent)
When the goal is to reconstruct a schematic (structural diagram, rack/ТШ layout, cable-run diagram), do NOT have the LLM write .drawio XML directly (it produces broken XML). Instead:
- Render the schematic page to PNG:
page.get_pixmap(dpi=150).save("/tmp/page.png"). - Send the PNG to a vision model describing nodes + edges.
- Have vision output a structured graph (JSON: nodes with labels, edges), then serialize to
.drawio(mxGraphModel) in code.
Honest capability table (set user expectations up front):
| Schematic type | draw.io rebuild |
|---|---|
| Structural diagram (nodes+links) | ✅ good — ideal candidate |
| Rack/ТШ layout (grid) | ✅ good (spec already parsed → grid) |
| Cable-run / fiber-splice diagram | ⚠️ medium — many tiny labels |
| Equipment/route layout plan | ⚠️ schematic only, NOT to scale |
| Isometric (axonometric) runs | ❌ poor — 3D iso doesn't fit draw.io |
Caveats to state:
- Local vision (qwen3-vl) reads diagrams but errs more on dense small-label sheets than cloud vision — local=privacy+free, cloud=accuracy. Hybrid allowed.
- Rendering DWG-render PDF → PNG loses geometry; output is a semantic diagram (right elements/links), not a pixel-exact copy.
- If pixel-exact matters, the real source is the
.dwgCAD files (needs DXF/ODA conversion — heavier). For "understand & rebuild", vision is enough. - Test on ONE simple representative sheet before scaling.
Pitfalls
- PEP-668: use a dedicated venv, never pollute system/Hermes site-packages.
- A 20MB PDF is not necessarily a scan — ALWAYS probe text layer first.
- pymupdf text layer from DWG renders is partial and spuriously ordered (rack U-number grids, SFP/QSF labels leak into spec text) — the parser will catch noise; filter obvious non-spec tokens.
- Never dump an extracted .txt into the LLM whole. Parse to structure, then summarize.
- Preserve the original text file and JSON next to the summary .md — they're the auditable source for the summary.
Verification
- After parsing a spec, spot-check 1-2 rows against the original PDF page (
read_fileat the known line) before trusting counts. - After any vision→drawio sheet, open the .drawio (or re-parse) to confirm valid XML / expected node count.