Initial commit: Hermes skill project-doc-processing

This commit is contained in:
estorozhenko
2026-09-06 13:51:20 +00:00
commit d53e4b88f6
4 changed files with 214 additions and 0 deletions
+50
View File
@@ -0,0 +1,50 @@
# Spec/структура parser — engineering PDF text layer
Pattern that worked on the wine-city СКС project (Спецификация оборудования ТШ шкафов из 90-page DWG-render PDF, 384KB extracted text). Reusable for any "оборудование по шкафам/узлам" spec table whose text layer is radially extracted.
## Goal
Turn the extracted .txt (with `===== СТРАНИЦА N =====` page markers) into structured items WITHOUT reading 2000 lines/spec into the LLM.
## Heuristic-обозначение parser
```python
import re
def parse_spec(lines, start, end):
"""items = [ {oboz, name, qty}, ... ]"""
items, i = [], start
while i < end:
s = lines[i-1].strip()
if not s:
i += 1; continue
# обозначение: one token, upper/alnum/dash/dot, >3 chars, not starting with digit
if re.fullmatch(r'[A-ZА-Я0-9][A-ZА-Я0-9\-.]+', s) and len(s) > 3 and not s[0].isdigit():
cur = {"oboz": s, "name": [], "qty": None}
j = i + 1; name_lines = []
while j < end and len(name_lines) < 8: # cap: наименование ≤8 строк
t = lines[j-1].strip()
if re.fullmatch(r'[A-ZА-Я0-9][A-ZА-Я0-9\-.]+', t) and len(t) > 3 and not t[0].isdigit():
break # next обозначение
if re.fullmatch(r'\d+', t):
cur["qty"] = int(t); j += 1; break # qty = bare int
if t and not re.fullmatch(r'\d+', t):
name_lines.append(t)
j += 1
cur["name"] = " ".join(name_lines)[:140]
items.append(cur); i = j
else:
i += 1
return items
```
## Known noise to filter (DWG-render text layer)
- Rack U-number grids: long runs of bare ints (1..47) — compiler may capture as qty=None items; ignore `qty is None` rows unless they carry a real name (they won't: name stays empty).
- SFP1..4 / QSFP1..6 / "резерв" / "220V Port 1..." label rows leak in from switch faceplate text. Drop tokens in {SFP*, QSFP*} and empty-name items.
- Real spec items are the ones with BOTH a multi-char обозначение AND a non-empty name AND a qty — keep those as the equipment table.
## Diff across nearly-identical sections (e.g. ТШ-A.1..A.4)
Since соседние spec-блоки почти одинаковы, don't re-read each: parse all to JSON, then compare counts per (oboz, name) to spot the deltas (A.1 vs A.2 vs A.3 vs A.4) and only narrate what differs.
## Finding section boundaries without full reads
Scan the .txt lines for header regexes once (cheap): `СПЕЦИФИКАЦИЯ`, `ВЕДОМОСТЬ`, `СХЕМА`, `ОБЩИЕ ДАННЫЕ`, `СОДЕРЖАНИЕ`; record their line numbers, then `read_file(offset=…)` ONLY the interesting ranges.
## Verification
Before trusting counts, `read_file` 1-2 rows at their known line and match против исходной таблицы on the actual PDF page (`page.get_text()` for the page). Spec numbers are consumable-accurate (закупочные количества) — worth the check.