3.1 KiB
Spec/структура parser — engineering PDF text layer
Pattern that worked on the wine-city СКС project (Спецификация оборудования ТШ шкафов из 90-page DWG-render PDF, 384KB extracted text). Reusable for any "оборудование по шкафам/узлам" spec table whose text layer is radially extracted.
Goal
Turn the extracted .txt (with ===== СТРАНИЦА N ===== page markers) into structured items WITHOUT reading 2000 lines/spec into the LLM.
Heuristic-обозначение parser
import re
def parse_spec(lines, start, end):
"""items = [ {oboz, name, qty}, ... ]"""
items, i = [], start
while i < end:
s = lines[i-1].strip()
if not s:
i += 1; continue
# обозначение: one token, upper/alnum/dash/dot, >3 chars, not starting with digit
if re.fullmatch(r'[A-ZА-Я0-9][A-ZА-Я0-9\-.]+', s) and len(s) > 3 and not s[0].isdigit():
cur = {"oboz": s, "name": [], "qty": None}
j = i + 1; name_lines = []
while j < end and len(name_lines) < 8: # cap: наименование ≤8 строк
t = lines[j-1].strip()
if re.fullmatch(r'[A-ZА-Я0-9][A-ZА-Я0-9\-.]+', t) and len(t) > 3 and not t[0].isdigit():
break # next обозначение
if re.fullmatch(r'\d+', t):
cur["qty"] = int(t); j += 1; break # qty = bare int
if t and not re.fullmatch(r'\d+', t):
name_lines.append(t)
j += 1
cur["name"] = " ".join(name_lines)[:140]
items.append(cur); i = j
else:
i += 1
return items
Known noise to filter (DWG-render text layer)
- Rack U-number grids: long runs of bare ints (1..47) — compiler may capture as qty=None items; ignore
qty is Nonerows unless they carry a real name (they won't: name stays empty). - SFP1..4 / QSFP1..6 / "резерв" / "220V Port 1..." label rows leak in from switch faceplate text. Drop tokens in {SFP*, QSFP*} and empty-name items.
- Real spec items are the ones with BOTH a multi-char обозначение AND a non-empty name AND a qty — keep those as the equipment table.
Diff across nearly-identical sections (e.g. ТШ-A.1..A.4)
Since соседние spec-блоки почти одинаковы, don't re-read each: parse all to JSON, then compare counts per (oboz, name) to spot the deltas (A.1 vs A.2 vs A.3 vs A.4) and only narrate what differs.
Finding section boundaries without full reads
Scan the .txt lines for header regexes once (cheap): СПЕЦИФИКАЦИЯ, ВЕДОМОСТЬ, СХЕМА, ОБЩИЕ ДАННЫЕ, СОДЕРЖАНИЕ; record their line numbers, then read_file(offset=…) ONLY the interesting ranges.
Verification
Before trusting counts, read_file 1-2 rows at their known line and match против исходной таблицы on the actual PDF page (page.get_text() for the page). Spec numbers are consumable-accurate (закупочные количества) — worth the check.